The salary range for this role is negotiable, the range being $400,000 - $600,000 per year in total compensation.
About Sezzle:
With a mission to financially empower the next generation, Sezzle is revolutionizing the shopping experience beyond payments, blending cutting-edge tech with seamless, interest-free installment plans that make shopping smarter and more accessible. We’re not just transforming payments; we’re redefining how people discover, interact with, and purchase the things they love while driving real impact on merchant sales through increased conversions and higher order values. As we continue to shape the future of fintech and retail, we’re building an innovative, dynamic team passionate about creating more than just a transaction but a truly unique shopping journey. If you’re excited about pushing boundaries in tech and delivering a game-changing experience for consumers and merchants alike, come join us at Sezzle and help create the future of shopping!
About the Role:
We are seeking an exceptional VP Engineering - Infrastructure & SRE to own the strategy, reliability, security posture, and evolution of the platform that powers Sezzle. This is a rare opportunity to lead infrastructure through a defining chapter: scaling a high-growth fintech while raising our operational rigor, resilience, and compliance posture to the standards of the most demanding financial institutions.
This is explicitly both a leadership job and a technical one. You will lead the teams responsible for our cloud infrastructure, Kubernetes platform, databases, networking, observability, and site reliability engineering, and you will stay hands-on in the systems yourself, working alongside the engineers you lead. Our stack runs on AWS, with workloads orchestrated on Kubernetes and data anchored in Aurora RDS (MySQL and Postgres). You should know these technologies deeply, not just manage people who do. You will set the technical vision, own the budget, and be directly accountable for availability, disaster recovery, and business continuity across a payments platform where downtime has immediate customer and financial impact.
That accountability is personal, not just organizational: you will manage the on-call rotation and take shifts in it. When Sezzle experiences a serious incident, up to and including a full outage, you are the kind of leader who can step into command, cut through noise with evidence-based triage, make fast, sound decisions under pressure with incomplete information, and drive the platform back to health. Your team should trust you at 3 AM as much as they do in a roadmap review.
Just as importantly, you will lead our infrastructure organization into an AI-boosted SRE era. We believe the next generation of infrastructure teams will pair strong engineers with AI agents and tooling for incident response, capacity planning, runbook automation, anomaly detection, and day-to-day operational toil. You won't just tolerate this shift; you'll drive it, with conviction and hands-on credibility.
This role reports to senior engineering leadership and partners closely with Security, Compliance, Engineering, Finance, and external auditors. Your work will directly shape whether Sezzle can operate at the highest standards of reliability and compliance without losing the speed and pragmatism of a startup.
Key Responsibilities
- Own the infrastructure vision, strategy, and multi-year roadmap: scale today's high-growth fintech platform while continually strengthening its resilience, controls, and audit-readiness.
- Lead, grow, and mentor the infrastructure, platform, and SRE organization, including hiring, career development, on-call health, and building a culture of operational excellence and blameless learning.
- Own reliability end-to-end: define and enforce SLOs and error budgets, mature incident management and postmortem practices, and be accountable for platform availability across the business.
- Manage, and participate in, the on-call rotation, and serve as senior incident commander for high-severity events: leading recovery from major degradations and full outages through rapid, evidence-based triage, decisive action under uncertainty, and clear communication to stakeholders throughout.
- Direct our AWS strategy, including account architecture, IAM and network design, multi-AZ/multi-region posture, service selection, and cost management (FinOps). You will own and defend the cloud budget.
- Own the Kubernetes platform as a product: cluster architecture, upgrade strategy, workload isolation, autoscaling, progressive delivery, and the developer experience of every team that ships on it.
- Own the database tier, centered on Aurora RDS (MySQL and Postgres): availability, performance, capacity, schema and migration safety practices, backup/restore verification, and encryption.
- Design, implement, and continuously test disaster recovery and business continuity: defined RTO/RPO targets per system tier, regular game days and failover exercises, and DR evidence that stands up to auditor scrutiny.
- Champion the AI-boosted SRE transformation: evaluate and deploy AI tooling and agents for incident triage, observability, runbook automation, and toil reduction; set standards for safe, auditable use of AI in production operations; and bring the team along through training and example.
- Partner with Security and Compliance to own infrastructure's role in PCI-DSS and SOC 2: control design and operation, evidence collection, segmentation, vulnerability and patch management, and audit support, with the maturity to meet the expectations of banking partners and financial-industry examinations.
- Drive infrastructure-as-code and platform automation as the default: everything reproducible, reviewed, and recoverable; nothing artisanal.
- Own vendor and technology strategy for the infrastructure domain: build-vs-buy decisions, vendor risk management, contract negotiation, and third-party resilience.
- Communicate crisply with executives, the board, and auditors, translating infrastructure risk, investment, and posture into business terms.
Minimum Requirements:
- 15+ years of combined experience across infrastructure, platform, site reliability, software development, or related engineering disciplines, with substantial depth in infrastructure, including 5+ years leading engineering teams.
- Deep, hands-on expertise with AWS: you have designed and operated production architectures across compute, networking (VPC design, Transit Gateway, PrivateLink), IAM, and multi-account organizations at scale.
- Deep, hands-on expertise with Kubernetes in production: cluster lifecycle management, workload architecture, scaling, and the operational realities of running business-critical services on it (EKS experience strongly preferred).
- Deep expertise with relational databases at scale, specifically RDS/Aurora (MySQL and/or Postgres): high availability, replication, failover, performance tuning, and backup/recovery you have personally verified under pressure.
- Proven ownership of disaster recovery and business continuity for a production platform: you have defined RTO/RPO targets, built the capability to meet them, and run real failover tests, not just written the document.
- Demonstrated AI-forward leadership: you actively use AI tooling in engineering or operations work today, have opinions grounded in practice about where it helps and where it doesn't, and have led (or are visibly leading) a team's adoption of AI-assisted workflows.
- Track record of operating a 24/7, high-availability platform where downtime has direct revenue or customer impact, including mature incident command and postmortem practices.
- Willingness to manage and participate in an on-call rotation, and demonstrated ability to lead recovery from a full production outage: forming and testing hypotheses from logs, metrics, and traces rather than guesswork, making the right call quickly with incomplete information, and knowing when to mitigate first and root-cause later.
- Still technical, by choice: you remain a credible hands-on engineer, comfortable in a terminal, reading dashboards, and reviewing designs, and you expect to stay that way. You will lead the team and work alongside it; this is not a delegation-only role.
- Experience owning significant cloud budgets and driving cost efficiency without sacrificing reliability.
- Strong grounding in infrastructure-as-code (Terraform or equivalent) and modern CI/CD practices.
- Demonstrated ability to hire, develop, and retain strong infrastructure and SRE talent, and to hold a high bar through growth.
- Bachelor's degree in Computer Science or a similar technical field (required).
Preferred Knowledge and Skills:
- Direct experience supporting PCI-DSS and SOC 2 programs from the infrastructure side: scoping and segmentation, control ownership, evidence automation, and working sessions with assessors and auditors.
- Experience in fintech, payments, or banking, especially in environments with heightened regulatory expectations (bank partnerships/sponsorships, FFIEC examinations, GLBA, or similar).
- Experience deploying AIOps or LLM-based tooling in production operations, such as AI-assisted incident response, intelligent alerting, automated runbooks, or agents (e.g., Claude Code or custom LLM integrations) embedded in SRE workflows, with sensible guardrails around safety and auditability.
- Experience with multi-region and active-active architectures, chaos engineering, and formal operational resilience programs.
- Proficiency with modern observability stacks (Prometheus, Grafana, Loki, Tempo, or commercial equivalents) and driving observability as a platform capability.
- Familiarity with service mesh, zero-trust networking, secrets management, and workload identity patterns.
- Experience with CI/CD pipelines, progressive delivery (canary/blue-green), and platform engineering / internal developer platform approaches.
- Experience presenting to boards, auditors, or examiners, and building the documentation and evidence culture that makes those conversations easy.
About You:
- You have relentlessly high standards - many people may think your standards are unreasonably high. You are continually raising the bar and driving those around you to deliver great results. You make sure that defects do not get sent down the line and that problems are fixed so they stay fixed.
- You’re not bound by convention - your success—and much of the fun—lies in developing new ways to do things
- You need action - speed matters in business. Many decisions and actions are reversible and do not need extensive study. We value calculated risk-taking.
- You earn trust - you listen attentively, speak candidly, and treat others respectfully.
- You have backbone; disagree, then commit - you can respectfully challenge decisions when you disagree, even when doing so is uncomfortable or exhausting. You have conviction and are tenacious. You do not compromise for the sake of social cohesion. Once a decision is determined, you commit wholly.
- You deliver results - you focus on the key inputs and deliver them with the right quality and in a timely fashion. Despite setbacks, you rise to the occasion and never settle.
Sezzle’s Technology Stack:
- Languages: Golang, Typescript, Python
- Frontend: Typescript - React and React Native
- Backend: Golang
- Database: MySQL, Postgres
- DevOps & Cloud: AWS, Kubernetes
- Version Control: Git
- CI/CD: Gitlab
- Testing: Developer and AI-driven, focus on automated end-to-end, integration, and unit tests
- Open Source: Sezzle is focused on using open source, and we build what we can before buying!
What Makes Working at Sezzle Awesome?
At Sezzle, we are more than just brilliant engineers, passionate data enthusiasts, out-of-the-box thinkers, and determined innovators; we are skilled musicians, yogis, cyclists, chefs, golfers, dog-lovers, and rock-climbers. We believe in surrounding ourselves with not only the best and the brightest individuals, but those that are unique and purpose-driven in all that they do. Our culture is not defined by a certain set of perks designed to give the illusion of the traditional startup culture, but rather, it is the visible example living in every employee that we hire.
Perks & Benefits:
- Unlimited PTO, volunteer hours and sabbatical
- Life, STD/LTD, medical, dental and vision insurance
- Highly discounted LifeTime gym membership
- 401k with match
- Collaborative fun co-workers
- The opportunity to join the fastest growing FinTech alongside a team of motivated and driven individuals
CCPA Disclosure: Sezzle Inc. is committed to protecting the privacy of our job applicants. In compliance with the California Consumer Privacy Act (CCPA), we inform California residents about the personal information we may collect, the purposes for its collection, and your rights under the CCPA. For details about the categories of personal information we collect and your rights under the CCPA, please visit the California Office of the Attorney General's CCPA page. By submitting your application, you acknowledge that you have read and understood this CCPA disclosure.
Equal Employment Opportunity: Sezzle Inc. is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate based on race, color, religion, sex, national origin, age, disability, genetic information, pregnancy, or any other legally protected status. Sezzle recognizes and values the importance of diversity and inclusion in enriching the employment experience of its employees and in supporting our mission.
#Li-remote #full-time