
Principal Infrastructure Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Brazil.
• Design, construct, manage, and enhance Sezzle’s infrastructure platform.
• Take ownership of intricate technical projects from architecture and prototyping to implementation, deployment, and ongoing management.
• Recognize system limitations and implement enhancements for traffic, data volume, workload complexity, performance, reliability, and cost-effectiveness.
• Create capacity models, conduct load and stress testing, and troubleshoot compute, networking, Kubernetes, and database bottlenecks.
• Develop resilient AWS account, IAM, network, and service architectures.
• Construct and manage Kubernetes platform architecture, lifecycle automation, workload isolation, resource allocation, autoscaling, upgrades, and deployment reliability.
• Scale and optimize Aurora RDS for MySQL and Postgres, focusing on queries, indexes, connections, replication, failover, schema modifications, and migrations.
• Define and measure service-level objectives and error budgets; implement failure isolation, backpressure, load shedding, and safe retry mechanisms.
• Engage in on-call rotations and lead technical recovery efforts during critical incidents and outages.
• Design, test, and document disaster recovery, backup, restore, and failover strategies.
• Develop infrastructure-as-code and operational automation for provisioning, configuration, deployments, upgrades, and recovery.
• Enhance observability through metrics, logs, traces, dashboards, and actionable alerts.
• Ensure safe infrastructure migrations through phased rollouts, validation, and rollback strategies.
• Increase cloud cost efficiency through resource right-sizing, utilization, autoscaling, storage optimization, and quantified savings.
• Build and assess AI-assisted tools for incident investigation, runbooks, anomaly analysis, and repetitive tasks.
• Draft architecture proposals, evaluate technology trade-offs, review shared infrastructure changes, and document system operations and failures.
• Report to engineering leadership while collaborating with application engineers, Security, and Compliance.
• Bachelor's degree in Computer Science or a related technical discipline (required).
• Over 12 years of experience in infrastructure, platform, site reliability, software development, or comparable engineering fields.
• Extensive production experience with AWS, covering compute, IAM, multi-account architectures, networking, VPC design, and private connectivity.
• Profound production expertise with Kubernetes; EKS strongly preferred.
• In-depth knowledge of RDS/Aurora for MySQL and/or Postgres at scale.
• Proven history of delivering infrastructure scaling enhancements personally.
• Strong coding and automation proficiency in Golang, Python, or similar programming languages.
• Experience with infrastructure-as-code using Terraform or a comparable tool.
• Solid understanding of systems fundamentals, including Linux, networking, DNS, TLS, storage, concurrency, and distributed-system failure modes.
• Experience in operating a 24/7 high-availability platform with a direct impact on customers or revenue.
• Willingness to participate in an on-call rotation and recover production systems under pressure.
• Experience in implementing and testing disaster recovery in line with defined recovery objectives.
• Hands-on experience with observability, load testing, capacity planning, and safe CI/CD practices.
• Active utilization of AI tools in engineering or operations.
• Ability to navigate ambiguous technical challenges through to production delivery while collaborating across engineering disciplines.
• Preferred: experience in fintech, payments, or banking; familiarity with multi-region architectures; chaos engineering; Prometheus, Grafana, Loki, Tempo; internal platform capabilities; AI-assisted operational automation.
• Competitive gross monthly salary ranging from $12,500 to $20,800 USD, depending on location and experience level.
• Open-source-focused technology environment.
VALR
First Circle
crewAI
Sezzle
Get handpicked remote jobs straight to your inbox weekly.