
Principal Infrastructure Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Mexico.
• Design, develop, manage, and enhance Sezzle’s infrastructure platform.
• Take ownership of intricate technical projects from architecture and prototyping to implementation, deployment, and ongoing management.
• Recognize system constraints and implement enhancements to handle traffic, data volume, and workload complexity.
• Create capacity models, conduct load and stress tests, identify bottlenecks, and confirm improvements in throughput, latency, saturation, and costs.
• Architect resilient AWS accounts, IAM configurations, networks, and service architectures.
• Construct and maintain Kubernetes clusters, including lifecycle automation, workload segregation, resource allocation, autoscaling, upgrades, and deployment reliability.
• Scale and optimize Aurora RDS for MySQL and Postgres, focusing on queries, indexes, connections, replication, failover, schema modifications, and migrations.
• Establish service-level objectives and error budgets; implement failure isolation, backpressure, load shedding, and safe retry mechanisms.
• Engage in on-call duties and lead technical recovery efforts during critical incidents and outages.
• Design and assess disaster recovery strategies, backups, restores, and failover procedures to meet recovery objectives.
• Develop infrastructure-as-code and operational automation solutions.
• Enhance observability through metrics, logs, traces, dashboards, and alert systems.
• Facilitate safe infrastructure migrations via phased rollouts, validation, and rollback strategies.
• Optimize cloud cost efficiency through resource sizing, utilization, autoscaling, and storage optimization.
• Develop and assess AI-assisted operational tools.
• Draft architectural proposals, evaluate technology trade-offs, review shared-infrastructure modifications, and document system operations and failure modes.
• Collaborate with application engineering, security, compliance, and engineering leadership teams.
• A Bachelor's degree in Computer Science or a related technical discipline (mandatory).
• Over 12 years of experience in infrastructure, platform, site reliability, software development, or equivalent engineering fields.
• Extensive production experience with AWS, including compute, IAM, multi-account architectures, VPC design, and private connectivity.
• Significant production expertise with Kubernetes; experience with EKS is highly preferred.
• In-depth knowledge of RDS/Aurora for MySQL and/or Postgres.
• Proven track record of implementing infrastructure scaling improvements.
• Proficient coding and automation abilities in Golang, Python, or similar programming languages.
• Experience with infrastructure-as-code using Terraform or equivalent tools.
• Strong understanding of systems fundamentals including Linux, networking, DNS, TLS, storage, concurrency, and distributed-system failure modes.
• Experience in operating 24/7 high-availability platforms.
• Willingness to be part of an on-call rotation.
• Experience in implementing and testing disaster recovery plans.
• Hands-on experience with observability, load testing, capacity planning, and safe CI/CD practices.
• Active engagement with AI tools in engineering or operations.
• Ability to navigate ambiguous technical challenges through production delivery and collaborate across various engineering disciplines.
• Preferred: experience in fintech, payments, or banking; knowledge of multi-region architectures; chaos engineering; familiarity with Prometheus, Grafana, Loki, Tempo; internal platform capabilities; and AI-assisted incident investigation or operational automation.
• Competitive gross monthly compensation ranging from $12,500 to $20,800 USD.
• An engineering environment focused on open-source development.
• Opportunity for remote work within Mexico / Latin America.
VALR
First Circle
crewAI
Sezzle
Get handpicked remote jobs straight to your inbox weekly.