
Principal Infrastructure Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Argentina.
• Design, construct, manage, and scale the platform that powers Sezzle, a fintech company specializing in interest-free installment payments and shopping technology.
• Take ownership of intricate infrastructure projects from architecture and prototyping through to implementation, production rollout, and ongoing maintenance.
• Analyze system limitations and implement enhancements for traffic, data volume, workload complexity, throughput, latency, resilience, and cost-effectiveness.
• Create robust AWS account, IAM, network, and service architectures, including multi-AZ or multi-region capabilities.
• Develop and manage the Kubernetes platform, encompassing lifecycle automation, workload isolation, resource allocation, autoscaling, upgrades, and deployment reliability.
• Scale and optimize Aurora RDS for MySQL and Postgres through query and index tuning, connection and replication enhancements, capacity planning, failover strategies, schema modifications, and migrations.
• Define and instrument service-level objectives and error budgets; implement failure isolation, backpressure, load shedding, and safe retries.
• Participate in on-call rotations and lead technical recovery efforts during critical incidents and outages.
• Design, test, and document disaster-recovery backup, restore, and failover procedures.
• Develop infrastructure-as-code and operational automation for provisioning, configuration, deployments, upgrades, and recovery processes.
• Enhance observability through metrics, logs, traces, dashboards, and actionable alerts.
• Execute safe infrastructure migrations with phased rollouts, validation, compatibility checks, and rollback strategies.
• Optimize cloud cost efficiency through resource right-sizing, utilization enhancements, autoscaling, and storage tuning.
• Develop and assess AI-assisted tools for incident investigation, runbooks, anomaly analysis, and toil reduction.
• Compose architecture proposals, evaluate technology trade-offs through prototypes and benchmarks, review shared-infrastructure changes, and document system operations and failures.
• Collaborate with application engineers, Security, Compliance, and engineering leadership to align infrastructure decisions with customer and business outcomes.
• Bachelor's degree in Computer Science or a related technical field (required).
• Over 12 years of experience in infrastructure, platform, site reliability, software development, or related engineering domains.
• Extensive production expertise with AWS, including compute, IAM, multi-account architectures, VPC design, and private connectivity.
• Profound production experience with Kubernetes; EKS experience is highly preferred.
• Significant expertise with RDS/Aurora MySQL and/or Postgres at scale.
• Proven track record of personally delivering infrastructure scaling enhancements.
• Strong coding and automation abilities using Golang, Python, or similar programming languages.
• Experience with infrastructure-as-code using Terraform or equivalent tools.
• Robust systems knowledge in Linux, networking, DNS, TLS, storage, concurrency, and distributed system failure modes.
• Experience managing a 24/7 high-availability platform.
• Willingness and capability to participate in an on-call rotation and address production incidents.
• Experience in implementing and testing disaster recovery against specified recovery objectives.
• Familiarity with observability, load testing, capacity planning, and safe CI/CD practices.
• Active engagement with AI tooling in engineering or operational contexts.
• Ability to navigate ambiguous technical challenges through production delivery and collaborate across engineering disciplines.
• Preferred: experience in fintech, payments, or banking; multi-region architectures; chaos engineering; tools such as Prometheus, Grafana, Loki, Tempo; internal platform and self-service tooling; AI-assisted incident investigation or operational automation.
• Competitive gross monthly compensation of $12,500–$20,800 USD, depending on location and experience.
• Remote work opportunity in Argentina / Latin America.
• Open-source-focused engineering environment.
VALR
First Circle
crewAI
Sezzle
Get handpicked remote jobs straight to your inbox weekly.