
Principal Infrastructure Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Netherlands.
• Take ownership of the technical architecture and advancement of Sezzle’s core infrastructure.
• Recognize system limitations, prioritize enhancements, and facilitate the handling of increased traffic, data volume, and workload complexity.
• Link business workflows to infrastructure enhancements across applications, data, and systems.
• Create capacity models, conduct load and stress testing, identify bottlenecks, and verify improvements in throughput, latency, saturation, and cost.
• Design and construct robust AWS account, IAM, network, and service architectures.
• Develop and manage the Kubernetes platform, including lifecycle automation, workload isolation, resource allocation, autoscaling, upgrades, and deployment reliability.
• Scale and optimize Aurora RDS for MySQL and Postgres, focusing on queries, indexes, connections, replication, failover, schema changes, and migrations.
• Establish and instrument service-level objectives and error budgets; implement failure isolation, backpressure, load shedding, and safe retries.
• Engage in on-call duties and lead technical recovery during significant incidents and outages.
• Design, test, and document disaster recovery, backup, restore, and failover procedures.
• Create infrastructure-as-code and operational automation for provisioning, configuration, deployment, upgrades, and recovery.
• Enhance observability through metrics, logs, traces, dashboards, and actionable alerts.
• Ensure safe infrastructure migrations through phased rollouts, validation, and rollback strategies.
• Boost cloud cost efficiency by optimizing resource sizing, utilization, autoscaling, and storage management.
• Develop and assess AI-assisted tools for incident investigation, runbooks, anomaly detection, and repetitive tasks.
• Draft architecture proposals, assess technology trade-offs through prototypes and benchmarks, and document system behavior and failure scenarios.
• Report to engineering leadership and work collaboratively with application engineering, Security, and Compliance teams.
• Bachelor’s degree in Computer Science or a related technical field (required).
• Over 12 years of experience in infrastructure, platform, site reliability, software development, or related engineering areas.
• Extensive production knowledge of AWS, encompassing compute, IAM, multi-account architectures, networking, VPC design, and private connectivity.
• Profound expertise in Kubernetes, including cluster lifecycle, scheduling, resource management, autoscaling, networking, and troubleshooting.
• In-depth knowledge of RDS/Aurora for MySQL and/or Postgres, including query performance, indexing, connection management, replication, high availability, failover, and backup/recovery.
• Proven history of delivering infrastructure scaling improvements personally.
• Strong coding and automation capabilities using Golang, Python, or similar programming languages.
• Experience with infrastructure-as-code tools such as Terraform or equivalents.
• Solid foundational knowledge in Linux, networking, DNS, TLS, storage, concurrency, and distributed-system failure modes.
• Experience managing a 24/7 high-availability platform with hands-on incident response and postmortem analysis.
• Willingness to be part of an on-call rotation.
• Experience in implementing and testing disaster recovery against specified recovery objectives.
• Practical knowledge of observability, load testing, capacity planning, and safe CI/CD practices.
• Active engagement with AI tools in engineering or operations.
• Ability to navigate ambiguous technical challenges through to production delivery and collaborate across engineering teams.
• Preferred: EKS, experience in fintech/payments/banking, multi-region architectures, chaos engineering, Prometheus, Grafana, Loki, Tempo, internal platform capabilities, self-service tools, and AI-assisted operational automation.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Generous paid time off and flexible work arrangements.
• Opportunities for professional development and continuous learning.
• Supportive company culture that values diversity and inclusion.
Sezzle
HavocAI
QuickNode ⚡
Sezzle
Get handpicked remote jobs straight to your inbox weekly.