
Principal Infrastructure Engineer
Posted 17 hours ago

Posted 17 hours ago
This is a fully remote position, open to applicants in France.
• Take ownership of the technical architecture and advancement of core infrastructure.
• Recognize system limitations and implement enhancements to support increased traffic, data volumes, and workload complexities.
• Link business workflows to their impacts across applications, data, and infrastructure.
• Create capacity models, conduct load and stress tests, diagnose bottlenecks, and validate improvements in throughput, latency, saturation, and costs.
• Design resilient architectures for AWS accounts, IAM, networks, and services.
• Develop and manage the Kubernetes platform, including lifecycle automation, workload isolation, autoscaling, upgrades, and deployment reliability.
• Scale and optimize Aurora RDS for MySQL and Postgres, focusing on queries, indexes, connections, replication, failover, schema changes, and migrations.
• Establish service-level objectives and error budgets; implement failure isolation, backpressure, load shedding, and safe retries.
• Engage in on-call duties and lead technical recovery during critical incidents.
• Design, test, and document disaster recovery, backup, restoration, and failover mechanisms.
• Construct infrastructure-as-code and automate operations.
• Enhance metrics, logs, traces, dashboards, and actionable alerts.
• Ensure safe infrastructure migrations with phased rollouts, validations, and rollback strategies.
• Boost cloud cost efficiency through resource sizing, utilization, autoscaling, and storage optimization.
• Develop and assess AI-assisted tools for incident investigation, runbooks, anomaly analysis, and toil reduction.
• Draft architecture proposals, assess technology tradeoffs, review shared infrastructure changes, and document system operations and failure modes.
• Collaborate with application engineering, Security, Compliance, and engineering leadership.
• A Bachelor's degree in Computer Science or a related technical field (required).
• Over 12 years of experience in infrastructure, platform, site reliability, software development, or related engineering disciplines.
• Extensive production experience with AWS, covering compute, IAM, multi-account architectures, VPC design, and private connectivity.
• Profound production knowledge of Kubernetes, including cluster lifecycle, scheduling, resource management, autoscaling, networking, and troubleshooting.
• In-depth expertise with RDS/Aurora MySQL and/or Postgres at scale.
• Proven track record of delivering infrastructure scaling improvements personally.
• Strong coding and automation capabilities using Golang, Python, or similar languages.
• Experience in infrastructure-as-code with Terraform or equivalent tools.
• Solid understanding of systems fundamentals in Linux, networking, DNS, TLS, storage, concurrency, and distributed system failure modes.
• Experience managing a 24/7 high-availability platform.
• Hands-on experience in incident response and postmortem remediation.
• Willingness to participate in an on-call rotation.
• Experience in implementing and testing disaster recovery against established recovery objectives.
• Practical knowledge of observability, load testing, capacity planning, and safe CI/CD practices.
• Active use of AI tools in engineering or operations.
• Capability to navigate ambiguous technical challenges from investigation through to production delivery.
• Preferred: Knowledge of EKS, fintech/payments/banking, multi-region architectures, chaos engineering, Prometheus, Grafana, Loki, Tempo, internal platform capabilities, self-service tools, and AI-assisted operational automation.
• Competitive gross monthly compensation ranging from $12,500 to $20,800 USD based on location and experience.
• Remote work opportunities.
• Chance to engage with cutting-edge fintech, cloud infrastructure, Kubernetes, AWS, and AI-assisted tools.
• Open-source-focused technology environment.
HavocAI
QuickNode ⚡
Sezzle
Sezzle
Get handpicked remote jobs straight to your inbox weekly.