
Principal Infrastructure Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Germany.
• Design, develop, manage, and scale the platform that powers Sezzle.
• Take ownership of intricate infrastructure projects from the architectural phase and prototyping through to implementation, deployment, and continuous operation.
• Recognize system constraints and implement enhancements to accommodate increased traffic, data volume, and workload complexity.
• Link business processes to infrastructure advancements across applications, data, and systems.
• Create capacity models, conduct load and stress testing, identify bottlenecks, and validate improvements in performance and cost.
• Design robust AWS accounts, IAM, network, and service architectures.
• Build and maintain the Kubernetes platform, including lifecycle automation, workload isolation, resource allocation, autoscaling, upgrades, and ensuring deployment reliability.
• Scale and optimize Aurora RDS for MySQL and Postgres, focusing on queries, indexes, connections, replication, failover, schema changes, and migrations.
• Establish service-level objectives and error budgets, implementing mechanisms for reliability.
• Participate in on-call duties and lead technical recovery efforts during critical incidents.
• Design, test, and document disaster recovery strategies.
• Develop infrastructure-as-code and automate operational processes.
• Enhance observability with metrics, logs, traces, dashboards, and alerts.
• Execute secure infrastructure migrations with phased rollouts and rollback strategies.
• Improve cost efficiency in cloud operations.
• Construct and assess AI-enhanced operational tools.
• Draft architecture proposals, assess trade-offs, review infrastructure modifications, and document system operations and failures.
• Bachelor’s degree in Computer Science or a related technical field (required).
• 12+ years of experience in infrastructure, platform, site reliability, software development, or related engineering fields.
• Extensive production experience with AWS, including compute, IAM, multi-account architectures, networking, VPC design, and private connectivity.
• In-depth production expertise with Kubernetes, covering cluster lifecycle, scheduling, resource management, autoscaling, networking, and troubleshooting.
• Strong expertise in RDS/Aurora MySQL and/or Postgres at scale.
• Proven track record of delivering infrastructure scaling enhancements personally.
• Excellent coding and automation capabilities in Golang, Python, or similar programming languages.
• Experience with infrastructure-as-code using Terraform or a comparable tool.
• Solid understanding of systems fundamentals including Linux, networking, DNS, TLS, storage, concurrency, and distributed system failure modes.
• Experience managing a 24/7 high-availability platform with active incident response and postmortem remediation.
• Willingness to engage in an on-call rotation.
• Experience implementing and testing disaster recovery strategies against specified recovery objectives.
• Practical experience in observability, load testing, capacity planning, and effective CI/CD practices.
• Active engagement with AI tooling in engineering or operations.
• Capability to navigate ambiguous technical challenges through to production delivery while collaborating across engineering disciplines.
• Competitive salary and performance bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible work schedule and remote work options.
• Opportunities for professional development and continuous learning.
• Collaborative and inclusive team culture.
Sezzle
HavocAI
QuickNode ⚡
Sezzle
Get handpicked remote jobs straight to your inbox weekly.