
Principal Infrastructure Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Turkey.
• Design, develop, manage, and enhance Sezzle’s infrastructure platform.
• Take ownership of intricate technical projects from architecture and prototyping to implementation, production deployment, and ongoing management.
• Recognize system limitations and implement enhancements to accommodate increased traffic, data volume, and workload complexity.
• Integrate business workflows with infrastructure advancements across applications, data, and systems.
• Construct capacity models, conduct load and stress testing, identify bottlenecks, and confirm improvements in throughput, latency, saturation, and costs.
• Design robust AWS account structures, IAM configurations, networks, and service architectures.
• Develop and manage the Kubernetes platform, ensuring lifecycle automation, workload isolation, autoscaling, upgrades, and deployment reliability.
• Scale and optimize Aurora RDS for MySQL and Postgres databases.
• Establish service-level objectives and error budgets; implement failure isolation, backpressure, load shedding, and safe retries.
• Engage in on-call duties and lead technical recovery efforts during critical incidents and outages.
• Create, test, and document disaster recovery, backup, restoration, and failover strategies.
• Develop infrastructure-as-code solutions and operational automation.
• Enhance observability through metrics, logs, traces, dashboards, and actionable alerts.
• Execute safe infrastructure migrations with phased rollouts, validation, and rollback strategies.
• Improve cloud cost efficiency through technical modifications.
• Develop and assess AI-assisted operational tools.
• Draft architecture proposals, evaluate technology trade-offs, review shared infrastructure changes, and document system operations and failure scenarios.
• Report to engineering leadership and collaborate with application engineers, Security, and Compliance teams.
• Bachelor's degree in Computer Science or a related technical field (mandatory).
• Over 12 years of experience in infrastructure, platform, site reliability, software development, or similar engineering fields.
• Extensive production experience with AWS, covering compute, IAM, multi-account architectures, networking, VPC design, and private connectivity.
• In-depth production expertise with Kubernetes, including cluster lifecycle management, scheduling, resource allocation, autoscaling, networking, and troubleshooting.
• Strong proficiency with RDS/Aurora MySQL and/or Postgres at scale.
• Proven track record of personally delivering improvements in infrastructure scaling.
• Proficient coding and automation skills in Golang, Python, or similar programming languages.
• Experience with infrastructure-as-code using Terraform or similar tools.
• Solid understanding of systems fundamentals including Linux, networking, DNS, TLS, storage, concurrency, and distributed system failure modes.
• Experience managing a 24/7 high-availability platform with direct customer or revenue impact.
• Willingness to participate in an on-call rotation.
• Experience in implementing and testing disaster recovery against established recovery objectives.
• Practical knowledge of observability, load testing, capacity planning, and safe CI/CD practices.
• Active usage of AI tools in engineering or operational contexts.
• Capability to navigate ambiguous technical challenges through to production delivery and collaborate across engineering disciplines.
• Preferred qualifications include experience in fintech, payments, or banking; familiarity with multi-region architectures; chaos engineering; tools such as Prometheus, Grafana, Loki, Tempo; internal platform capabilities; and AI-assisted incident investigation or operational automation.
• Competitive gross monthly salary ranging from $12,500 to $20,800 USD.
• Remote working opportunity in Türkiye.
• Chance to collaborate with a global fintech and retail technology organization.
• Exposure to AWS, Kubernetes, Aurora RDS, AI-assisted infrastructure, and SRE tooling.
• An open-source-focused engineering environment.
VALR
First Circle
crewAI
Sezzle
Get handpicked remote jobs straight to your inbox weekly.