
Senior Platform Engineer – AI Native
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United Kingdom.
• Lead the design, execution, and delivery of intricate infrastructure projects.
• Develop and scale infrastructure tailored for AI workloads, focusing on cost efficiency, API reliability, and platform support for LLM-powered features.
• Troubleshoot cluster behavior under load, manage service-to-service traffic, address event-backbone consumer lag, and enhance database performance.
• Oversee the infrastructure-as-code framework, maintaining visibility on drift and ensuring module standards are met.
• Enhance delivery pipelines by implementing policy-as-code gates, progressive delivery, and ensuring changes are tested, reviewed, and auditable.
• Contribute to internal AI tools, including the MCP server and agentic workflows that minimize operational toil.
• Develop cost control measures through right-sizing, reserved capacity, and eliminating waste.
• Mentor colleagues through code pairing, review processes, and knowledge sharing.
• Collaborate with the infrastructure team to address security and governance matters.
• Achieve outcomes such as decreased operational effort, enhanced platform reliability, systemic incident resolutions, comprehensive infrastructure projects, reconciled drift, and adherence to established standards.
• Proficient software engineering capabilities in at least one programming language, with the flexibility to transition between languages.
• Experience in managing production infrastructure on a major cloud platform, including handling incidents.
• Extensive knowledge of infrastructure-as-code principles, encompassing state and module design.
• Practical experience with Kubernetes in production settings, including troubleshooting problematic clusters.
• Familiarity with CI/CD processes, particularly in pipeline design.
• Critical and proficient use of AI tools, along with an understanding of their limitations.
• Capability to work remotely in an asynchronous-first work environment.
• Composed and systematic incident response with clear communication during high-pressure situations.
• Availability overlap with the UK team within GMT+0 to GMT+4.
• Desirable: experience in freight, logistics, supply chain, B2B SaaS, or operational technology.
• Desirable: experience with workflow automation or orchestration tools, such as n8n.
• Desirable: familiarity with service mesh, event streaming at scale, or managed Postgres-compatible databases.
• Desirable: knowledge of identity and access management, SSO, secret management, or SRE methodologies.
• Desirable: experience working with LLM APIs, MCP servers, or agentic tooling.
• Fully remote work options.
• Async-first working culture.
• Flexible location, with consideration for GMT+0 to GMT+4 overlap with the UK team.
• Opportunities for pairing, code reviews, and knowledge sharing.
• Mentorship and professional development through enhanced judgment.
Agility Technologies Inc
American College of Education
First Due
Faire
Get handpicked remote jobs straight to your inbox weekly.