
Senior DevOps / Platform Engineer, AI Infrastructure
Posted 7 hours ago

Posted 7 hours ago
This is a fully remote position, open to applicants in Germany.
• Assume technical responsibility for MAIA’s production infrastructure.
• Manage and enhance self-administered Linux systems utilizing both virtual machines and dedicated servers.
• Oversee containers, networks, reverse proxies, and API gateways.
• Enhance Infrastructure as Code, GitHub Actions workflows, deployment methodologies, automated checks, versioning, and rollback functionalities.
• Ensure deployments are secure, repeatable, and user-friendly for engineers.
• Manage and optimize PostgreSQL in production, including performance assessments, connection pooling, capacity planning, backups, and verified restore procedures.
• Develop a self-hosted observability stack that integrates metrics, logs, traces, and actionable alerts.
• Fortify infrastructure security through IAM, least privilege access, secrets management, TLS, vulnerability scanning, and patch management.
• Implement technical controls for ISO 27001, ensuring continuous generation of clear and auditable evidence.
• Shape AI infrastructure by integrating and evaluating model and inference providers.
• Assess providers based on reliability, latency, throughput, cost, and operational effort.
• Investigate self-hosted LLM inference, potentially incorporating GPU infrastructure and vLLM.
• Enhance incident response and operational resilience through root-cause analysis, runbooks, documentation, and sustainable improvements.
• Advance the internal developer experience by minimizing manual tasks and establishing clear interfaces and workflows.
• Ensure transparency in infrastructure and inference costs to guide build, buy, and hosting decisions.
• Evaluate the existing platform, pinpoint risks, clarify options, and implement improvements for reliable production operations.
• Collaborate with software engineers, the CTO, company leadership, and remote team members.
• Function as a hands-on senior individual contributor without managerial responsibilities.
• Several years of experience managing production SaaS systems on Linux servers, either directly or through your team.
• Experience solely with fully managed hyperscaler services is insufficient.
• Production experience with Docker and Docker Compose.
• Knowledge of reverse proxies or API gateways such as Traefik or Kong.
• Practical experience with PostgreSQL, including personally configured and tested backups and restores, performance analysis, and connection pooling.
• Experience in building and maintaining CI/CD pipelines using GitHub Actions, GitLab CI, or similar systems.
• Proficiency in managing infrastructure using Terraform, Pulumi, or similar Infrastructure as Code tools.
• Experience in operating an observability stack employing Grafana, Loki, Prometheus, Sentry, or similar technologies.
• Understanding of IAM, least privilege, secrets management, vulnerability scanning, and patching.
• Capability to design infrastructure that is reliable for customers and user-friendly for engineers.
• Fluent communication in English.
• Proficiency in German is advantageous but not mandatory.
• Understanding of how RAG systems function and their common failures in production.
• Ability to discuss trade-offs in LLM serving, including latency, throughput, reliability, cost, and operational complexity.
• Familiarity with models, inference providers, and the broader GenAI market.
• Strong foundation in platform engineering and a technical curiosity to deepen expertise in AI infrastructure.
• Ability to independently manage critical business production systems.
• Capability to identify and prioritize infrastructure tasks without detailed guidance.
• Ability to lead technical enhancements across team boundaries.
• Reliable written and remote communication skills.
• Accountability for changes after deployment.
• Helpful but not required: experience with Hetzner, NixOS, self-hosted Supabase, ISO 27001 controls and audit evidence, SRE practices, GPUs, or LLM serving technologies like vLLM.
• Gross annual salary ranging from €75,000 to €85,000, based on experience and role scope.
• Opportunity to participate in our VSOP.
• Permanent, full-time employment.
• Flexible working hours.
• Fully remote work from anywhere within Germany.
• Regular opportunities to meet and collaborate with the team in Leipzig, with travel and accommodation expenses covered.
• Access to a WellPass fitness membership.
NavAide
Share
Lime
Convoso
Get handpicked remote jobs straight to your inbox weekly.