Senior DevOps / Platform Engineer, AI Infrastructure

atMAIARemoteDE flagGermanyFull-timePlatform EngineerSenior€75k – €85k/year

Posted 7 hours ago

This is a fully remote position, open to applicants in Germany.

📋 Description

• Assume technical responsibility for MAIA’s production infrastructure.

• Manage and enhance self-administered Linux systems utilizing both virtual machines and dedicated servers.

• Oversee containers, networks, reverse proxies, and API gateways.

• Enhance Infrastructure as Code, GitHub Actions workflows, deployment methodologies, automated checks, versioning, and rollback functionalities.

• Ensure deployments are secure, repeatable, and user-friendly for engineers.

• Manage and optimize PostgreSQL in production, including performance assessments, connection pooling, capacity planning, backups, and verified restore procedures.

• Develop a self-hosted observability stack that integrates metrics, logs, traces, and actionable alerts.

• Fortify infrastructure security through IAM, least privilege access, secrets management, TLS, vulnerability scanning, and patch management.

• Implement technical controls for ISO 27001, ensuring continuous generation of clear and auditable evidence.

• Shape AI infrastructure by integrating and evaluating model and inference providers.

• Assess providers based on reliability, latency, throughput, cost, and operational effort.

• Investigate self-hosted LLM inference, potentially incorporating GPU infrastructure and vLLM.

• Enhance incident response and operational resilience through root-cause analysis, runbooks, documentation, and sustainable improvements.

• Advance the internal developer experience by minimizing manual tasks and establishing clear interfaces and workflows.

• Ensure transparency in infrastructure and inference costs to guide build, buy, and hosting decisions.

• Evaluate the existing platform, pinpoint risks, clarify options, and implement improvements for reliable production operations.

• Collaborate with software engineers, the CTO, company leadership, and remote team members.

• Function as a hands-on senior individual contributor without managerial responsibilities.


⛳️ Requirements

• Several years of experience managing production SaaS systems on Linux servers, either directly or through your team.

• Experience solely with fully managed hyperscaler services is insufficient.

• Production experience with Docker and Docker Compose.

• Knowledge of reverse proxies or API gateways such as Traefik or Kong.

• Practical experience with PostgreSQL, including personally configured and tested backups and restores, performance analysis, and connection pooling.

• Experience in building and maintaining CI/CD pipelines using GitHub Actions, GitLab CI, or similar systems.

• Proficiency in managing infrastructure using Terraform, Pulumi, or similar Infrastructure as Code tools.

• Experience in operating an observability stack employing Grafana, Loki, Prometheus, Sentry, or similar technologies.

• Understanding of IAM, least privilege, secrets management, vulnerability scanning, and patching.

• Capability to design infrastructure that is reliable for customers and user-friendly for engineers.

• Fluent communication in English.

• Proficiency in German is advantageous but not mandatory.

• Understanding of how RAG systems function and their common failures in production.

• Ability to discuss trade-offs in LLM serving, including latency, throughput, reliability, cost, and operational complexity.

• Familiarity with models, inference providers, and the broader GenAI market.

• Strong foundation in platform engineering and a technical curiosity to deepen expertise in AI infrastructure.

• Ability to independently manage critical business production systems.

• Capability to identify and prioritize infrastructure tasks without detailed guidance.

• Ability to lead technical enhancements across team boundaries.

• Reliable written and remote communication skills.

• Accountability for changes after deployment.

• Helpful but not required: experience with Hetzner, NixOS, self-hosted Supabase, ISO 27001 controls and audit evidence, SRE practices, GPUs, or LLM serving technologies like vLLM.


🏝️ Benefits

• Gross annual salary ranging from €75,000 to €85,000, based on experience and role scope.

• Opportunity to participate in our VSOP.

• Permanent, full-time employment.

• Flexible working hours.

• Fully remote work from anywhere within Germany.

• Regular opportunities to meet and collaborate with the team in Leipzig, with travel and accommodation expenses covered.

• Access to a WellPass fitness membership.

People also viewed

NavAide9 hours ago

Power Platform Developer, DOD Clearance Required

US flagCalifornia OnlyFull-timePlatform Engineer$98k – $140k/year
ApplyView job
Share14 hours ago

Platform Engineer

ES flagSpain, +6 more countriesFull-timePlatform Engineer
ApplyView job
Lime15 hours ago

Staff Software Engineer, Platform Engineering

CA flagCanada OnlyFull-timePlatform EngineerC$172k – C$237k/year
ApplyView job
Convoso16 hours ago

Senior Platform Engineer

US flagCalifornia OnlyFull-timePlatform Engineer$190k – $210k/year
ApplyView job
Cortex20 hours ago

Platform Engineer – SRE, Mid-Level

BR flagBrazil OnlyFull-timePlatform Engineer
ApplyView job
Sigma Software Group1 day ago

Senior MLOps, ML Platform Engineer

ML flagMali, +1 more countryFull-timePlatform Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers