Remotery

Lead DevOps Engineer

Posted Jun 26

This is a fully remote position, open to applicants in Germany.

📋 Description

• Take charge of and enhance Noxtua's infrastructure spanning OTC and our self-hosted GPU servers — ensuring effective architecture, dependable operation, and cost management.

• Lead and expand a team of 4–5 DevOps engineers, establishing technical direction, fostering their growth, and embodying a strong sense of ownership.

• Manage our self-hosted GPU server fleet — including provisioning, driver installation, hardening, and connectivity through Ansible — while overseeing provider SLAs to ensure heavy AI workloads operate reliably.

• Create and sustain infrastructure automation utilizing Infrastructure as Code (Terraform & Ansible).

• Operate our container platform on Kubernetes, assist teams with Docker, and maintain the stability, accessibility, and security of our services (APIs).

• Establish and uphold monitoring and alerting systems (e.g., Prometheus, Grafana) to guarantee system reliability and performance.

• Design and manage CI/CD pipelines while collaborating with development and AI teams to automate deployments and facilitate AI-driven workloads.


⛳️ Requirements

• Leadership: Proven experience in leading or mentoring a team, determining technical direction, and balancing hands-on operations with personnel responsibilities.

• Server fleet management: Experience managing a fleet of servers with an understanding of the underlying methodologies — beyond merely renting cloud instances.

• Familiarity with GPU servers is a significant advantage, though not mandatory.

• Strong expertise in Linux and Bash, along with proficiency in a scripting language such as Python.

• Demonstrated success in designing, operating, and managing costs for cloud-based architectures — preferably OTC (Open Telecom Cloud), or applicable experience from AWS, Azure, or Google Cloud — with solid networking knowledge (DNS, OSI model).

• A strong emphasis on automating provisioning and configuration using Terraform and Ansible.

• Proficient in containerizing applications with Docker and executing them at scale on Kubernetes.

• Capable of setting up and maintaining monitoring/alerting tools (e.g., Prometheus, Grafana), aggregating data, visualizing insights, and deriving actionable steps.


🏝️ Benefits

• 100% remote work available (with a German residence), other countries upon request.

• Flexible working hours.

• Vacation: 26 days + December 24th & 31st off, plus an additional vacation day for each year of employment (up to 30 days).

• Discounts: e.g., Urban Sports Club Membership, depending on location.

• Equipment: Laptop (Lenovo or Mac), along with a €1,000 net home office setup budget (paid with your first salary).

People also viewed

CVS Health4 hours ago

Salesforce DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$83.4k – $166.9k/year
ApplyView job
Devoteam5 hours ago

Data, AWS DevSecOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Aspirion6 hours ago

Senior DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Goodgame Studios6 hours ago

Senior Agentic Engineer – Java Backend, DevOps

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Instacart6 hours ago

Site Reliability Engineer II

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$133k – $169k/year
ApplyView job
Logicalis Spain7 hours ago

DevOps Engineer

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€40k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers