Remotery

DevOps Engineer, AI Inference

Posted 1 day ago

This is a fully remote position, open to applicants in Cyprus, +4 more states.

📋 Description

• Design, develop, and sustain infrastructure for AI inference workloads, which includes GPU scheduling, model deployment pipelines, and data access patterns in on-premises environments.

• Create and oversee monitoring and observability tools for AI inference platforms, including dashboards, alerts, and runbooks focused on model health and system performance.

• Collaborate with machine learning engineers and platform teams to architect systems for AI workloads.

• Integrate inference runtimes effectively.

• Evaluate AI inference performance at scale.


⛳️ Requirements

• This role is exclusively available under an employment (labor) agreement.

• In-depth understanding of Kubernetes architecture, encompassing CNI, CSI, operators, ingress/gateway, and control plane components.

• Practical experience in operating and troubleshooting production Kubernetes clusters.

• Strong skills in Linux and networking troubleshooting, including DNS, routing, firewalling, TLS, MTU, connectivity, and performance issues.

• Capability to develop automation and operational tools utilizing Python, Go, or Bash.

• Familiarity with Terraform, Ansible, or similar Infrastructure as Code/configuration management tools.

• Experience with VictoriaMetrics/Grafana or comparable monitoring, alerting, and troubleshooting solutions.

• Solid experience with Git-based workflows and CI/CD pipelines.

• Knowledge of Cluster API or similar technologies for Kubernetes cluster lifecycle management.

• Hands-on experience with the operation or administration of Slurm clusters.

• Understanding of Argo CD, GitOps workflows, Helm, or Helmfile.

• Background in working with managed platforms, PaaS, or cloud services.

• Exposure to bare metal, GPU, HPC, or other high-performance computing environments.

• Familiarity with the NVIDIA GPU stack, RDMA/InfiniBand, or high-performance networking.

• Understanding of OpenStack or similar cloud infrastructure platforms.

• Practical experience in developing Kubernetes operators or controllers.


🏝️ Benefits

• Competitive compensation.

• Flexible working hours with hybrid or remote options, depending on your role.

• Opportunity to work from anywhere in the world for up to 45 days per year.

• Private medical insurance for you and your family.*

• Additional paid vacation and sick leave days.*

• Support for important life moments and celebrations.

• Language courses to facilitate connection and growth.

• Modern, inviting offices equipped with snacks, drinks, and entertainment.*

• Team sports and social activities.*

People also viewed

CWILL14 hours ago

DevOps/SRE Engineer, Bilingual Mandarin

US flagCalifornia, +4 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$100k – $130k/year
ApplyView job
a3715 hours ago

Forward Deployed DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GT16 hours ago

Site Reliability Engineer, SRE

PL flagPoland, +2 more statesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group16 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Applaudo16 hours ago

Google Cloud DevOps Engineer – Temporary Contract

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Branch16 hours ago

Cloud Operations Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers