
DevOps Engineer, AI Inference
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Cyprus, +4 more states.
• Design, develop, and sustain infrastructure for AI inference workloads, which includes GPU scheduling, model deployment pipelines, and data access patterns in on-premises environments.
• Create and oversee monitoring and observability tools for AI inference platforms, including dashboards, alerts, and runbooks focused on model health and system performance.
• Collaborate with machine learning engineers and platform teams to architect systems for AI workloads.
• Integrate inference runtimes effectively.
• Evaluate AI inference performance at scale.
• This role is exclusively available under an employment (labor) agreement.
• In-depth understanding of Kubernetes architecture, encompassing CNI, CSI, operators, ingress/gateway, and control plane components.
• Practical experience in operating and troubleshooting production Kubernetes clusters.
• Strong skills in Linux and networking troubleshooting, including DNS, routing, firewalling, TLS, MTU, connectivity, and performance issues.
• Capability to develop automation and operational tools utilizing Python, Go, or Bash.
• Familiarity with Terraform, Ansible, or similar Infrastructure as Code/configuration management tools.
• Experience with VictoriaMetrics/Grafana or comparable monitoring, alerting, and troubleshooting solutions.
• Solid experience with Git-based workflows and CI/CD pipelines.
• Knowledge of Cluster API or similar technologies for Kubernetes cluster lifecycle management.
• Hands-on experience with the operation or administration of Slurm clusters.
• Understanding of Argo CD, GitOps workflows, Helm, or Helmfile.
• Background in working with managed platforms, PaaS, or cloud services.
• Exposure to bare metal, GPU, HPC, or other high-performance computing environments.
• Familiarity with the NVIDIA GPU stack, RDMA/InfiniBand, or high-performance networking.
• Understanding of OpenStack or similar cloud infrastructure platforms.
• Practical experience in developing Kubernetes operators or controllers.
• Competitive compensation.
• Flexible working hours with hybrid or remote options, depending on your role.
• Opportunity to work from anywhere in the world for up to 45 days per year.
• Private medical insurance for you and your family.*
• Additional paid vacation and sick leave days.*
• Support for important life moments and celebrations.
• Language courses to facilitate connection and growth.
• Modern, inviting offices equipped with snacks, drinks, and entertainment.*
• Team sports and social activities.*
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.