
Senior DevOps Engineer, AI Platform
Posted Sep 2

Posted Sep 2
This is a fully remote position, open to applicants in Argentina.
• Develop and maintain infrastructure that supports AI platforms, agent runtimes, web applications, backend services, APIs, and shared platform functionalities.
• Convert technical designs for applications and platforms into production-ready cloud infrastructure with minimal oversight.
• Design, provision, manage, and troubleshoot Kubernetes environments, focusing on Azure Kubernetes Service and Oracle Kubernetes Engine.
• Facilitate AI workloads, including LiteLLM-based gateways, Python agent runtimes, RAG workers, MCP services, background workers, and asynchronous processing pipelines.
• Design and oversee ingress and egress networking, load balancers, DNS, TLS, private connectivity, routing, NAT, firewalls, network policies, and inter-service communication.
• Construct and manage infrastructure for web applications and backend services, encompassing APIs, databases, caches, queues, scheduled jobs, and event-driven workloads.
• Create and maintain CI/CD pipelines utilizing Jenkins and Bitbucket, integrating Docker, Helm, Kubernetes, ArgoCD, and container registries.
• Automate the provisioning and configuration of infrastructure using Terraform, Helm, Kubernetes manifests, Python, Bash, and related tools.
• Implement comprehensive observability through metrics, logs, distributed tracing, dashboards, alerts, health checks, and SLOs.
• Take ownership of production readiness, incident troubleshooting, root cause analysis, scalability, reliability, and optimizing infrastructure costs.
• Develop reusable infrastructure patterns that enable engineering teams to launch new services efficiently and consistently.
• Determine necessary cloud resources, Kubernetes configurations, namespaces, scaling models, and supporting services based on technical designs.
• Configure ingress and egress, private connectivity, DNS, TLS, service communication, identities, and secrets.
• Provision and manage dependencies such as PostgreSQL, Redis, RabbitMQ, storage, and other shared services.
• Build Jenkins and Bitbucket CI/CD workflows for building, testing, container publishing, deployment, validation, and rollback.
• Define observability metrics, health checks, dashboards, alerts, capacity monitoring, and operational runbooks prior to production launch.
• Oversee infrastructure delivery throughout UAT and production, collaborating with architects and engineers on design trade-offs.
• A minimum of 7 years of experience in DevOps, SRE, Platform Engineering, Cloud Infrastructure, or a related field.
• Extensive hands-on experience managing production Kubernetes environments with in-depth knowledge of networking, scheduling, storage, autoscaling, security, and troubleshooting.
• Strong proficiency with Microsoft Azure, including AKS, networking, identity, storage, and monitoring.
• Experience with Oracle Cloud Infrastructure (OCI) is preferred, or a proven ability to work across various cloud providers.
• In-depth knowledge of cloud networking, including virtual networks, subnets, routing, NAT, load balancers, private networking, DNS, TLS, firewalls, ingress, and egress.
• Significant experience with Jenkins, Bitbucket, Docker, Terraform, Helm, Kubernetes, and Infrastructure as Code practices.
• Proven track record supporting production web applications and backend services, including REST APIs, microservices, background workers, and asynchronous architectures.
• Hands-on experience with databases, caching, and messaging systems such as PostgreSQL, Redis, RabbitMQ, or equivalent technologies.
• Experience in implementing production observability using tools like OpenTelemetry, Grafana, Prometheus, Sentry, cloud monitoring, or similar solutions.
• Strong skills in Linux, systems management, and production troubleshooting.
• Familiarity with Python, particularly in developing backend services using frameworks such as FastAPI.
• Knowledge of at least one additional programming language, such as C#, Java, Go, JavaScript, or TypeScript.
• Understanding of HTTP, HTTPS, DNS, TCP/IP, proxies, authentication, APIs, connection pooling, caching, concurrency, queues, retries, dead letter queues, and asynchronous processing.
• Ability to analyze application logs and stack traces to diagnose latency, memory, CPU, connection, and dependency issues.
• Competitive salary and comprehensive benefits package.
• Opportunities for professional development and continuous learning.
• Flexible working hours and remote work options.
• Collaborative and innovative work environment.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.