Senior DevOps Engineer, AI Platform

Posted Sep 2

This is a fully remote position, open to applicants in Argentina.

📋 Description

• Develop and maintain infrastructure that supports AI platforms, agent runtimes, web applications, backend services, APIs, and shared platform functionalities.

• Convert technical designs for applications and platforms into production-ready cloud infrastructure with minimal oversight.

• Design, provision, manage, and troubleshoot Kubernetes environments, focusing on Azure Kubernetes Service and Oracle Kubernetes Engine.

• Facilitate AI workloads, including LiteLLM-based gateways, Python agent runtimes, RAG workers, MCP services, background workers, and asynchronous processing pipelines.

• Design and oversee ingress and egress networking, load balancers, DNS, TLS, private connectivity, routing, NAT, firewalls, network policies, and inter-service communication.

• Construct and manage infrastructure for web applications and backend services, encompassing APIs, databases, caches, queues, scheduled jobs, and event-driven workloads.

• Create and maintain CI/CD pipelines utilizing Jenkins and Bitbucket, integrating Docker, Helm, Kubernetes, ArgoCD, and container registries.

• Automate the provisioning and configuration of infrastructure using Terraform, Helm, Kubernetes manifests, Python, Bash, and related tools.

• Implement comprehensive observability through metrics, logs, distributed tracing, dashboards, alerts, health checks, and SLOs.

• Take ownership of production readiness, incident troubleshooting, root cause analysis, scalability, reliability, and optimizing infrastructure costs.

• Develop reusable infrastructure patterns that enable engineering teams to launch new services efficiently and consistently.

• Determine necessary cloud resources, Kubernetes configurations, namespaces, scaling models, and supporting services based on technical designs.

• Configure ingress and egress, private connectivity, DNS, TLS, service communication, identities, and secrets.

• Provision and manage dependencies such as PostgreSQL, Redis, RabbitMQ, storage, and other shared services.

• Build Jenkins and Bitbucket CI/CD workflows for building, testing, container publishing, deployment, validation, and rollback.

• Define observability metrics, health checks, dashboards, alerts, capacity monitoring, and operational runbooks prior to production launch.

• Oversee infrastructure delivery throughout UAT and production, collaborating with architects and engineers on design trade-offs.


⛳️ Requirements

• A minimum of 7 years of experience in DevOps, SRE, Platform Engineering, Cloud Infrastructure, or a related field.

• Extensive hands-on experience managing production Kubernetes environments with in-depth knowledge of networking, scheduling, storage, autoscaling, security, and troubleshooting.

• Strong proficiency with Microsoft Azure, including AKS, networking, identity, storage, and monitoring.

• Experience with Oracle Cloud Infrastructure (OCI) is preferred, or a proven ability to work across various cloud providers.

• In-depth knowledge of cloud networking, including virtual networks, subnets, routing, NAT, load balancers, private networking, DNS, TLS, firewalls, ingress, and egress.

• Significant experience with Jenkins, Bitbucket, Docker, Terraform, Helm, Kubernetes, and Infrastructure as Code practices.

• Proven track record supporting production web applications and backend services, including REST APIs, microservices, background workers, and asynchronous architectures.

• Hands-on experience with databases, caching, and messaging systems such as PostgreSQL, Redis, RabbitMQ, or equivalent technologies.

• Experience in implementing production observability using tools like OpenTelemetry, Grafana, Prometheus, Sentry, cloud monitoring, or similar solutions.

• Strong skills in Linux, systems management, and production troubleshooting.

• Familiarity with Python, particularly in developing backend services using frameworks such as FastAPI.

• Knowledge of at least one additional programming language, such as C#, Java, Go, JavaScript, or TypeScript.

• Understanding of HTTP, HTTPS, DNS, TCP/IP, proxies, authentication, APIs, connection pooling, caching, concurrency, queues, retries, dead letter queues, and asynchronous processing.

• Ability to analyze application logs and stack traces to diagnose latency, memory, CPU, connection, and dependency issues.


🏝️ Benefits

• Competitive salary and comprehensive benefits package.

• Opportunities for professional development and continuous learning.

• Flexible working hours and remote work options.

• Collaborative and innovative work environment.

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers