
Cloud DevOps Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Argentina, +4 more countries.
• Take charge of the production infrastructure for AI systems, which includes managing clusters, deployment pipelines, and monitoring.
• Provision, upgrade, network, and configure production Kubernetes clusters.
• Create reproducible infrastructure as code using Terraform or a similar tool and identify infrastructure drift.
• Develop and maintain dependable CI/CD pipelines along with standardized container image workflows.
• Implement monitoring and alerting systems, investigate incidents, identify root causes, and execute preventive measures.
• Provide the necessary infrastructure for AI workloads, including inference services, their scaling, and cost profiles.
• Analyze infrastructure expenses and capacity requirements, including projected costs at increased traffic levels.
• Secure Linux, Kubernetes, containers, and service meshes, focusing on secrets, access controls, and audit trails.
• Create internal tools and documentation while supporting developers and QA during release cycles.
• Collaborate within client environments, repositories, cloud accounts, and change processes as needed.
• Operate within Azumo's SOC 2-certified environment and address engagement requirements such as HIPAA.
• Utilize AI-assisted engineering tools and automated codebase audits to evaluate security, cost, and architectural findings.
• Over 5 years of experience as a DevOps, Site Reliability Engineer (SRE), or systems engineer managing production infrastructure.
• Proficient in Linux administration, networking, Git, and scripting in Bash, along with Python or Go.
• Experience with production Kubernetes, including the provisioning, upgrading, and debugging of clusters.
• Familiarity with infrastructure as code practices using Terraform or equivalent, including state management, modules, and environment parity.
• Proven experience in building and maintaining CI/CD pipelines with GitHub Actions, GitLab, or similar tools.
• Responsible for managing container images.
• Experience deploying on cloud platforms like AWS, Azure, or GCP, including managed Kubernetes services and relevant cost models.
• Monitoring and incident response experience with tools like Datadog, CloudWatch, or similar.
• Proven experience in managing incidents from alert to root cause analysis and implementing preventive changes.
• Expertise in security hardening of Linux, containers, and Kubernetes.
• Experience in managing infrastructure costs and capacity.
• Active utilization of AI-assisted coding tools such as Claude Code, Cursor, or GitHub Copilot in your work.
• Proficient in written and spoken English at a C1 level or higher.
• Ability to communicate technical trade-offs directly to clients.
• Bachelor's degree in Computer Science, a related field, or equivalent professional experience.
• Preferred: experience with service mesh and traffic management technologies like Istio, Linkerd, or equivalent.
• Preferred: familiarity with Helm, Kustomize, or similar templating and environment configuration tools.
• Preferred: experience with production-scale database operations using PostgreSQL, MongoDB, RDS, DynamoDB, or equivalent.
• Preferred: experience running or scaling inference workloads, assessing cost and latency trade-offs against hosted APIs.
• Preferred: experience delivering under compliance frameworks such as SOC 2 or HIPAA.
• Preferred: cloud certifications, contributions to open-source infrastructure, or published technical writings.
• Fully remote-first culture (work from anywhere in Latin America).
• Paid time off (PTO).
• Observance of U.S. Holidays.
• Comprehensive AI training and certification.
• Supported career development with mentorship.
• Profit-sharing opportunities.
• Compensation in U.S. dollars.
• Maternity coverage.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.