
Senior DevOps Engineer – Production Support
Posted Aug 27

Posted Aug 27
This is a fully remote position, open to applicants in Brazil, +4 more countries.
• Oversee essential production systems, such as Azure Kubernetes Service (AKS), microservices, and CI/CD pipelines, utilizing dashboards and proactive alerting mechanisms.
• Serve as the main technical point of contact for live production incidents and Slack escalations.
• Conduct swift triage, root-cause analysis, and resolution of production incidents.
• Maintain, enhance, and optimize internal runbooks and standard operating procedures (SOPs).
• Manage and support deployments across both production and non-production environments, ensuring compliance with SLAs and corporate response times.
• Collaborate with DevOps and software engineering teams to address recurring systemic issues and bolster platform reliability.
• Create and implement automation scripts for repetitive operational tasks to minimize manual effort.
• Work in conjunction with US-based teams operating within Central Time during core hours.
• 6+ years of demonstrated experience in DevOps, Cloud Infrastructure, or high-pressure Production Support roles.
• Strong and comprehensive understanding of Microsoft Azure fundamentals, particularly in Compute, Networking, and Azure Monitoring ecosystems.
• Practical operational experience with Kubernetes, focusing on Azure Kubernetes Service (AKS) operations, log analysis, and cluster scaling.
• Proficient with modern monitoring and observability tools such as Azure Monitor, Grafana, Prometheus, or similar platforms.
• Familiarity with structured incident management, escalation processes, and stringent SLA guidelines.
• Intermediate to advanced scripting skills in Bash, PowerShell, or Python.
• Significant experience working independently in Agile teams within fully remote environments.
• Outstanding verbal and written communication skills in English.
• All interviews, technical documentation, and daily communications will be conducted in English.
• Nice to have: hands-on experience with CI/CD pipelines like GitHub Actions and Jenkins.
• Nice to have: practical familiarity with Infrastructure as Code concepts and tools like Terraform and Bicep.
• Nice to have: previous involvement in 24/7 mission-critical or high-availability infrastructure settings.
• Nice to have: knowledge of ITIL frameworks or structured enterprise incident management systems.
• 100% remote work.
• Full-time vendor contract directly with Inallmedia.com.
• Time zone alignment with Central Time (CT) ±2 hours.
NVIDIA
EXL
Revecore
Get handpicked remote jobs straight to your inbox weekly.