
Senior DevOps Engineer
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in Brazil.
• Deploy, manage, and enhance a microservices-oriented platform operating within Kubernetes clusters on AWS, GCP, and on-premises Rancher.
• Support and operate GPU-accelerated ML inference services utilizing Triton Inference Server and vLLM on RunPod, Scaleway, and Nebius.
• Create and uphold Docker images for all microservices while ensuring a reliable service lifecycle.
• Maintain and scale both development and production Kubernetes clusters.
• Engage in deployment troubleshooting, incident investigation, and performance optimization.
• Develop, sustain, and advance custom Helm charts for every service.
• Design and manage CI/CD pipelines utilizing GitHub and GitLab for on-premises client deployments.
• Ensure platform adherence to SOC 2 standards and enhance security and compliance procedures.
• Manage cluster access through NetBird VPN and implement role-based access control via group policies.
• Deploy and oversee infrastructure using Terraform and Ansible Infrastructure as Code (IaC) practices.
• Develop and consistently enhance observability systems incorporating Grafana, Prometheus, and the ELK stack.
• Continuously refine infrastructure across IaC, IAM, observability, and CI/CD.
• Work with technologies such as Python, Kubernetes, Linux, Docker, GitHub CI/CD, PostgreSQL, ClickHouse, Kafka, Superset, Terraform, and Ansible.
• A minimum of 5 years of experience in a DevOps and/or Site Reliability Engineering position.
• Strong hands-on experience in Linux system administration.
• Extensive experience in deploying, managing, and scaling Kubernetes in both cloud and bare-metal settings.
• Deep knowledge and practical experience with at least one leading cloud provider, preferably Google Cloud Platform.
• Experience with ML inference on GPU/CPU is a significant advantage.
• Proven track record of implementing SRE methodologies and constructing observability stacks using Grafana, Prometheus, and Loki.
• Strong commitment to GitOps, Infrastructure as Code (IaC), and CI/CD principles.
• Advanced proficiency in Terraform, Ansible, and Python.
• Capability to operate in high-uncertainty environments and quickly adapt to new technologies and patterns.
• Ability to analyze beyond DevOps functions and effectively debug and comprehend the product.
• Capacity to select technologies and architectural strategies based on long-term objectives.
• English proficiency is implied through private English lessons, though no explicit language requirement is stated.
• Award-winning AI products and a state-of-the-art technology stack.
• Fully remote work environment.
• 21 vacation days, in addition to public holidays and 5 sick days.
• Private English lessons through Preply.
• Fast-paced startup culture with enterprise-level stability.
• Rapid career advancement opportunities.
• Genuine ownership and direct impact from your work.
Koniag Government Services
FP Markets (First Prudential Markets)
Modern Campus
InRule
Get handpicked remote jobs straight to your inbox weekly.