
Senior Platform DevOps Engineer
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Colombia, +1 more country.
• Oversee, operate, and troubleshoot Kubernetes production environments and workloads.
• Develop, maintain, and enhance infrastructure utilizing Terraform and Infrastructure as Code methodologies.
• Manage application deployments, upgrades, configuration modifications, and intricate deployment lifecycles.
• Utilize GitOps-based deployment processes and tools such as Argo CD.
• Create and sustain Python scripts and tools for platform automation and operational workflows.
• Diagnose and resolve production incidents across infrastructure, applications, and platform services.
• Enhance platform reliability, scalability, observability, and operational efficiency.
• Support scaling and autoscaling strategies for production Kubernetes workloads.
• Collaborate with development and platform teams to address infrastructure and deployment challenges.
• Engage in technical decision-making and identify risks, reliability issues, and opportunities for improvement.
• Uphold high engineering and quality standards, providing technical pushback when necessary.
• Take ownership of platform initiatives and guide issues to resolution with minimal supervision.
• Extensive hands-on experience managing Kubernetes in production settings.
• Proficiency in Kubernetes deployments, scaling, troubleshooting, and operational management.
• Practical experience with Terraform for provisioning and managing cloud infrastructure.
• Capability to read and write Python for scripting, automation, troubleshooting, and platform tooling.
• Familiarity with AWS cloud infrastructure, preferably including EKS or similar managed Kubernetes environments.
• Knowledge of GitOps practices and deployment tools such as Argo CD.
• Experience in managing complex application deployment and upgrade lifecycles.
• Proven track record in troubleshooting, triaging, and supporting production incidents.
• Strong grasp of infrastructure reliability, scalability, and operational best practices.
• Excellent problem-solving abilities and the skill to independently investigate complex production issues.
• Strong sense of ownership and accountability, capable of functioning effectively with limited supervision.
• Quality-first approach with the judgment to balance delivery speed, reliability, and long-term maintainability.
• Exceptional communication and collaboration skills.
• Preferred: Experience with Helm and Kubernetes package/deployment management.
• Preferred: Knowledge of PyTorch and Hugging Face Transformers.
• Preferred: Experience in supporting GPU-based workloads or ML inference platforms.
• Preferred: Familiarity with NVIDIA Triton Inference Server.
• Preferred: Experience with Chainguard, distroless container images, Trivy, or reducing container vulnerabilities.
• Preferred: Experience in implementing or enhancing Kubernetes autoscaling solutions.
• Preferred: Knowledge of streaming or messaging platforms like Apache Kafka or similar technologies.
• Preferred: Familiarity with Elasticsearch or ArangoDB.
• Preferred: Experience troubleshooting complex service-to-service networking.
• Preferred: Exposure to OpenShift, IL5, FedRAMP, or similarly regulated environments.
• Preferred: Familiarity with AI/ML or agentic AI development environments.
• Remote work arrangement.
• Full-time employment.
Koniag Government Services
FP Markets (First Prudential Markets)
Modern Campus
InRule
Get handpicked remote jobs straight to your inbox weekly.