Senior Platform DevOps Engineer

Posted 3 days ago

This is a fully remote position, open to applicants in Colombia, +1 more country.

📋 Description

• Oversee, operate, and troubleshoot Kubernetes production environments and workloads.

• Develop, maintain, and enhance infrastructure utilizing Terraform and Infrastructure as Code methodologies.

• Manage application deployments, upgrades, configuration modifications, and intricate deployment lifecycles.

• Utilize GitOps-based deployment processes and tools such as Argo CD.

• Create and sustain Python scripts and tools for platform automation and operational workflows.

• Diagnose and resolve production incidents across infrastructure, applications, and platform services.

• Enhance platform reliability, scalability, observability, and operational efficiency.

• Support scaling and autoscaling strategies for production Kubernetes workloads.

• Collaborate with development and platform teams to address infrastructure and deployment challenges.

• Engage in technical decision-making and identify risks, reliability issues, and opportunities for improvement.

• Uphold high engineering and quality standards, providing technical pushback when necessary.

• Take ownership of platform initiatives and guide issues to resolution with minimal supervision.


⛳️ Requirements

• Extensive hands-on experience managing Kubernetes in production settings.

• Proficiency in Kubernetes deployments, scaling, troubleshooting, and operational management.

• Practical experience with Terraform for provisioning and managing cloud infrastructure.

• Capability to read and write Python for scripting, automation, troubleshooting, and platform tooling.

• Familiarity with AWS cloud infrastructure, preferably including EKS or similar managed Kubernetes environments.

• Knowledge of GitOps practices and deployment tools such as Argo CD.

• Experience in managing complex application deployment and upgrade lifecycles.

• Proven track record in troubleshooting, triaging, and supporting production incidents.

• Strong grasp of infrastructure reliability, scalability, and operational best practices.

• Excellent problem-solving abilities and the skill to independently investigate complex production issues.

• Strong sense of ownership and accountability, capable of functioning effectively with limited supervision.

• Quality-first approach with the judgment to balance delivery speed, reliability, and long-term maintainability.

• Exceptional communication and collaboration skills.

• Preferred: Experience with Helm and Kubernetes package/deployment management.

• Preferred: Knowledge of PyTorch and Hugging Face Transformers.

• Preferred: Experience in supporting GPU-based workloads or ML inference platforms.

• Preferred: Familiarity with NVIDIA Triton Inference Server.

• Preferred: Experience with Chainguard, distroless container images, Trivy, or reducing container vulnerabilities.

• Preferred: Experience in implementing or enhancing Kubernetes autoscaling solutions.

• Preferred: Knowledge of streaming or messaging platforms like Apache Kafka or similar technologies.

• Preferred: Familiarity with Elasticsearch or ArangoDB.

• Preferred: Experience troubleshooting complex service-to-service networking.

• Preferred: Exposure to OpenShift, IL5, FedRAMP, or similarly regulated environments.

• Preferred: Familiarity with AI/ML or agentic AI development environments.


🏝️ Benefits

• Remote work arrangement.

• Full-time employment.

People also viewed

Koniag Government Services1 day ago

Architect/DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
FP Markets (First Prudential Markets)2 days ago

Senior DevOps Engineer

AM flagArmenia, +4 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Modern Campus2 days ago

Senior DevOps Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
InRule2 days ago

Site Reliability Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Thumbtack2 days ago

Senior Software Engineer, Site Reliability Engineering

US flagUnited States, +38 more locationsFull-timeDevOps & Site Reliability Engineer (SRE)$179.4k – $272.8k/year
ApplyView job
Thumbtack2 days ago

Senior Software Engineer, Site Reliability Engineering

CA flagCanada, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)C$180.2k – C$233.2k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers