
Senior Software Engineer, DevOps
Posted 19 hours ago

Posted 19 hours ago
This is a fully remote position, open to applicants in Washington.
• Develop and manage Amazon EKS clusters and self-hosted Kubernetes clusters on virtualized infrastructure.
• Take ownership of cluster provisioning, add-ons, upgrades, and ongoing operations.
• Construct AWS infrastructure using Python with AWS CDK across a multi-account organization.
• Oversee networking, IAM, DNS, compute, containers, databases, and storage infrastructures.
• Collaborate with Digital, Data Engineering, and other teams to provide environments, access, and pipelines.
• Assist teams in troubleshooting production issues.
• Create and maintain CI/CD workflows, reusable GitHub Actions, container build and release patterns, and GitOps deployments to Kubernetes.
• Engage in the on-call rotation and respond to incidents promptly.
• Document postmortems and ensure follow-through on incident resolutions.
• Develop dashboards and alerts for monitoring.
• Create self-service tooling and implement repeatable automation.
• Write runbooks, how-to guides, and design documentation.
• Report directly to the leader of the DevOps team.
• Minimum of 6 years of experience in DevOps, site reliability, platform, or infrastructure engineering, with hands-on responsibility for production systems.
• Practical experience in building and operating production workloads on AWS, covering IAM, VPC networking, EC2, EKS or ECS, Lambda, RDS, S3, Route 53, and CloudWatch.
• Production experience with Kubernetes is essential, ideally with both EKS and self-managed clusters.
• Proficient in Python programming.
• Experience with infrastructure as code is required; familiarity with AWS CDK is preferred.
• Capable of reading, understanding, and making specific modifications to TypeScript.
• Experienced in creating pipelines with GitHub Actions or similar tools.
• Knowledge in building and securing container images.
• Familiarity with deploying using Helm and GitOps tools such as Argo CD or Flux.
• Experience with monitoring and logging tools like Prometheus, Grafana, OpenTelemetry, or CloudWatch.
• Participation in an on-call rotation for production systems is required.
• Strong Linux administration skills.
• Understanding of TCP/IP, DNS, load balancing, TLS certificates, and VPN connectivity.
• Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
• Candidates must have authorization to work in the United States on a permanent basis.
• AWS Associate or Professional certification, Rancher experience, or CKA certification are advantageous.
• Ability to thrive in a fast-paced, resource-limited startup environment.
• Medical insurance
• Dental insurance
• Vision insurance
• Life insurance
• Disability insurance
• Vacation
• 401k
• Eligibility for equity program
• Eligibility for discretionary annual incentive program
Funding Xchange
Leidos
LeoLabs
OpenRouter
Get handpicked remote jobs straight to your inbox weekly.