Site Reliability Engineer

atKong Inc.RemoteUS flagWashingtonFull-timeDevOps & Site Reliability Engineer (SRE)Mid-levelSenior$123k – $150k/year

Posted Aug 25

This is a fully remote position, open to applicants in Washington.

📋 Description

• Manage and expand Kong’s global SaaS platform, Konnect, across various regions and cloud environments.

• Construct, automate, and sustain Kubernetes-based infrastructure and deployment workflows utilizing Terraform/Terragrunt, Helm, and ArgoCD.

• Design, oversee, and optimize data and caching layers across multiple regions, including PostgreSQL, Redis, ClickHouse, and Druid.

• Operate and enhance Kong Gateway and Kong Mesh environments that support hybrid and distributed architectures.

• Develop and maintain CI/CD pipelines along with GitOps workflows.

• Improve observability and readiness for incident response utilizing tools such as Datadog, Prometheus, Grafana, and Thanos.

• Establish and monitor service-level objectives.

• Partner with development and security teams to ensure SaaS services operate in alignment with reliability, security, and regulatory standards.

• Engage in a global 24/7 on-call rotation.

• Refine operational playbooks and enhance postmortem practices.

• Lead and contribute to initiatives aimed at scaling that enhance elasticity, reliability, and cost-effectiveness.


⛳️ Requirements

• Bachelor’s degree in Computer Science or equivalent practical experience.

• Demonstrated experience in managing SaaS or PaaS systems at an enterprise level within secure, multi-region, multi-tenant environments.

• Extensive knowledge of Kubernetes, including troubleshooting cluster/networking issues and designing for fault tolerance and scalability.

• Strong expertise in Terraform or Terragrunt.

• Experience with CI/CD pipelines and GitOps workflows, including ArgoCD, Atlantis, and Helm.

• Proficiency in programming languages such as Go, Python, or Bash.

• Comprehensive understanding of Linux/Unix systems, DNS, TLS/SSL, HTTP, load balancers, and distributed systems.

• Experience with API gateway and service mesh technologies.

• Familiarity with Kafka and observability tools like Datadog, Prometheus, and Grafana.

• Experience in a 24/7/365 production support environment.

• Must be legally authorized to work in the country where the position is located.

• Must disclose any current or future sponsorship requirements.


🏝️ Benefits

• Healthcare benefits.

• 401(k) plan.

• Short-term disability benefits.

• Long-term disability benefits.

• Basic life insurance.

• AD&D insurance.

• Additional rewards may be available depending on the applicable plan and role.

People also viewed

Bet On Talent20 hours ago

Senior DevOps Engineer

EuropeFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Virtasant21 hours ago

Build & Release Support Engineer – CI/CD

MX flagMexico, +5 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ookla21 hours ago

Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$90k – $100k/year
ApplyView job
opinov821 hours ago

Senior DevOps Engineer, Media and Advertising Industry

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GFT Technologies21 hours ago

DevOps Specialist

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
RELX1 day ago

Senior Site Reliability Engineer II

US flagNorth Carolina, +3 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$104.9k – $174.7k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers