Senior DevOps Engineer

Posted Aug 28

This is a fully remote position, open to applicants in United States.

📋 Description

• Develop and enhance cloud architecture on GCP, encompassing networking, interconnects, IAM, and high-availability topology, implemented as Terraform code in accordance with GitOps principles.

• Create and manage CI/CD pipelines that plan, review, test, and safely deploy Infrastructure as Code (IaC) changes, incorporating Policy-as-Code guardrails, drift detection, and gradual rollout strategies.

• Promote Platform-as-a-Product by establishing self-service capabilities and optimal paths for engineers.

• Enhance observability across metrics, logs, traces, and alerting using tools like Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.

• Manage GKE clusters and infrastructure services, including Helm-packaged workloads, RabbitMQ, IBM MQ, and various data stores.

• Engage in the Follow-The-Sun on-call model; assess alerts, participate in incident declaration, lead debugging and escalation efforts, and facilitate blameless post-mortems and follow-up actions.

• Integrate SRE methodologies such as SLIs/SLOs, error budgets, and capacity planning into Core Infrastructure operations, collaborating closely with SRE teams.


⛳️ Requirements

• Over 5 years of experience in a DevOps, Platform/Infrastructure, or SRE role, demonstrating a successful history of managing large-scale, high-availability, high-performance systems in production environments.

• Extensive hands-on experience in designing cloud architecture on Google Cloud Platform (GCP), focusing on landing zones, networking, IAM, and high-availability topology.

• Strong expertise in Infrastructure-as-Code using Terraform, organizing large codebases across multiple environments, with GitOps as a foundational principle and a default mindset of least privilege.

• Proven ability to construct CI/CD pipelines for IaC, including automated plan/apply, code review, Policy-as-Code, drift detection, and secure rollout processes.

• Considerable experience in production using Kubernetes (preferably GKE) and in packaging/deploying workloads with Helm.

• Solid understanding of cloud fundamentals and L3/L4-L7 networking (VPCs, routing, load balancing, DNS, TLS, interconnects), along with the ability to debug cross-service connectivity issues.

• Practical experience with a modern observability stack including Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager, covering metrics, logs, traces, and alerting.

• Operator-level familiarity with data stores like PostgreSQL and message brokers such as RabbitMQ and RedPanda, capable of operating and troubleshooting them in production.

• A solid grasp of SRE practices, including SLOs/error budgets and capacity planning, alongside a Platform-as-a-Product mindset.

• Proficient in incident management, including incident declaration, structured debugging under pressure, escalation, accurate documentation, and post-mortems that foster real change.

• Willingness and capability to participate in a Follow-The-Sun on-call rotation during APAC hours, while effectively collaborating in a distributed, async-first team with strong written communication skills.

• Bonus: Familiarity with Policy-as-Code and IaC quality tools (e.g., OPA/Conftest, Checkov, tflint, Atlantis, or similar).

• Bonus: Experience in managing Terraform state, module registries, and versioning at scale across multiple teams.

• Bonus: Background in developing self-service developer platforms and internal golden paths (e.g., with Backstage, Tilt, or similar).

• Bonus: Experience with the Alloy collector and incident management tools such as Rootly.

• Bonus: Working knowledge of Go for automation and tooling.

• Bonus: Strong fundamentals in Linux (Debian/Ubuntu) and container technologies (Docker/containerd).

• Bonus: Experience with security and compliance in regulated environments (SOC 2, secrets management, audit logging).

• Bonus: Familiarity with trading, brokerage, or other regulated fintech sectors, particularly in low-latency systems.


🏝️ Benefits

• Competitive Salary & Stock Options

• Health Benefits

• New Hire Home-Office Setup: One-time USD $500

• Monthly Stipend: USD $150 per month via a Brex Card

People also viewed

Truelogic Software11 hours ago

Senior DevOps Engineer – Wealth Management Fintech

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Cadwell14 hours ago

Cloud Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $130k/year
ApplyView job
Raya14 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Arize AI16 hours ago

DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$150k – $185k/year
ApplyView job
Capgemini18 hours ago

DevOps Engineer

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Cast & Crew18 hours ago

Staff DevOps Engineer

US flagCalifornia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$190k – $235k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers