
Senior DevOps Engineer
Posted Aug 28

Posted Aug 28
This is a fully remote position, open to applicants in United States.
• Develop and enhance cloud architecture on GCP, encompassing networking, interconnects, IAM, and high-availability topology, implemented as Terraform code in accordance with GitOps principles.
• Create and manage CI/CD pipelines that plan, review, test, and safely deploy Infrastructure as Code (IaC) changes, incorporating Policy-as-Code guardrails, drift detection, and gradual rollout strategies.
• Promote Platform-as-a-Product by establishing self-service capabilities and optimal paths for engineers.
• Enhance observability across metrics, logs, traces, and alerting using tools like Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.
• Manage GKE clusters and infrastructure services, including Helm-packaged workloads, RabbitMQ, IBM MQ, and various data stores.
• Engage in the Follow-The-Sun on-call model; assess alerts, participate in incident declaration, lead debugging and escalation efforts, and facilitate blameless post-mortems and follow-up actions.
• Integrate SRE methodologies such as SLIs/SLOs, error budgets, and capacity planning into Core Infrastructure operations, collaborating closely with SRE teams.
• Over 5 years of experience in a DevOps, Platform/Infrastructure, or SRE role, demonstrating a successful history of managing large-scale, high-availability, high-performance systems in production environments.
• Extensive hands-on experience in designing cloud architecture on Google Cloud Platform (GCP), focusing on landing zones, networking, IAM, and high-availability topology.
• Strong expertise in Infrastructure-as-Code using Terraform, organizing large codebases across multiple environments, with GitOps as a foundational principle and a default mindset of least privilege.
• Proven ability to construct CI/CD pipelines for IaC, including automated plan/apply, code review, Policy-as-Code, drift detection, and secure rollout processes.
• Considerable experience in production using Kubernetes (preferably GKE) and in packaging/deploying workloads with Helm.
• Solid understanding of cloud fundamentals and L3/L4-L7 networking (VPCs, routing, load balancing, DNS, TLS, interconnects), along with the ability to debug cross-service connectivity issues.
• Practical experience with a modern observability stack including Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager, covering metrics, logs, traces, and alerting.
• Operator-level familiarity with data stores like PostgreSQL and message brokers such as RabbitMQ and RedPanda, capable of operating and troubleshooting them in production.
• A solid grasp of SRE practices, including SLOs/error budgets and capacity planning, alongside a Platform-as-a-Product mindset.
• Proficient in incident management, including incident declaration, structured debugging under pressure, escalation, accurate documentation, and post-mortems that foster real change.
• Willingness and capability to participate in a Follow-The-Sun on-call rotation during APAC hours, while effectively collaborating in a distributed, async-first team with strong written communication skills.
• Bonus: Familiarity with Policy-as-Code and IaC quality tools (e.g., OPA/Conftest, Checkov, tflint, Atlantis, or similar).
• Bonus: Experience in managing Terraform state, module registries, and versioning at scale across multiple teams.
• Bonus: Background in developing self-service developer platforms and internal golden paths (e.g., with Backstage, Tilt, or similar).
• Bonus: Experience with the Alloy collector and incident management tools such as Rootly.
• Bonus: Working knowledge of Go for automation and tooling.
• Bonus: Strong fundamentals in Linux (Debian/Ubuntu) and container technologies (Docker/containerd).
• Bonus: Experience with security and compliance in regulated environments (SOC 2, secrets management, audit logging).
• Bonus: Familiarity with trading, brokerage, or other regulated fintech sectors, particularly in low-latency systems.
• Competitive Salary & Stock Options
• Health Benefits
• New Hire Home-Office Setup: One-time USD $500
• Monthly Stipend: USD $150 per month via a Brex Card
Truelogic Software
Cadwell
Raya
Arize AI
Get handpicked remote jobs straight to your inbox weekly.