Remotery

Senior DevOps Engineer

Posted Jul 19

This is a fully remote position, open to applicants in Japan.

📋 Description

• Develop and enhance our cloud architecture on Google Cloud Platform (GCP) - including networking, interconnects, Identity and Access Management (IAM), and high-availability topology - while codifying it entirely with Terraform, adhering to GitOps as a foundational principle.

• Create and manage the CI/CD pipelines that plan, review, test, and safely implement Infrastructure as Code (IaC) changes - establishing Policy-as-Code guardrails, drift detection, and progressive rollout to ensure infrastructure modifications are deployed as reliably as application code.

• Propel Platform-as-a-Product initiatives: establish self-service capabilities and streamlined pathways so engineers can provision what they require through a golden path instead of relying on hand-offs.

• Enhance our observability framework - encompassing metrics, logs, traces, and alerting through Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager - ensuring the platform is straightforward to manage and comprehend.

• Manage our Google Kubernetes Engine (GKE) clusters and the infrastructure services operating on them - including Helm-packaged workloads, message brokers (RabbitMQ, IBM MQ), and various data stores.

• Engage in our Follow-The-Sun on-call model: monitor and triage alerts, participate in and declare incidents, lead structured debugging and escalation, and conduct blameless post-mortems with actionable outcomes that effectively close the loop.

• Integrate Site Reliability Engineering (SRE) practices - including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets, as well as capacity planning - into how Core Infrastructure is built and maintained, collaborating closely with our SRE team.


⛳️ Requirements

• 5+ years of experience in a DevOps, Platform/Infrastructure, or SRE role, demonstrating a successful history of managing large-scale, high-availability, and high-performance systems in a production environment.

• Extensive hands-on experience in designing cloud architecture on Google Cloud Platform (GCP) as the primary cloud - including landing zones, networking, IAM, and high-availability topology.

• Strong Infrastructure-as-Code capabilities using Terraform, structuring large codebases across multiple environments, with GitOps as a foundational principle and a least-privilege approach as standard.

• Demonstrated experience in building CI/CD pipelines for IaC - automating plan/apply, conducting code reviews, implementing Policy-as-Code, drift detection, and ensuring safe rollouts.

• Significant hands-on experience with Kubernetes (preferably GKE) and packaging/deploying workloads using Helm.

• Solid understanding of cloud and L3/L4-L7 networking fundamentals (VPCs, routing, load balancing, DNS, TLS, interconnects) and proficiency in debugging cross-service connectivity issues.

• Practical experience with a modern observability stack - including Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager - focusing on metrics, logs, traces, and alerting.

• Operator-level knowledge of data stores such as PostgreSQL and message brokers (e.g., RabbitMQ, RedPanda), with the ability to operate and troubleshoot them in production settings.

• A solid understanding of SRE practices - such as SLOs, error budgets, and capacity planning - along with a Platform-as-a-Product perspective.

• Comprehensive understanding of incident management processes: participating in and declaring incidents, structured debugging under pressure, escalation, thorough documentation, and conducting post-mortems that drive meaningful changes.

• Willingness and capability to participate in a Follow-The-Sun on-call rotation during APAC hours, and to collaborate effectively within a distributed, asynchronous team with strong written communication skills.


🏝️ Benefits

• Competitive Salary & Stock Options

• Health Benefits

• New Hire Home-Office Setup: One-time USD $500

• Monthly Stipend: USD $150 per month via a Brex Card

People also viewed

The CodestJul 26

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
IRIUMJul 26

Ingeniero/a Cloud DevOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€33k – €40k/year
ApplyView job
SólidesJul 26

Senior DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ResilincJul 25

Junior/Senior Site Reliability Engineer – Night Shift

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity GroupJul 25

Senior SRE / DevOps Engineer

Anywhere in the WorldFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
HOESSLER & HOESSLERJul 25

DevOps Software Engineer – Career Ambitions

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€65k – €75k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers