
Senior SRE / Platform Engineer
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in Germany.
• Advance our Kubernetes platform: Assess and incorporate technologies like Kubernetes Gateway API and service mesh patterns, while steering platform development across more than 10 engineering teams.
• Elevate observability standards: Propel organization-wide adoption of OpenTelemetry for distributed tracing and metrics, assisting teams in establishing significant SLOs.
• Design multi-region architecture and manage data residency: Facilitate our transition from an EU-centric presence to a global, multi-cloud framework that meets disaster recovery and data residency standards.
• Oversee cloud cost and efficiency at scale: Ensure our petabyte-scale infrastructure remains cost-effective, secure, and well-monitored.
• Enhance tooling: Develop self-service AWS account provisioning, implement guardrails, and create AI-assisted automations to enable engineering teams to manage infrastructure safely and efficiently at scale.
• A minimum of 5 years of professional experience in SRE, platform, or infrastructure engineering.
• Background in software development: You come from a software development background and transitioned to SRE. You write production-quality code in at least one language such as Python, Go, Rust, or Java.
• Strong understanding of systems: You possess a solid grasp of Linux internals and distributed systems, enabling you to debug intricate production issues.
• Practical cloud and infrastructure experience: Familiarity with AWS (or GCP), declarative infrastructure (Terraform), gitops workflows (ArgoCD), and container orchestration (Kubernetes).
• Experience with observability and reliability: You have worked with OpenTelemetry, Prometheus, distributed tracing, monitoring, and established meaningful SLOs/SLIs.
• Depth in production debugging: You can analyze complex failures, communicate effectively during incidents, and translate findings into lasting enhancements.
• Awareness of security and compliance: You recognize how infrastructure choices impact access control, auditability, disaster recovery, logging, and compliance standards like SOC 2.
• Excellent communication skills: You can articulate trade-offs to engineering teams and facilitate the adoption of improved platform practices with minimal friction.
• Join a committed, supportive team with limitless growth opportunities and potential for leadership.
• Quickly make an impact by sharing ideas and contributing to innovative, goal-driven projects.
• Work within a diverse, inclusive environment alongside colleagues from over 35 countries.
• Enjoy flexible hours and the option to work remotely from any location worldwide.
• Access comprehensive health benefits, retirement plans, paid time off, and wellness support.
• Relish fresh office lunches or gift cards as a remote employee.
• Advance your professional development through online/offline learning, language courses, and tech talks.
• Engage in team events, join support groups, and contribute to our ESG and DE&I initiatives.
• Participate in enjoyable team challenges and competitions to foster excitement and team spirit.
Quantiphi
Encompass Corporation
Shippit
Group O
Get handpicked remote jobs straight to your inbox weekly.