
Staff Software Engineer, Platform Infrastructure
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Ireland.
• Establish Site Reliability Engineering (SRE) as a key discipline within the Platform Infrastructure Engineering team and the wider Internal Platform Group.
• Define and promote the adoption of Service Level Objectives (SLOs) and Service Level Indicators (SLIs), along with reliability standards and incident response methodologies.
• Oversee and enhance observability patterns utilizing Prometheus metrics, OpenTelemetry tracing, and structured logging.
• Design and implement service templates, Terraform modules, and infrastructure patterns for Go APIs, Command Line Interfaces (CLIs), Cloud Run services, and Google Kubernetes Engine (GKE) workloads.
• Contribute to migration efforts and standardization initiatives involving GCP Secret Manager, GitHub Actions, Cloud Build, and Cloud SQL.
• Advocate for GCP-native managed and serverless services as alternatives to self-hosted infrastructure.
• Implement security-first practices related to Identity-Aware Proxy (IAP), Identity and Access Management (IAM), VPC design, networking, and identity management.
• Support compliance efforts across various frameworks including SOC 2, ISO 27001, PCI DSS, and others.
• Mentor engineers to elevate the technical capabilities across the Platform Infrastructure Engineering (PIE) team and related groups.
• Manage development, testing, operations, and support for systems built under a comprehensive DevOps model.
• Participate in the on-call rotation following an initial onboarding period of approximately 3–6 months.
• Alleviate on-call responsibilities through automation, runbooks, and improvements in reliability.
• Collaborate with teams across North America and Europe during cross-timezone standups and incident responses.
• In-depth knowledge of SLOs, SLIs, error budgets, toil reduction, and the operationalization of SRE practices.
• Proficient in GCP services, including Cloud Run, GKE, Cloud SQL, GCP Secret Manager, IAP/IAM, and networking.
• Hands-on experience in implementing metrics, traces, and logs using Prometheus, OpenTelemetry, and structured logging.
• Practical experience utilizing Grafana in a production environment.
• Strong skills in Terraform, including the design of reusable modules.
• Experience in setting the technical direction for a platform or a team.
• Ability to translate ambiguous reliability objectives into concrete architectural solutions.
• Effective in influencing engineers who do not report directly to you.
• Capable of clearly communicating reliability risks, architectural decisions, and lessons learned from incident postmortems.
• 8+ years of experience in building and operating production systems, with substantial experience in infrastructure, SRE, or platform engineering.
• Proven SRE experience in implementing SLO frameworks, incident management, fostering an on-call culture, and achieving measurable reliability improvements.
• Strong practical proficiency in Go, version 1.21 or higher.
• Familiarity with Python as a secondary programming language.
• Extensive practical experience with GCP; comparable experience with other cloud platforms will be considered if there is a clear willingness to learn GCP.
• Hands-on experience in building and maintaining Terraform modules.
• Solid operational expertise with GKE or similar Kubernetes platforms.
• Production experience using Grafana, Prometheus, and OpenTelemetry.
• Practical experience with IAP, IAM, VPC design, and compliance frameworks such as SOC 2 or PCI DSS.
• Experience in developing reusable infrastructure patterns or developer platforms.
• Visa sponsorship is currently not available.
• Competitive compensation package with equity options.
• Comprehensive vacation policy offering 28 days of holiday.
• Private medical and dental insurance coverage.
• Life and critical illness insurance benefits.
• Attractive workplace pension scheme.
• Access to an Employee Resource Platform.
• High-quality equipment provided for work purposes.
• Monthly allowance for wellness and reading materials.
• Access to LinkedIn Learning for ongoing professional development.
• Opportunities for team-based and company-wide events and activities.
EasyLlama - HR & Compliance Training For Modern Teams
Coinbase
Stitch Fix
Ambry Genetics
Get handpicked remote jobs straight to your inbox weekly.