
Site Reliability Engineer III
Posted 13 hours ago

Posted 13 hours ago
This is a fully remote position, open to applicants in Colorado, +6 more states.
• Develop and sustain infrastructure that empowers developers to deploy reliably at scale.
• Oversee the infrastructure platform, deployment automation, and observability using infrastructure-as-code methodologies.
• Implement, monitor, and manage highly available systems leveraging Terraform, CockroachDB, GKE/Kubernetes, Cloud SQL, Bigtable, Google Composer/Airflow, Cloud Storage, BigQuery, Pub/Sub, Cloud Run, and other GCP services.
• Enhance and manage a substantial, mature Terraform codebase.
• Evaluate systems and propose enhancements to performance, availability, and cost-effectiveness.
• Automate manual processes to reduce toil.
• Create and sustain integrations with monitoring and alerting systems such as Google Cloud Monitoring, Prometheus, OpenTelemetry, Checkly, and Rootly.
• Lead incident response best practices for engineering teams on-call.
• Engage in the SRE team's on-call rotation for core infrastructure support.
• Collaborate on architectural strategies and directions for services and initiatives.
• Report directly to the Principal Site Reliability Engineer.
• B.S. or M.S. in computer science or a related discipline, or equivalent experience.
• Minimum of 5+ years of experience, including at least 3+ years supporting production systems.
• Strong passion and experience with Kubernetes, networking, and infrastructure-as-code practices.
• Proficiency with Terraform/OpenTofu.
• Familiarity with at least one major cloud platform.
• Assess technologies and solutions based on merit, reliability, performance, and ease of debugging.
• Practical experience with SQL, NoSQL, and object storage, including making selections based on access patterns and scalability requirements.
• Strong foundational knowledge in computer science.
• Demonstrated ownership of work and platform responsibilities.
• Familiarity with Google Cloud Platform.
• Excellent troubleshooting and issue analysis capabilities.
• Experience with high-throughput, low-latency services.
• Proven experience collaborating with a distributed team.
• Knowledge of IAM, auditing, and cloud security management.
• Experience with GIS mapping systems and tiles.
• Familiarity with Claude Code.
• Experience with Airflow or similar ETL systems.
• Competitive salaries, annual bonuses, equity, and opportunities for advancement.
• Comprehensive health benefits, including a no-monthly-cost medical plan.
• Parental leave plan of fully paid 5 or 13 weeks.
• 401k matching at 100% for the first 3% you contribute and 50% for contributions between 3-5%.
• Company-wide outdoor adventures and exceptional outdoor industry perks.
• Annual “Get Out, Get Active” funds to support your active lifestyle both in and outside the gym.
• Flexible time-off package that encompasses PTO, STO, VTO, quiet weeks, and floating holidays.
• Common share options with a vesting schedule.
• Potential annual bonus of 10% based on company performance.
Akamai Technologies
BeyondTrust
Cencora
PrizePicks
Get handpicked remote jobs straight to your inbox weekly.