
Senior Infrastructure Engineer, SRE
Posted 8 hours ago

Posted 8 hours ago
This is a fully remote position, open to applicants in California, +2 more states.
• Develop and enhance the reliability and resilience of systems and services.
• Set up Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for essential services and user experiences.
• Collaborate with ownership teams to review SLIs, SLOs, and error budgets.
• Take charge of and advance the disaster recovery strategy, including recovery objectives, failover and restoration pathways, and regular practice exercises.
• Work alongside product engineering teams to assist them in managing and operating their services.
• Advance observability platforms and standards in areas such as metrics, tracing, and logging.
• Enhance instrumentation, alert quality, and observability cost efficiency.
• Strengthen incident management practices by adjusting paging thresholds, maintaining runbooks, and following up on post-mortem actions.
• Contribute to Cloud Infrastructure initiatives, including infrastructure build-outs and managing platform backlogs.
• Engage in a shared on-call rotation, one week every six weeks.
• Collaborate with engineering and internal support teams.
• Lead or assist in modernizing reliability and observability, internal tooling, game days, chaos experiments, disaster recovery exercises, and optimizing observability costs.
• Over 5 years of practical experience in cloud or infrastructure engineering, with significant involvement in reliability and production operations at scale.
• Established SLIs and SLOs for actual production services.
• Hands-on experience with a production observability platform; Datadog is highly preferred.
• Proficient in coding with Python, Go, TypeScript, or similar languages.
• Experience with Terraform in production environments.
• Comfortable operating within AWS.
• Developed or managed a disaster recovery plan, including recovery goals, failover and restoration procedures, and practice drills.
• On-call experience for services that you have helped build.
• Understanding of effective alerting practices.
• Health, Dental & Vision Plans.
• Competitive Pay.
• 401k Matching.
• Unlimited PTO.
• Daily Lunch (in-office only).
• Snacks & Coffee (in-office only).
• Commuter benefits (in-office only).
• Bonus.
Leidos
Precise Software Solutions, Inc.
Coinbase
Hitachi Solutions America
Get handpicked remote jobs straight to your inbox weekly.