
Senior Cloud Engineer – SRE
Posted Sep 15

Posted Sep 15
This is a fully remote position, open to applicants in United States.
• Design, develop, and sustain automation and tools aimed at minimizing operational toil.
• Enhance reliability practices in collaboration with Product and Engineering leadership.
• Engage in the incident-management lifecycle, covering detection, engagement, escalation, mitigation, stakeholder communication, post-incident review, and follow-up on corrective actions.
• Utilize alert, incident, support, and SLO trends to promote proactive risk mitigation.
• Establish feedback loops that link incident insights to engineering standards, service maturity, product priorities, and vendor actions.
• Work alongside application engineering teams to enhance developer experience and decrease toil.
• Collaborate with service owners to ensure production-readiness standards are met prior to release phases.
• Provide SRE expertise on capacity planning, resilience testing, game days, disaster recovery readiness, and the modernization of fragile or legacy workloads.
• Document essential customer workflows, outline health expectations, identify dependencies, and align reliability investments with business objectives.
• Engage in cross-team reliability initiatives and influence outcomes without direct authority.
• Foster strategic vendor partnerships in observability, incident response, and cloud infrastructure.
• Implement AI-assisted workflows for alert triage, incident mitigation, postmortems, trend analysis, capacity planning, SLO analysis, and self-service knowledge management.
• Enhance service metadata, observability data, incident records, runbooks, architecture documentation, and the quality of corrective actions.
• Proven experience in a production-facing SRE role within a complex SaaS environment.
• Strong technical acumen in distributed systems, multi-cloud environments, Kubernetes, networking, infrastructure technologies, GitOps, and CI/CD.
• Demonstrated experience in establishing and advancing SRE principles and practices, including SLOs, error budgets, observability, capacity planning, incident response, and toil reduction.
• Proficiency in utilizing observability data to address high-severity incidents and conduct root-cause analysis during postmortems.
• Experience in leading high-pressure incidents and effectively communicating with technical teams, executives, customer-facing stakeholders, and third-party vendors.
• Willingness to participate in a 12-hour follow-the-sun on-call rotation.
• Capability to responsibly implement AI-assisted and agentic workflows while ensuring qualified individuals remain involved in decision-making for production-impacting actions.
• Eligibility for an incentive plan in addition to the base salary.
• An inclusive workplace culture.
• A globally distributed work environment.
Salve.Inno
General Dynamics Information Technology
Seismic
Get handpicked remote jobs straight to your inbox weekly.