
Senior Cloud Engineer – SRE
Posted Sep 15

Posted Sep 15
This is a fully remote position, open to applicants in United States.
• Design, develop, and manage automation and tools aimed at minimizing operational toil.
• Collaborate to enhance reliability practices alongside Product and Engineering leaders.
• Engage in and continually refine the incident-management lifecycle, which encompasses detection, escalation, mitigation, communication, post-incident analysis, and corrective measures.
• Take part in a 12-hour follow-the-sun on-call rotation as a member of the Global SRE team.
• Leverage alert, incident, support, and SLO trends to promote proactive risk mitigation.
• Establish feedback mechanisms that link incident learnings to engineering standards, service maturity, product priorities, and vendor actions.
• Collaborate with application engineering teams to enhance developer experience and minimize toil.
• Verify that production-readiness standards are fulfilled prior to release stages.
• Offer SRE insights on capacity planning, resilience testing, game day exercises, disaster-recovery preparedness, and the modernization of outdated or fragile workloads.
• Document essential customer workflows, define health expectations, identify dependencies, and align reliability investments with business objectives.
• Engage in cross-team reliability initiatives and influence outcomes without direct authority.
• Cultivate strategic vendor relationships in observability, incident response, and cloud infrastructure.
• Embrace AI-assisted and agentic workflows for alert triage, incident resolution, postmortems, trend analysis, capacity planning, SLO evaluation, and self-service knowledge management.
• Ensure qualified human oversight for actions impacting production.
• Enhance service metadata, observability data, incident records, runbooks, architectural documentation, and the quality of corrective actions.
• Proven experience in a production-facing SRE role within a complex SaaS environment.
• Strong technical acumen across distributed systems, multi-cloud environments, Kubernetes, networking, infrastructure technologies, GitOps, and CI/CD.
• Experience in establishing and advancing SRE practices, including SLOs, error budgets, observability, capacity planning, incident response, and eliminating toil.
• Expertise in utilizing observability data to address high-severity incidents and conduct root-cause analysis during postmortems.
• Experience leading high-pressure incidents and effectively communicating with technical teams, executives, customer-facing stakeholders, and third-party vendors.
• A curious, respectful, and collaborative approach when discussing solutions, processes, and proposals.
• Eligibility for an incentive plan in addition to the base salary.
• Comprehensive benefits and perks for personal well-being (specific benefits vary by country; details available on the Global Benefits page).
• Reasonable accommodations provided during the application or recruitment process.
Salve.Inno
General Dynamics Information Technology
Seismic
Get handpicked remote jobs straight to your inbox weekly.