
Principal Cloud Engineer
Posted 23 hours ago

Posted 23 hours ago
This is a fully remote position, open to applicants in United States.
• Design, develop, and sustain automation and tools to minimize operational challenges.
• Enhance reliability practices in collaboration with Product & Engineering leaders.
• Engage in the incident-management lifecycle, covering detection, engagement, escalation, mitigation, stakeholder communication, post-incident review, and follow-up on corrective actions.
• Analyze alert, incident, support, and SLO trends to transition from reactive measures to proactive risk management.
• Establish feedback loops that connect incident insights to engineering standards, service maturity, product priorities, and vendor actions.
• Work alongside application engineering teams to enhance developer experience and reduce operational toil.
• Collaborate with service owners to define production readiness standards.
• Offer SRE expertise on capacity planning, resilience testing, game days, disaster recovery readiness, and modernization of fragile or legacy workloads.
• Document essential customer workflows, set health expectations, identify dependencies, and align reliability investments with business objectives.
• Engage in cross-team reliability initiatives and influence results without direct authority.
• Cultivate strategic vendor partnerships in observability, incident response, and cloud infrastructure.
• Implement AI-assisted and agentic workflows for alert triage, incident mitigation, postmortems, trend analysis, capacity planning, SLO analysis, and self-service knowledge.
• Enhance service metadata, observability data, incident records, runbooks, architecture documentation, and the quality of corrective actions.
• Proven experience in a production-facing SRE role within a complex SaaS environment.
• Strong technical acumen across distributed systems, multi-cloud environments, Kubernetes, networking, infrastructure technologies, GitOps, and CI/CD.
• Demonstrated experience in establishing and advancing SRE principles and practices, including SLOs, error budgets, observability, capacity planning, incident response, and toil reduction.
• Proficient in utilizing observability data to address high-severity incidents and investigate root causes during postmortems.
• Experience leading high-pressure incidents while effectively communicating with technical teams, executives, customer-facing stakeholders, and third-party vendors.
• Capability to participate in a 12-hour follow-the-sun on-call rotation.
• Ability to responsibly integrate AI-assisted and agentic workflows while ensuring qualified personnel remain involved in decision-making for production-impacting actions.
• Participation in Seismic's incentive plans in addition to base salary.
• Remote work arrangement.
• 12-hour follow-the-sun on-call rotation.
Salve.Inno
General Dynamics Information Technology
Seismic
Get handpicked remote jobs straight to your inbox weekly.