
Principal Cloud Engineer
Posted 19 hours ago

Posted 19 hours ago
This is a fully remote position, open to applicants in United States.
• Design, develop, and maintain automation tools to minimize operational toil.
• Enhance reliability practices in collaboration with Product & Engineering leaders.
• Engage in the incident-management lifecycle, encompassing detection, engagement, escalation, mitigation, stakeholder communication, post-incident review, and follow-through on corrective actions.
• Ensure that incident practices prioritize customer needs, are data-driven, blameless, and consistently applied across teams.
• Analyze alert, incident, support, and SLO trends to transition from reactive responses to proactive risk management.
• Establish feedback loops that connect incident learning with engineering standards, service maturity, product priorities, and vendor actions.
• Work alongside application engineering teams to enhance the developer experience and decrease operational toil.
• Collaborate with service owners to establish production-readiness standards, capacity planning, resilience testing, game days, disaster-recovery preparedness, and modernization of fragile or legacy workloads.
• Document essential customer workflows, define health expectations, identify dependencies, and align reliability investments with business goals.
• Engage in cross-team reliability initiatives and influence results without direct authority.
• Build relationships with vendors in observability, incident response, and cloud infrastructure.
• Embrace AI-assisted and autonomous workflows for alert triage, incident mitigation, postmortems, trend analysis, capacity planning, SLO analysis, and self-service knowledge, ensuring qualified personnel remain involved in decision-making.
• Enhance service metadata, observability data, incident records, runbooks, architecture documentation, and the quality of corrective actions.
• Proven experience in a production-facing SRE role supporting a complex SaaS environment.
• Strong technical acumen in distributed systems, multi-cloud environments, Kubernetes, networking, infrastructure technologies, GitOps, and CI/CD.
• Demonstrated experience in establishing and advancing SRE principles and practices, including SLOs, error budgets, observability, capacity planning, incident response, and toil elimination.
• Proficiency in utilizing observability data to manage high-severity incidents and investigate root causes during postmortems.
• Experience leading teams through high-pressure incidents and communicating effectively with technical teams, executives, customer-facing stakeholders, and third-party vendors.
• Willingness to participate in a 12-hour follow-the-sun on-call rotation.
• Ability to collaborate effectively across Product & Engineering, Security, and customer-facing teams.
• Benefits and perks catering to the entire self; availability may vary by country.
• Eligibility for participation in Seismic's incentive plans in addition to the base salary.
• Reasonable accommodations are available during the application or recruitment process.
Salve.Inno
General Dynamics Information Technology
Latitude IT Solutions | SDVOSB
Get handpicked remote jobs straight to your inbox weekly.