
Senior Site Reliability Engineer
Posted Jun 18

Posted Jun 18
This is a fully remote position, open to applicants in Idaho.
• Act as the primary owner for the reliability, availability, performance, operability, and capacity of one or more production services.
• Deploy, operate, maintain, and consistently enhance production services within Autodesk GovCloud environments.
• Collaborate with engineering teams to ensure services are designed with a focus on reliability, scalability, security, and operability.
• Define and implement reliability practices, including SLOs/SLIs, error budgets, production readiness reviews, service evaluations, and operational health assessments.
• Create automation to enhance deployment safety, operational efficiency, incident response, and service recovery.
• Design, develop, and sustain software, automation, and tools that bolster the reliability, scalability, and efficiency of production systems.
• Implement and enhance monitoring, alerting, logging, tracing, and observability capabilities across the supported services.
• Lead and engage in incident response, troubleshooting, and post-incident reviews focused on learning and continuous improvement.
• Develop and maintain operational documentation, runbooks, and recovery procedures.
• Scale and improve resilience testing and Gameday practices to validate system behavior, recovery capabilities, and operational readiness.
• Continuously identify and eliminate operational toil through software engineering, automation, and process enhancements.
• Ensure supported services remain compliant with Autodesk's security, privacy, and regulatory requirements, including FedRAMP and related controls where applicable.
• Participate in a 24x7 on-call rotation for production services.
• Operate effectively in a fast-paced environment while contributing to the establishment and maturation of operational excellence practices for Autodesk GovCloud.
• B.S. or higher in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
• Over 7 years of experience in Site Reliability Engineering, Software Engineering, Platform Engineering, Cloud Infrastructure, or Production Operations.
• Experience in operating and supporting customer-facing production services in large-scale cloud environments.
• Strong grasp of reliability engineering principles, including SLOs/SLIs, observability, incident management, capacity planning, production readiness, and automation.
• Familiarity with AWS, Azure, or other public cloud platforms.
• Proficient in developing automation using languages such as Python, Go, Java, PowerShell, Bash, or similar.
• Experience with Infrastructure as Code, CI/CD pipelines, deployment automation, and modern cloud operations practices.
• Understanding of security, compliance, and operational risk management in production environments.
• Excellent written and verbal communication skills.
• Health and financial benefits.
• Time away.
• Everyday wellness.
CVS Health
Devoteam
Aspirion
Goodgame Studios
Get handpicked remote jobs straight to your inbox weekly.