
Senior DevOps Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in United States.
• Oversee the design of reliability, availability, scalability, and recovery for essential systems.
• Establish and refine SLOs, SLIs, and error budget methodologies across various services.
• Recognize systemic reliability threats and lead cross-team corrective actions.
• Shape application and platform architecture to enhance operational results.
• Act as the technical lead during significant incidents and intricate outages.
• Conduct root cause analyses and propose remedial measures.
• Enhance incident response protocols, tools, and documentation.
• Create and execute automation to minimize operational toil on a large scale.
• Develop and sustain shared SRE tools and platforms.
• Establish engineering standards for reliability-centric code and operational methodologies.
• Assess and refine CI/CD, deployment, and rollback strategies.
• Collaborate with Release and Change Management to streamline release processes through automation.
• Conduct risk assessments for significant changes and releases.
• Ensure compliance requirements are fulfilled while preserving engineering speed.
• Act as the reliability authority for decisions regarding release readiness.
• Provide mentorship to junior SREs and engineers through technical advice and evaluations.
• Foster a culture of reliability throughout engineering and product teams.
• Bachelor’s degree in Computer Science, Engineering, or a related discipline.
• 6–10+ years of experience in SRE, software engineering, platform, or DevOps positions.
• Hands-on experience in conducting root cause analyses on incidents and documenting SRE systems and their utilization.
• Proficient programming skills with professional experience in several programming languages.
• Extensive experience with AWS and distributed systems.
• In-depth knowledge of observability, ITSM, and principles of reliability engineering.
• Demonstrated ability to operate effectively in complex, regulated settings.
• Experience with observability tools encompassing metrics, logs, and tracing.
• Familiarity with CI/CD pipelines and deployment automation techniques.
• Background in root cause analysis investigations and documentation.
• Knowledge of containerization and orchestration technologies.
• Strong troubleshooting and analytical capabilities.
• Preferred: experience with IaC (CloudFormation), application frameworks, application servers/containers, relational and non-relational databases, ORM/drivers, Agile/Scrum, ITSM tools, ITIL-aligned change and release processes, and security compliance frameworks.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible work hours and remote work opportunities.
• Professional development and continuous learning programs.
• Generous paid time off and holidays.
• Collaborative and innovative work environment.
Fundrise
Verity Group
Méliuz
ODILO
Get handpicked remote jobs straight to your inbox weekly.