
Senior Site Reliability Engineer
Posted Jul 25

Posted Jul 25
This is a fully remote position, open to applicants in Florida, +1 more state.
• Design, develop, and manage highly available and scalable systems within AWS.
• Write, maintain, and review Terraform scripts for infrastructure provisioning and management.
• Take ownership of monitoring, alerting, and observability enhancements using tools such as Grafana, Pingdom, and Uptrends.
• Engage in a rotating on-call schedule, addressing production incidents and driving resolutions.
• Lead incident response efforts, conduct root cause analyses, and facilitate post-incident reviews with an emphasis on prevention and automation.
• Define and oversee Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.
• Construct and enhance CI/CD pipelines and operational workflows utilizing Azure DevOps and GitHub.
• Collaborate directly with application teams to enhance reliability, performance, and deployability.
• Automate manual operational processes to minimize toil.
• Maintain clear, actionable runbooks and documentation in Confluence.
• Monitor work, incidents, and operational improvements through Jira and ServiceNow.
• Mentor fellow engineers and contribute to establishing SRE standards and best practices.
• Over 5 years of practical experience in Site Reliability Engineering (SRE), DevOps, or Infrastructure Engineering roles.
• Strong production experience with AWS.
• Extensive hands-on experience with Terraform in real-world scenarios.
• Experience in operating monitoring and uptime platforms like Grafana, Pingdom, and Uptrends.
• Proficient in Linux systems, networking, and troubleshooting techniques.
• Experience in supporting production systems through incident response and on-call duties.
• Proficient with GitHub and contemporary Git workflows.
• Experience in building or maintaining CI/CD pipelines using Azure DevOps.
• Familiarity with IT Service Management (ITSM) and incident workflows utilizing ServiceNow.
• Excellent written communication skills with experience in documenting systems and processes in Confluence.
• Ability to work autonomously in a remote or hybrid setting.
• Competitive salary along with comprehensive benefits.
• Flexible work location options, including hybrid or fully remote arrangements.
• Genuine ownership of production systems and their reliability outcomes.
• A culture that prioritizes automation, continuous learning, and improvement.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.