Site Reliability Engineer

Posted Aug 25

This is a fully remote position, open to applicants in United States.

📋 Description

• Develop, manage, and expand production systems within Microsoft Azure, encompassing App Service, Networking, WAF, CosmosDB, and associated infrastructure.

• Engage in an on-call rotation to triage, respond to, and resolve production incidents effectively.

• Lead or participate in blameless postmortems following incidents.

• Establish and monitor SLOs/SLIs and error budgets in collaboration with engineering and product teams.

• Design and sustain observability and APM tools, which include metrics, logs, tracing, dashboards, and alert systems.

• Optimize alert thresholds and runbooks to minimize noise and reduce mean time to resolution.

• Decrease operational toil through scripting, automated recovery systems, and standardized processes.

• Create and uphold CI/CD pipelines that facilitate safe and frequent software releases.

• Provision and manage infrastructure as code using Bicep.

• Adhere to change control and version control practices.

• Ensure application infrastructure complies with security and compliance standards such as SOC 2 and GovRAMP.

• Implement security best practices in infrastructure design and change management.

• Collaborate with the agile development team to convert business requirements into reliable technical solutions.

• Actively participate in technical design discussions.

• Generate clear documentation, which includes diagrams, runbooks, and architectural notes.

• Remain updated on Azure capabilities, industry standards, and SRE best practices, providing recommendations to the team.

• Undertake additional duties as assigned.


⛳️ Requirements

• A minimum of 3 years of experience in an IT Operations, DevOps, or SRE role.

• Practical technical experience with Microsoft Azure in a production setting.

• Familiarity with infrastructure as code, specifically Terraform and/or Bicep.

• Proficient in at least one scripting or programming language, such as Python or TypeScript.

• Experience working with REST and/or GraphQL APIs.

• Proven experience in defining and tracking KPIs/SLOs for web-based applications.

• Comfortable participating in an on-call rotation.

• Strong understanding of networking fundamentals, HTTP/S, and observability principles.

• Capable of evaluating various technical approaches and advising on the most effective solutions.

• Strong independent problem-solving capabilities coupled with effective collaboration in a team setting.

• Familiarity with the software development lifecycle and programming standards.

• Exemplary professional communication skills, particularly under pressure during incidents.

• Preferred: Experience with compliance audits (SOC 2 Type 2, GovRAMP).

• Preferred: Knowledge of security frameworks (NIST, ISO 27001).

• Preferred: AZ-104 certification or equivalent experience in Azure networking.

• Preferred: Experience with Node.js.

• Preferred: Familiarity with low-code platforms (Power Apps, Logic Apps).

• Preferred: Acquaintance with Scrum/Agile methodologies and associated tools (Confluence, JIRA, Git, Jenkins, Bamboo, TFS).

• Preferred: Ability to translate business requirements into application/site behavior modifications.

• Must be eligible to work in the U.S.


🏝️ Benefits

• Eligible for the company's commission plan.

• Profit sharing opportunities.

• Conditional employment offer subject to the successful completion of a drug screening and a fingerprint-based background check.

• Accommodations available during the application or selection process.

• Equal Opportunity Employer dedicated to fostering an inclusive environment.

People also viewed

Cisco1 day ago

Site Reliability Engineer – FEDRAMP

US flagAlabama, +25 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$128.6k – $184.9k/year
ApplyView job
Stefanini Brasil1 day ago

Integration Architect – DevSecOps, OpenShift, Kubernetes

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Study Now1 day ago

DevOps Engineer

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)£1,170 – £1,950/month
ApplyView job
Spring Financial1 day ago

DevOps Engineer II – Contract

MX flagMexico OnlyFreelanceDevOps & Site Reliability Engineer (SRE)$612k – $857k/year
ApplyView job
Endeavor1 day ago

DevOps Lead

US flagConnecticut, +3 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$112.5k – $150k/year
ApplyView job
MOXFIVE1 day ago

Senior DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$110k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers