
Site Reliability Engineer
Posted Aug 25

Posted Aug 25
This is a fully remote position, open to applicants in United States.
• Develop, manage, and expand production systems within Microsoft Azure, encompassing App Service, Networking, WAF, CosmosDB, and associated infrastructure.
• Engage in an on-call rotation to triage, respond to, and resolve production incidents effectively.
• Lead or participate in blameless postmortems following incidents.
• Establish and monitor SLOs/SLIs and error budgets in collaboration with engineering and product teams.
• Design and sustain observability and APM tools, which include metrics, logs, tracing, dashboards, and alert systems.
• Optimize alert thresholds and runbooks to minimize noise and reduce mean time to resolution.
• Decrease operational toil through scripting, automated recovery systems, and standardized processes.
• Create and uphold CI/CD pipelines that facilitate safe and frequent software releases.
• Provision and manage infrastructure as code using Bicep.
• Adhere to change control and version control practices.
• Ensure application infrastructure complies with security and compliance standards such as SOC 2 and GovRAMP.
• Implement security best practices in infrastructure design and change management.
• Collaborate with the agile development team to convert business requirements into reliable technical solutions.
• Actively participate in technical design discussions.
• Generate clear documentation, which includes diagrams, runbooks, and architectural notes.
• Remain updated on Azure capabilities, industry standards, and SRE best practices, providing recommendations to the team.
• Undertake additional duties as assigned.
• A minimum of 3 years of experience in an IT Operations, DevOps, or SRE role.
• Practical technical experience with Microsoft Azure in a production setting.
• Familiarity with infrastructure as code, specifically Terraform and/or Bicep.
• Proficient in at least one scripting or programming language, such as Python or TypeScript.
• Experience working with REST and/or GraphQL APIs.
• Proven experience in defining and tracking KPIs/SLOs for web-based applications.
• Comfortable participating in an on-call rotation.
• Strong understanding of networking fundamentals, HTTP/S, and observability principles.
• Capable of evaluating various technical approaches and advising on the most effective solutions.
• Strong independent problem-solving capabilities coupled with effective collaboration in a team setting.
• Familiarity with the software development lifecycle and programming standards.
• Exemplary professional communication skills, particularly under pressure during incidents.
• Preferred: Experience with compliance audits (SOC 2 Type 2, GovRAMP).
• Preferred: Knowledge of security frameworks (NIST, ISO 27001).
• Preferred: AZ-104 certification or equivalent experience in Azure networking.
• Preferred: Experience with Node.js.
• Preferred: Familiarity with low-code platforms (Power Apps, Logic Apps).
• Preferred: Acquaintance with Scrum/Agile methodologies and associated tools (Confluence, JIRA, Git, Jenkins, Bamboo, TFS).
• Preferred: Ability to translate business requirements into application/site behavior modifications.
• Must be eligible to work in the U.S.
• Eligible for the company's commission plan.
• Profit sharing opportunities.
• Conditional employment offer subject to the successful completion of a drug screening and a fingerprint-based background check.
• Accommodations available during the application or selection process.
• Equal Opportunity Employer dedicated to fostering an inclusive environment.
Cisco
Stefanini Brasil
Study Now
Spring Financial
Get handpicked remote jobs straight to your inbox weekly.