
Site Reliability Engineer
Posted Jul 29

Posted Jul 29
This is a fully remote position, open to applicants in Illinois.
• Oversees post-incident investigations for the Site Reliability team.
• Conducts thorough post-incident analyses to determine root causes and formulates preventive measures.
• Prepares clear and insightful Root Cause Analyses (RCAs) for client delivery.
• Trains colleagues on effectively utilizing observability tools during incident and performance investigations.
• Ensures visibility for all stakeholders throughout the entire Site Reliability process.
• Works with cross-functional teams to implement system improvements that boost scalability and stability.
• Creates client-oriented dashboards and alerts to proactively detect performance issues.
• Monitors and consistently enhances our time-to-resolution metrics.
• Maintains and configures essential observability tools to guarantee optimal performance and availability of key metrics and data for incident response and performance investigations.
• Provides actionable feedback to the Observability and Engineering teams to enhance MELT and development practices.
• Aids in the creation of automation tools to streamline incident response.
• Proactively works to avert incidents and minimize their impact on our platform.
• Collaborates with the broader Cloud Operations, SRE, Engineering teams, and the business as a whole to improve our SaaS platforms.
• Additional duties as assigned.
• Bachelor's degree in Computer Science or a related field (or equivalent experience).
• Over 5 years of demonstrated experience in a Site Reliability Engineering position.
• In-depth knowledge of SRE best practices and incident management protocols.
• Extensive experience with New Relic, Data Dog, SumoLogic, or similar observability tools.
• Proficient in reading and writing code (e.g., JavaScript, .NET, SQL).
• Understanding of cloud platforms (e.g., AWS, Azure) and architectural patterns.
• Excellent problem-solving abilities with a data-driven approach to incident analysis.
• Prior experience in a Public Cloud environment (AWS preferred).
• Experience in troubleshooting C#/.NET-based web applications to identify bugs and performance issues.
• Strong understanding of SaaS operations.
• Ability to thrive in ambiguous situations and varying levels of operational maturity.
• Advanced written and verbal communication skills.
• Preferred troubleshooting skills in Windows and SQL Server.
• Familiarity with Continuous Integration and Continuous Delivery (CI/CD) pipelines preferred.
• Experience in an Infrastructure as Code (IaC) environment preferred.
• Medical and Dental coverage available for employees, dependents, domestic partners, and spouses.
• Paid Time Off – Flexible options plus 10 paid company holidays where applicable.
• Fully Paid by Origami Risk – Vision insurance, Short & Long-Term Disability Insurance, and Basic Life Insurance.
• Generous family leave options—including adoption and foster care placements.
• Pre-Tax Savings Accounts – Flexible Spending Account, Health Savings Account, Commuter Benefits, Dependent Care Savings Account.
• Retirement Savings – 401(k) with company match up to 4%.
• Employee Assistance Program (EAP) – Confidential & Free support offered to colleagues facing personal or work-related challenges.
• Education Assistance Program – to support colleagues pursuing industry and role-specific certifications.
• Wellness Benefits – reimbursement program to encourage healthy habits and enhance colleague productivity and stress management.
• Additional coverages available – Pet Insurance, Critical Illness Insurance, and Voluntary Life & AD&D coverage.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.