Remotery

Site Reliability Engineer

Posted Jul 26

This is a fully remote position, open to applicants in United States.

📋 Description

• Take responsibility for the daily reliability of our .NET/C# services operating on Windows.

• Engage in the on-call rotation for production trading systems and spearhead incident responses during service outages.

• Analyze production incidents, conduct root cause analyses, and implement preventive measures to eliminate recurring problems.

• Develop and maintain Grafana dashboards, Prometheus alerts, and operational health metrics across applications, infrastructure, and databases.

• Enhance .NET services with improved telemetry, metrics, logging, and visibility into service performance and customer impact.

• Define, execute, and oversee Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.

• Diagnose issues across .NET/C# applications, Windows Server, Aurora PostgreSQL databases, and AWS infrastructure.

• Enhance deployment safety, automate releases, and refine rollback strategies.

• Collaborate with developers to boost application operability, resilience, and fault isolation.

• Automate operational processes through scripting and infrastructure automation.

• Produce and maintain runbooks, operational documentation, and incident response protocols.

• Commit to continuous improvement in monitoring, alert quality, automation, and platform reliability.


⛳️ Requirements

• 3–5 years of experience in Site Reliability Engineering or a related field.

• Proficient in debugging and supporting .NET/C# applications in a production environment.

• Practical experience with Windows Server environments.

• Strong skills in PowerShell scripting.

• Familiarity with Python or Bash.

• Experience with Grafana, Prometheus, and Loki (or similar monitoring and observability tools).

• Knowledge of modern CI/CD pipelines.

• Understanding of deployment strategies, release automation, and rollback procedures.

• Experience with AWS.

• Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools.

• Proficient in troubleshooting and supporting Aurora PostgreSQL or other relational database systems.

• Practical knowledge of SLIs & SLOs, Error Budgets, Incident Response, Root Cause Analysis (RCA), and Alert Design.


🏝️ Benefits

• Work on mission-critical trading infrastructure that has a direct impact on customers.

• Tackle challenging reliability and scalability issues in a real-time environment.

• Build top-notch observability, automation, and deployment practices.

• Collaborate with seasoned engineers within a modern engineering culture.

• Shape reliability strategies and engineering best practices across the platform.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers