
Site Reliability Engineer
Posted Jul 26

Posted Jul 26
This is a fully remote position, open to applicants in United States.
• Take responsibility for the daily reliability of our .NET/C# services operating on Windows.
• Engage in the on-call rotation for production trading systems and spearhead incident responses during service outages.
• Analyze production incidents, conduct root cause analyses, and implement preventive measures to eliminate recurring problems.
• Develop and maintain Grafana dashboards, Prometheus alerts, and operational health metrics across applications, infrastructure, and databases.
• Enhance .NET services with improved telemetry, metrics, logging, and visibility into service performance and customer impact.
• Define, execute, and oversee Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
• Diagnose issues across .NET/C# applications, Windows Server, Aurora PostgreSQL databases, and AWS infrastructure.
• Enhance deployment safety, automate releases, and refine rollback strategies.
• Collaborate with developers to boost application operability, resilience, and fault isolation.
• Automate operational processes through scripting and infrastructure automation.
• Produce and maintain runbooks, operational documentation, and incident response protocols.
• Commit to continuous improvement in monitoring, alert quality, automation, and platform reliability.
• 3–5 years of experience in Site Reliability Engineering or a related field.
• Proficient in debugging and supporting .NET/C# applications in a production environment.
• Practical experience with Windows Server environments.
• Strong skills in PowerShell scripting.
• Familiarity with Python or Bash.
• Experience with Grafana, Prometheus, and Loki (or similar monitoring and observability tools).
• Knowledge of modern CI/CD pipelines.
• Understanding of deployment strategies, release automation, and rollback procedures.
• Experience with AWS.
• Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools.
• Proficient in troubleshooting and supporting Aurora PostgreSQL or other relational database systems.
• Practical knowledge of SLIs & SLOs, Error Budgets, Incident Response, Root Cause Analysis (RCA), and Alert Design.
• Work on mission-critical trading infrastructure that has a direct impact on customers.
• Tackle challenging reliability and scalability issues in a real-time environment.
• Build top-notch observability, automation, and deployment practices.
• Collaborate with seasoned engineers within a modern engineering culture.
• Shape reliability strategies and engineering best practices across the platform.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.