Site Reliability Engineer – Low-Latency Trading Systems

Posted Sep 18

This is a fully remote position, open to applicants in United States.

📋 Description

• Take charge of production reliability for real-time trading services, which encompass trading engines, execution gateways, market data ingestion, and PnL/reconciliation pipelines.

• Manage and enhance multi-region Kubernetes clusters on AWS EKS utilizing GitOps with Flux and SOPS-encrypted secrets.

• Develop and optimize Prometheus metrics and alerting, along with Datadog logs, dashboards, and SLOs.

• Enhance deployment safety through progressive rollouts, configuration reload behavior, and protective measures for live trading.

• Troubleshoot production incidents related to stale market data, exchange rate limits, WebSocket disconnects, order-lifecycle desynchronization, and trading-path latency regressions.

• Strengthen market data ingestion from Databento and venue-native REST/WebSocket feeds by implementing staleness detection, failover mechanisms, and replay capabilities.

• Create reconciliation and data-integrity tools across live gauges, Postgres, and S3 Parquet data lakes.

• Engage in an on-call rotation covering US equity market hours and 24/7 crypto venues.


⛳️ Requirements

• Over 5 years of experience in SRE, production engineering, or infrastructure roles, specifically supporting real-time or latency-sensitive systems.

• Strong programming skills in Go or Rust, with a willingness to work in both languages.

• Extensive hands-on experience with Kubernetes and AWS, particularly in managing stateful, latency-sensitive workloads in a production environment.

• Proficient in PromQL, structured-log analysis, and skilled in designing high-signal, low-noise alerts.

• Solid understanding of Linux internals and networking fundamentals.

• Capability to investigate p99 regressions through the kernel, NIC, or GC.

• Experience in incident response with clear communication during and post-incident.

• Must reside in the United States.


🏝️ Benefits

• Competitive compensation package that includes potential future token rights and/or equity, based on preferences.

• Comprehensive medical, vision, and dental coverage.

• Flexible vacation policy (PTO).

• Remote-first work environment.

• Opportunity to contribute to shaping the company’s vision, culture, and design practices.

• Collaboration with top-tier colleagues and industry-leading experts.

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers