
Site Reliability Engineer – Low-Latency Trading Systems
Posted Sep 18

Posted Sep 18
This is a fully remote position, open to applicants in United States.
• Take charge of production reliability for real-time trading services, which encompass trading engines, execution gateways, market data ingestion, and PnL/reconciliation pipelines.
• Manage and enhance multi-region Kubernetes clusters on AWS EKS utilizing GitOps with Flux and SOPS-encrypted secrets.
• Develop and optimize Prometheus metrics and alerting, along with Datadog logs, dashboards, and SLOs.
• Enhance deployment safety through progressive rollouts, configuration reload behavior, and protective measures for live trading.
• Troubleshoot production incidents related to stale market data, exchange rate limits, WebSocket disconnects, order-lifecycle desynchronization, and trading-path latency regressions.
• Strengthen market data ingestion from Databento and venue-native REST/WebSocket feeds by implementing staleness detection, failover mechanisms, and replay capabilities.
• Create reconciliation and data-integrity tools across live gauges, Postgres, and S3 Parquet data lakes.
• Engage in an on-call rotation covering US equity market hours and 24/7 crypto venues.
• Over 5 years of experience in SRE, production engineering, or infrastructure roles, specifically supporting real-time or latency-sensitive systems.
• Strong programming skills in Go or Rust, with a willingness to work in both languages.
• Extensive hands-on experience with Kubernetes and AWS, particularly in managing stateful, latency-sensitive workloads in a production environment.
• Proficient in PromQL, structured-log analysis, and skilled in designing high-signal, low-noise alerts.
• Solid understanding of Linux internals and networking fundamentals.
• Capability to investigate p99 regressions through the kernel, NIC, or GC.
• Experience in incident response with clear communication during and post-incident.
• Must reside in the United States.
• Competitive compensation package that includes potential future token rights and/or equity, based on preferences.
• Comprehensive medical, vision, and dental coverage.
• Flexible vacation policy (PTO).
• Remote-first work environment.
• Opportunity to contribute to shaping the company’s vision, culture, and design practices.
• Collaboration with top-tier colleagues and industry-leading experts.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.