
Site Reliability Engineer
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in California, +2 more states.
• Take ownership of uptime, service level agreements (SLAs), and service level objectives (SLOs) across API and platform services.
• Establish reliability and operational best practices for developers.
• Develop and maintain observability through logging, metrics, and tracing.
• Design and implement high-availability deployment and containerization strategies.
• Analyze and optimize Python applications for throughput, latency, and cost efficiency.
• Troubleshoot production issues at the operating system level using tools such as strace, perf, eBPF, ss, and gdb.
• Enhance deployment and continuous integration/continuous deployment (CI/CD) workflows.
• Engage in on-call rotation, lead incident response efforts, and conduct post-incident reviews.
• Identify necessary fixes and manage projects from conception to completion.
• Work with petabyte-scale data processing, customer management and billing, and query systems that support APIs.
• Mid-level or senior individual contributor.
• Full-time experience in Site Reliability Engineering (SRE), DevOps, or backend engineering, ideally within a trading firm, tech company, or high-growth startup.
• Practical experience with observability tools for logging, metrics, and tracing, including Prometheus, OpenTelemetry, VictoriaMetrics, Jaeger, Logstash, Loki, or Vector.
• Experience with containerization and high-availability deployment technologies such as Docker, Podman, Docker Compose, Docker Swarm, Kubernetes, or k3s.
• Strong command of Python, encompassing application development and performance tuning.
• Comfortable utilizing Linux debugging and profiling tools like strace, perf, eBPF, ss, and gdb.
• Proven track record of making a measurable impact in a recent position.
• Familiarity with alerting and incident response best practices is advantageous.
• Knowledge of configuration management or infrastructure-as-code tools such as Ansible or Terraform is beneficial.
• Experience with HTTP benchmarking, load testing, and capacity planning is a plus.
• Skills in database schema design and query optimization are desirable.
• Strong communication skills and a solid work ethic suitable for a remote environment.
• A keen interest in financial data or algorithmic trading.
• Equal employment opportunities and protections against discrimination.
• Employment accommodations available upon request.
• Option to opt-out of AI-powered Talent Matching.
Koniag Government Services
FP Markets (First Prudential Markets)
Modern Campus
InRule
Get handpicked remote jobs straight to your inbox weekly.