Site Reliability Engineer

Posted 3 days ago

This is a fully remote position, open to applicants in California, +2 more states.

📋 Description

• Take ownership of uptime, service level agreements (SLAs), and service level objectives (SLOs) across API and platform services.

• Establish reliability and operational best practices for developers.

• Develop and maintain observability through logging, metrics, and tracing.

• Design and implement high-availability deployment and containerization strategies.

• Analyze and optimize Python applications for throughput, latency, and cost efficiency.

• Troubleshoot production issues at the operating system level using tools such as strace, perf, eBPF, ss, and gdb.

• Enhance deployment and continuous integration/continuous deployment (CI/CD) workflows.

• Engage in on-call rotation, lead incident response efforts, and conduct post-incident reviews.

• Identify necessary fixes and manage projects from conception to completion.

• Work with petabyte-scale data processing, customer management and billing, and query systems that support APIs.


⛳️ Requirements

• Mid-level or senior individual contributor.

• Full-time experience in Site Reliability Engineering (SRE), DevOps, or backend engineering, ideally within a trading firm, tech company, or high-growth startup.

• Practical experience with observability tools for logging, metrics, and tracing, including Prometheus, OpenTelemetry, VictoriaMetrics, Jaeger, Logstash, Loki, or Vector.

• Experience with containerization and high-availability deployment technologies such as Docker, Podman, Docker Compose, Docker Swarm, Kubernetes, or k3s.

• Strong command of Python, encompassing application development and performance tuning.

• Comfortable utilizing Linux debugging and profiling tools like strace, perf, eBPF, ss, and gdb.

• Proven track record of making a measurable impact in a recent position.

• Familiarity with alerting and incident response best practices is advantageous.

• Knowledge of configuration management or infrastructure-as-code tools such as Ansible or Terraform is beneficial.

• Experience with HTTP benchmarking, load testing, and capacity planning is a plus.

• Skills in database schema design and query optimization are desirable.

• Strong communication skills and a solid work ethic suitable for a remote environment.

• A keen interest in financial data or algorithmic trading.


🏝️ Benefits

• Equal employment opportunities and protections against discrimination.

• Employment accommodations available upon request.

• Option to opt-out of AI-powered Talent Matching.

People also viewed

Koniag Government Services2 days ago

Architect/DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
FP Markets (First Prudential Markets)2 days ago

Senior DevOps Engineer

AM flagArmenia, +4 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Modern Campus2 days ago

Senior DevOps Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
InRule2 days ago

Site Reliability Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Thumbtack2 days ago

Senior Software Engineer, Site Reliability Engineering

US flagUnited States, +38 more locationsFull-timeDevOps & Site Reliability Engineer (SRE)$179.4k – $272.8k/year
ApplyView job
Thumbtack2 days ago

Senior Software Engineer, Site Reliability Engineering

CA flagCanada, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)C$180.2k – C$233.2k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers