Remotery

Site Reliability Engineer

Posted Jul 31

This is a fully remote position, open to applicants in United States, +7 more states.

📋 Description

• Take ownership of reliability objectives across our backend/API, worker services, applications, and CV pipeline; MTD, MTM, MTR, and ensure thorough investigation of root causes.

• Enhance our monitoring and alerting systems, creating automated remediation processes, so that on-call responsibilities are managed through automation rather than increasing headcount.

• Collaborate with our engineering teams to develop agents that can triage alerts and manage routine remediation tasks.

• Strengthen and optimize our GCP infrastructure (Cloud Run, Cloud SQL, GCS) for both cost-effectiveness and performance as load increases.

• Manage database scaling and performance, including connection pooling, query optimization, indexing, read replicas, and capacity planning, to prevent Postgres from becoming a bottleneck as data volume expands.

• Enhance the reliability of our ML training and monitoring infrastructure in collaboration with the CV/ML team.

• Conduct blameless postmortems and focus on implementing solutions for root causes rather than merely addressing symptoms.

• Participate in the on-call rotation.


⛳️ Requirements

• A minimum of 4 years of experience in an SRE, infrastructure, or backend engineering position with production on-call responsibilities.

• Extensive experience with a major cloud service provider (GCP preferred); including compute, managed databases, object storage, and networking.

• Proven experience in constructing monitoring/alerting/observability stacks (Grafana, Prometheus, Zabbix, Datadog, or similar tools).

• Strong scripting and automation capabilities (Python, Bash, or equivalent).

• Proficient with containerized workloads (Docker) and CI/CD pipelines.

• Demonstrated history of decreasing incident volumes or enhancing reliability metrics — not just reacting to incidents.

• Excellent communication abilities, comfortable engaging with both technical and non-technical stakeholders, adept at recognizing when and how to escalate urgency, and capable of fostering strong working relationships across teams.

• Proficient communication skills in English — capable of writing clearly and engaging effectively in asynchronous conversations.


🏝️ Benefits

• Health insurance

• 401(k) matching

• Flexible work hours

• Paid time off

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers