
Site Reliability Engineer
Posted Jul 31

Posted Jul 31
This is a fully remote position, open to applicants in United States, +7 more states.
• Take ownership of reliability objectives across our backend/API, worker services, applications, and CV pipeline; MTD, MTM, MTR, and ensure thorough investigation of root causes.
• Enhance our monitoring and alerting systems, creating automated remediation processes, so that on-call responsibilities are managed through automation rather than increasing headcount.
• Collaborate with our engineering teams to develop agents that can triage alerts and manage routine remediation tasks.
• Strengthen and optimize our GCP infrastructure (Cloud Run, Cloud SQL, GCS) for both cost-effectiveness and performance as load increases.
• Manage database scaling and performance, including connection pooling, query optimization, indexing, read replicas, and capacity planning, to prevent Postgres from becoming a bottleneck as data volume expands.
• Enhance the reliability of our ML training and monitoring infrastructure in collaboration with the CV/ML team.
• Conduct blameless postmortems and focus on implementing solutions for root causes rather than merely addressing symptoms.
• Participate in the on-call rotation.
• A minimum of 4 years of experience in an SRE, infrastructure, or backend engineering position with production on-call responsibilities.
• Extensive experience with a major cloud service provider (GCP preferred); including compute, managed databases, object storage, and networking.
• Proven experience in constructing monitoring/alerting/observability stacks (Grafana, Prometheus, Zabbix, Datadog, or similar tools).
• Strong scripting and automation capabilities (Python, Bash, or equivalent).
• Proficient with containerized workloads (Docker) and CI/CD pipelines.
• Demonstrated history of decreasing incident volumes or enhancing reliability metrics — not just reacting to incidents.
• Excellent communication abilities, comfortable engaging with both technical and non-technical stakeholders, adept at recognizing when and how to escalate urgency, and capable of fostering strong working relationships across teams.
• Proficient communication skills in English — capable of writing clearly and engaging effectively in asynchronous conversations.
• Health insurance
• 401(k) matching
• Flexible work hours
• Paid time off
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.