Site Reliability Engineer

Posted Sep 1

This is a fully remote position, open to applicants in Massachusetts.

📋 Description

• Diagnose intricate issues within Linux systems, networking, and distributed services.

• Develop software and automation solutions to minimize operational challenges, enhance efficiency, and avert recurring problems.

• Create and utilize AI-enhanced tools for incident investigation, identifying operational patterns, reducing toil, and improving reliability.

• Leverage data analysis, network diagnostics, and debugging tools to pinpoint enhancements in performance and reliability.

• Establish and refine monitoring, alerting, SLIs, and SLOs for essential services.

• Contribute to root cause investigations, post-incident evaluations, and long-term corrective strategies.

• Collaborate with Engineering teams to enhance system design, ensure deployment safety, and bolster operational readiness.

• Engage in an on-call rotation and lead efforts during incident responses.

• Propel timely service restoration, facilitate effective communication, and drive post-incident improvement initiatives.


⛳️ Requirements

• Relevant experience along with a Bachelor's degree in Computer Engineering, Computer Science, or a related field.

• Experience in supporting large-scale distributed systems.

• Knowledge of Linux and networking, including routing, DNS, firewalls, TCP/IP, and L7 traffic management.

• Proficient in a programming language such as Python or Go.

• Familiarity with observability tools like Prometheus, Grafana, Loki, ELK/OpenSearch, or similar platforms.

• Experience with infrastructure automation or configuration management tools such as Terraform, Ansible, Salt, or comparable solutions.

• Understanding of container technologies like Docker or Podman.

• Familiarity with orchestration platforms such as Kubernetes or Nomad.


🏝️ Benefits

• Annual bonus or incentives.

• Equity awards.

• Employee Stock Purchase Plan (ESPP).

• Healthcare benefits.

• 401K savings plan.

• Company holidays.

• Vacation in the form of PTO.

• Sick time.

• Parental leave.

• Employee assistance program.

• Mental wellness support.

• Financial wellness support.

• Flexible work arrangements: at home, in an office, or a combination of both.

People also viewed

In All Media18 hours ago

DevOps Engineer – Cloud

BR flagBrazil, +5 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group18 hours ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Fingerprint21 hours ago

Senior Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$152k – $205k/year
ApplyView job
Endava23 hours ago

Senior DevOps Engineer, Terraform

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
CVS Health1 day ago

Staff DevSecOps Engineer, Health

US flagConnecticut, +3 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$130.3k – $260.6k/year
ApplyView job
GoFasti1 day ago

Senior DevOps Engineer

Latin AmericaFull-timeDevOps & Site Reliability Engineer (SRE)$5,000 – $6,000/month
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers