Remotery

Site Reliability Engineer – L3 Support

Posted Jul 28

This is a fully remote position, open to applicants in Kansas, +3 more states.

📋 Description

• Oversee the health, availability, performance, and security of production services.

• Proactively detect emerging issues utilizing telemetry, logs, metrics, and distributed tracing.

• Investigate, troubleshoot, and resolve intricate production incidents across both application and infrastructure layers.

• Serve as the L3 escalation point for operational issues that cannot be resolved by L1 or L2 support.

• Participate in an on-call rotation for urgent production incidents.

• Lead incident response efforts, including coordination, communication, and post-incident evaluations.

• Conduct root cause analysis and ensure corrective measures are taken to prevent future occurrences.

• Create and maintain operational runbooks, dashboards, alerts, and standard operating procedures.

• Enhance platform observability by improving monitoring, alerting, dashboards, and service-level indicators.

• Collaborate closely with software engineering teams to boost service reliability, scalability, and resilience.

• Identify chances to automate operational tasks and eliminate repetitive manual processes.

• Support production deployments, infrastructure modifications, and maintenance tasks.

• Assist with disaster recovery drills, resilience testing, and operational readiness assessments.

• Ensure that operational activities adhere to FedRAMP High security and compliance standards.

• Contribute to continuous improvement efforts in reliability, performance, and operational excellence.


⛳️ Requirements

• U.S. Citizenship (required)

• 3–6 years of experience in Site Reliability Engineering, Production Engineering, DevOps, Platform Engineering, or a senior production support position

• Experience with mission-critical cloud-based production systems

• Strong knowledge of Linux operating systems and networking principles

• Experience troubleshooting distributed applications operating in Kubernetes

• Familiarity with public cloud platforms, preferably AWS

• Experience with infrastructure as code and configuration management

• Proficient scripting or programming skills (e.g., Python, Bash, PowerShell, Go, or similar)

• Experience using monitoring and observability tools such as Prometheus, Grafana, CloudWatch, Datadog, Splunk, or OpenTelemetry

• Proficient in analyzing application logs, metrics, and traces to diagnose production issues

• Understanding of incident management, problem management, and root cause analysis

• Strong analytical and troubleshooting capabilities

• Excellent written and verbal communication skills.


🏝️ Benefits

• Hybrid Work Model & a Business Casual Dress Code, including jeans

• 401k Matching Program

• Professional Development Reimbursement

• Flexible Personal/Vacation Time Off

• Sick Leave

• Paid Holidays

• Medical, Dental, Vision

• Employee Assistance Program

• Parental Leave

• Discounts on fitness clubs, travel, and more!

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers