
Site Reliability Engineer – L3 Support
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in Kansas, +3 more states.
• Oversee the health, availability, performance, and security of production services.
• Proactively detect emerging issues utilizing telemetry, logs, metrics, and distributed tracing.
• Investigate, troubleshoot, and resolve intricate production incidents across both application and infrastructure layers.
• Serve as the L3 escalation point for operational issues that cannot be resolved by L1 or L2 support.
• Participate in an on-call rotation for urgent production incidents.
• Lead incident response efforts, including coordination, communication, and post-incident evaluations.
• Conduct root cause analysis and ensure corrective measures are taken to prevent future occurrences.
• Create and maintain operational runbooks, dashboards, alerts, and standard operating procedures.
• Enhance platform observability by improving monitoring, alerting, dashboards, and service-level indicators.
• Collaborate closely with software engineering teams to boost service reliability, scalability, and resilience.
• Identify chances to automate operational tasks and eliminate repetitive manual processes.
• Support production deployments, infrastructure modifications, and maintenance tasks.
• Assist with disaster recovery drills, resilience testing, and operational readiness assessments.
• Ensure that operational activities adhere to FedRAMP High security and compliance standards.
• Contribute to continuous improvement efforts in reliability, performance, and operational excellence.
• U.S. Citizenship (required)
• 3–6 years of experience in Site Reliability Engineering, Production Engineering, DevOps, Platform Engineering, or a senior production support position
• Experience with mission-critical cloud-based production systems
• Strong knowledge of Linux operating systems and networking principles
• Experience troubleshooting distributed applications operating in Kubernetes
• Familiarity with public cloud platforms, preferably AWS
• Experience with infrastructure as code and configuration management
• Proficient scripting or programming skills (e.g., Python, Bash, PowerShell, Go, or similar)
• Experience using monitoring and observability tools such as Prometheus, Grafana, CloudWatch, Datadog, Splunk, or OpenTelemetry
• Proficient in analyzing application logs, metrics, and traces to diagnose production issues
• Understanding of incident management, problem management, and root cause analysis
• Strong analytical and troubleshooting capabilities
• Excellent written and verbal communication skills.
• Hybrid Work Model & a Business Casual Dress Code, including jeans
• 401k Matching Program
• Professional Development Reimbursement
• Flexible Personal/Vacation Time Off
• Sick Leave
• Paid Holidays
• Medical, Dental, Vision
• Employee Assistance Program
• Parental Leave
• Discounts on fitness clubs, travel, and more!
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.