
Site Reliability Engineer
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in California.
• Establish and execute SLIs/SLOs for essential services.
• Direct incident response and oversee cross-team mitigation initiatives.
• Carry out blameless postmortems and ensure all corrective measures are implemented.
• Conduct production readiness assessments for new services and features.
• Recognize systemic risks and promote preventive enhancements.
• Design and enhance monitoring systems, alerting mechanisms, and dashboards (Prometheus, Grafana, etc.).
• Develop internal tools for tracking and reporting reliability.
• Automate recurring operational tasks.
• Enhance CI/CD reliability along with release processes.
• Collaborate with engineering teams to bolster system resilience.
• Offer expertise on fault tolerance, scalability, and failure management.
• Engage in architectural discussions with a focus on reliability.
• Over 5 years of experience in SRE, Reliability Engineering, or Production Engineering.
• Proficient in Linux systems and networking.
• Experience overseeing containerized production environments.
• Strong knowledge of distributed systems and their failure modes.
• Proven ability to define and manage SLIs/SLOs.
• Demonstrated experience in incident response and leading postmortems.
• Strong skills in scripting or programming.
• Familiarity with monitoring and alerting systems.
• Outstanding written communication abilities.
• Successful completion of a background check.
• Preferred: Experience with GPU infrastructure or AI/ML platforms.
• Experience enhancing reliability in high-growth or large-scale settings.
• Knowledge of GPU observability tools.
• Experience with Infrastructure as Code.
• Background in startup environments.
• Experience creating internal reliability platforms or frameworks.
• Meaningful equity in a rapidly growing company—everyone on the team receives stock options, allowing you to share in the company's success as your contributions drive growth.
• Comprehensive medical, dental, and vision plans.
• Flexible PTO—take the time you need to rejuvenate.
• Most positions are remote-first, with an inclusive and collaborative team using Slack as the primary communication tool.
• Become part of a passionate team at the forefront of AI infrastructure—where culture, learning, and ownership are central to our scaling efforts.
TEKsystems
TEKsystems
Level Data
Level Data
Get handpicked remote jobs straight to your inbox weekly.