
Senior Site Reliability Engineer
Posted Jul 2

Posted Jul 2
This is a fully remote position, open to applicants in Connecticut, +6 more states.
• Architect, construct, and maintain the foundational shared platforms that engineers utilize daily: GCP infrastructure, Kubernetes, networking, routing, CI/CD, and observability.
• Identify and resolve issues within complex distributed systems operating under high request volumes.
• Ensure system observability and evaluate the performance of our technology stack.
• Engage in ongoing projects such as modernizing our edge, caching, and gateway layers using Fastly, while enhancing observability across the platform.
• Elevate reliability standards through improved dashboards, alert severity protocols, paging guidelines, on-call preparedness, and incident management.
• Aim to make deployments seamless: establish golden paths, conduct production readiness assessments, ensure safe rollouts, and create effective automation to streamline the shipping process for engineers.
• Mentor fellow engineers and elevate the technical standards through code reviews, design assessments, and collaborative work.
• Contribute to our on-call rotation and assist in the successful implementation of our developer on-call program.
• Located in the United States, with a reasonable overlap with European engineering hours.
• Proficiency with SRE/DevOps tools, methodologies, and culture.
• Over 5 years of experience participating in an SRE on-call rotation.
• Analytical mindset for designing, diagnosing, and optimizing infrastructure.
• Experience managing scalable, highly available, cloud-based applications, particularly with high request volumes and customer-facing uptime obligations.
• Familiarity with Kubernetes for orchestrating, scaling, and overseeing containerized applications in cloud environments.
• Experience in developing CI/CD pipelines.
• Knowledge of an observability stack (such as Prometheus, etc.).
• Comfortable working across CDNs, edge networks, gateways, and caching layers, or enthusiastic about deepening knowledge in these areas.
• You enhance on-call processes and reliability by creating systems, standards, and feedback mechanisms that progressively improve production health.
• Adept at managing incidents and outages, possessing a practical and thoughtful communication style for high-pressure situations.
• Open-minded yet thoughtful approach towards new technologies.
• A team of highly skilled, inspiring, and supportive professionals.
• Opportunities to work with real infrastructure scale and engage in meaningful, hands-on tasks that transform operations.
• A positive, flexible, and trust-based work environment that fosters long-term personal and professional growth.
• A diverse, global group of colleagues and customers from various cultures.
• Comprehensive health plans and additional perks.
• A healthy work-life balance that accommodates individual and family needs.
• Competitive stock options program along with location-based salary.
Ontrac Solutions
CyberSheath
Ontrac Solutions
NVIDIA
Get handpicked remote jobs straight to your inbox weekly.