
Staff Site Reliability Engineer
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in Connecticut.
• Design, enhance, and sustain scalable CI/CD pipelines and deployment workflows.
• Create dependable staging and development environments that adhere to production standards.
• Develop and oversee observability practices, which include monitoring, alerting, dashboards, and SLO frameworks.
• Collaborate with engineering teams to facilitate platform modernization and ensure infrastructure reliability.
• Promote Infrastructure-as-Code (IaC) standards and maintain multi-cloud operational consistency across AWS and GCP.
• Offer technical leadership regarding infrastructure architecture, reliability, and operational best practices.
• Assist in incident response, system reliability, and operational readiness initiatives.
• Leverage AI-assisted development tools to aid in infrastructure analysis and improvement efforts.
• A minimum of 10 years of experience as a Staff SRE or in a senior-level infrastructure engineering role supporting large-scale production systems.
• Extensive knowledge of AWS and GCP cloud platforms.
• Understanding of queue-based or event-driven architectures and autoscaling technologies.
• Experience in designing and enhancing CI/CD infrastructure and automated deployment processes.
• Practical experience with Kubernetes operations, focusing on scalability and workload reliability.
• Familiarity with observability tools, monitoring strategies, and practices for incident management.
• Proficient in Infrastructure-as-Code tools like Terraform or similar technologies.
• Excellent communication and collaboration abilities across engineering teams.
• Experience with AI-assisted development tools such as Claude Code, Codex, or other comparable technologies.
• Equal opportunity workplace.
• Affirmative action employer.
• Commitment to disability accommodation and inclusion.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.