
Senior Site Reliability Engineer
Posted Jul 17

Posted Jul 17
This is a fully remote position, open to applicants in United States.
• Take ownership of and continuously enhance the observability strategy for your team, which includes monitoring, alerting, dashboards, distributed tracing, log aggregation, and SLI/SLO/SLA frameworks.
• Manage the entire incident management process: detection, triage, communication, resolution, and conducting blameless post-mortems that foster lasting improvements.
• Act as a leader in Reliability Engineering for your team, providing strong technical guidance, sound judgment, and a clear voice on reliability throughout the SDLC.
• Design and sustain autonomous systems for the development, deployment, testing, and operation of all Filevine products with minimal human oversight.
• Continuously improve CI/CD pipelines, automation scripts, playbooks, and tools to minimize toil and expedite resolution time.
• Identify and proactively address gaps in system availability, performance, and security while maintaining the overall security posture.
• Mentor engineers by dedicating significant time to their growth, sharing knowledge proactively, and enhancing their capabilities.
• Document processes, architecture, procedures, and best practices, taking full responsibility for the documentation related to the technologies in your domain and actively closing gaps for fellow SREs.
• Engage in a 24/7 on-call rotation for production support and emergency response, ensuring clear communication with technical and management stakeholders at all levels.
• Develop roadmaps for the technologies and workflows you oversee, providing the team with a clear direction for ongoing improvements.
• A minimum of 8 years of hands-on technical experience in software engineering, infrastructure, or operations roles, with at least 5 years focused specifically on Site Reliability Engineering.
• A comprehensive and expert-level SRE skill set, demonstrating proficiency in monitoring/alerting, incident response, capacity planning, performance optimization, CI/CD, and reliability engineering best practices.
• Extensive hands-on experience with New Relic or a similar observability platform, with a preference for candidates who have led the adoption or migration of observability platforms at scale.
• Proven experience in managing incident management programs, including on-call processes, escalation design, post-mortem culture, and measurable improvements in MTTR/MTTD.
• Strong skills in Python, Bash, PowerShell, and other common SRE scripting and automation technologies.
• Expert-level ability to design, build, and maintain autonomous systems for software build, deployment, testing, monitoring, and operations.
• Proficient hands-on experience with AWS (EC2, EKS/Kubernetes, CloudWatch, Lambda, S3, IAM) and the broader cloud-native ecosystem.
• A strong communicator who proactively updates stakeholders, operates transparently, and can demystify technical complexities for product and management audiences.
• A proven record of mentoring engineers, driving initiatives to completion, and significantly enhancing the capabilities of those around them.
• A Bachelor's degree in Computer Science, Information Systems, or a related field; equivalent certifications (e.g., AWS certifications, Google Cloud Professional); or substantial comparable direct work experience.
• Medical, Dental, & Vision Insurance (for full-time employees)
• Competitive & Fair Pay
• Maternity & Paternity Leave (for full-time employees)
• Short & Long-Term Disability
• Opportunity to learn from a dedicated leadership team
• Top-of-the-line company swag
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.