
Site Reliability Engineer
Posted Jun 23

Posted Jun 23
This is a fully remote position, open to applicants in United States.
• Play a key role in the operation, maintenance, and ongoing enhancement of Accela's production cloud environments.
• Assist in platform modernization projects, focusing on containerization, cloud-native technologies, and automation initiatives.
• Oversee platform health, availability, performance, and capacity utilizing advanced observability and monitoring tools.
• Engage in incident response activities, addressing production issues and participating in Root Cause Analysis efforts.
• Create and sustain automation, tooling, and scripts that enhance reliability, scalability, deployment efficiency, and operational effectiveness.
• Aid in the implementation and tracking of service level objectives (SLOs), service level agreements (SLAs), and operational metrics.
• Collaborate with Development, DevOps, Database Engineering, and Security teams to pinpoint and resolve reliability, performance, and scalability issues.
• Support platform deployments, conduct operational readiness reviews, and engage in change management processes.
• Contribute to observability efforts through monitoring, logging, metrics collection, and distributed tracing.
• Assist with compliance-related operational tasks in accordance with SOC 2, HIPAA, FedRAMP, StateRAMP, and PCI-DSS standards.
• Take part in post-incident evaluations and contribute to corrective and preventive measures to enhance platform stability.
• A minimum of 4 years of experience in Site Reliability Engineering, Cloud Operations, Systems Engineering, DevOps, Software Engineering, or a similar technical field.
• Proven experience in supporting cloud-based SaaS environments, ideally within Microsoft Azure.
• Proficiency with Kubernetes and containerized application ecosystems.
• Competence in scripting and automation using Python, PowerShell, Bash, or comparable languages.
• Experience in troubleshooting distributed systems across application, infrastructure, networking, and operating system layers.
• Knowledge of monitoring, logging, metrics, and observability platforms.
• Strong analytical and problem-solving abilities with a methodical approach to troubleshooting and Root Cause Analysis.
• Familiarity with Incident, Problem, and Change Management processes.
• Excellent written and verbal communication skills, with the capability to collaborate effectively across cross-functional teams.
• Experience with Git and GitHub workflows.
• Flexible time off
• Comprehensive medical, dental, and vision plans
• Family planning benefits
• 401(k) retirement savings plan with company match
• Health savings account with company contributions
• Flexible spending account
• Life, accident, and disability coverage
• Business travel insurance
• Employee assistance programs
• Other well-being benefits
CVS Health
Devoteam
Aspirion
Goodgame Studios
Get handpicked remote jobs straight to your inbox weekly.