Senior SRE Engineer – CCaaS

Posted Sep 15

This is a fully remote position, open to applicants in Canada.

📋 Description

• Ensure the CCaaS platform and its supporting services maintain 24x7 availability, reliability, performance, security, and operational excellence.

• Define and oversee Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.

• Monitor the health, performance, capacity, and stability of the platform.

• Create and sustain monitoring dashboards, alerts, and health checks.

• Configure and optimize observability tools such as CloudWatch, Splunk, Dynatrace, Grafana, and others.

• Develop operational runbooks and troubleshooting documentation.

• Serve as the escalation point for critical incidents and service disruptions.

• Lead incident triage, impact assessments, communication efforts, and recovery activities.

• Conduct root cause analyses and identify preventive measures.

• Drive initiatives to reduce Mean Time to Recovery (MTTR), Mean Time to Detection (MTTD), and recurring incidents.

• Minimize operational toil through automation and create self-healing and automated remediation solutions.

• Build scripts and tools leveraging Python, PowerShell, AWS Lambda, and other automation frameworks.

• Automate operational checks, deployment validations, monitoring, and reporting tasks.

• Provide support for Amazon Connect, AWS services, APIs, middleware, routing, IVR, and integration components.

• Manage platform configurations, certificates, service accounts, and connectivity requirements.

• Assist in disaster recovery, backup, restoration, and failover testing procedures.

• Conduct capacity planning and performance optimization efforts.

• Participate in release planning, deployment validation, and production implementations.

• Review changes to assess operational risks and reliability impacts.

• Support production deployments and rollback strategies.

• Ensure operational readiness prior to launch.

• Assist with vulnerability remediation, audits, security events, and compliance requirements.

• Analyze operational metrics to identify areas for improvement.

• Promote Site Reliability Engineering (SRE) best practices across CCaaS teams.

• Engage in architecture reviews focused on scalability and resilience.

• Mentor developers and operations teams on SRE methodologies.

• Collaborate with Development, DevOps, Infrastructure, Security, Business, and Vendor teams.

• Debug production issues across various services and technology stack levels.

• Enhance service health visibility through metrics, logs, and traces.

• Assess the cost and business impact of SLA breaches and downtime.


⛳️ Requirements

• Generally, 5–6 years of relevant experience is required.

• A post-secondary degree in a related field or an equivalent combination of education and experience is necessary.

• Extensive education and business experience have provided deep knowledge and technical proficiency.

• Foundational understanding of DevOps principles is expected.

• Foundational knowledge of cybersecurity and privacy concepts, principles, and solutions is required.

• Intermediate proficiency in IT Infrastructure Library (ITIL) is necessary.

• Intermediate proficiency in Robotic Process Automation (RPA) is required.

• Intermediate cloud computing skills are necessary.

• Intermediate knowledge of configuration management is expected.

• Intermediate proficiency in container orchestration is required.

• Intermediate skills in system design and implementation are necessary.

• Intermediate proficiency in incident management is required.

• Advanced skills in alerting and log configuration with OpenSearch and CloudWatch are essential.

• Advanced proficiency in Dynatrace or similar Application Performance Management (APM) tools, including configuration and dashboarding, is required.

• Advanced automation skills and experience with automation pipelines are necessary.

• Advanced proficiency in automated testing is expected.

• Strong verbal and written communication, collaboration, analytical and problem-solving skills, along with data-driven decision-making abilities, are essential.


🏝️ Benefits

• Performance-based incentives are available.

• Discretionary bonuses are offered.

• Health insurance coverage is provided.

• Tuition reimbursement is available.

• Accident and life insurance are included.

• Retirement savings plans are offered.

• Comprehensive training and coaching opportunities are provided.

• Managerial support is available.

• Opportunities for network-building are encouraged.

• Accommodations are available upon request during the selection process.

People also viewed

Horizon3.ai14 hours ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH14 hours ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM14 hours ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies15 hours ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)15 hours ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC15 hours ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers