
Senior SRE Engineer – CCaaS
Posted Sep 15

Posted Sep 15
This is a fully remote position, open to applicants in Canada.
• Ensure the CCaaS platform and its supporting services maintain 24x7 availability, reliability, performance, security, and operational excellence.
• Define and oversee Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
• Monitor the health, performance, capacity, and stability of the platform.
• Create and sustain monitoring dashboards, alerts, and health checks.
• Configure and optimize observability tools such as CloudWatch, Splunk, Dynatrace, Grafana, and others.
• Develop operational runbooks and troubleshooting documentation.
• Serve as the escalation point for critical incidents and service disruptions.
• Lead incident triage, impact assessments, communication efforts, and recovery activities.
• Conduct root cause analyses and identify preventive measures.
• Drive initiatives to reduce Mean Time to Recovery (MTTR), Mean Time to Detection (MTTD), and recurring incidents.
• Minimize operational toil through automation and create self-healing and automated remediation solutions.
• Build scripts and tools leveraging Python, PowerShell, AWS Lambda, and other automation frameworks.
• Automate operational checks, deployment validations, monitoring, and reporting tasks.
• Provide support for Amazon Connect, AWS services, APIs, middleware, routing, IVR, and integration components.
• Manage platform configurations, certificates, service accounts, and connectivity requirements.
• Assist in disaster recovery, backup, restoration, and failover testing procedures.
• Conduct capacity planning and performance optimization efforts.
• Participate in release planning, deployment validation, and production implementations.
• Review changes to assess operational risks and reliability impacts.
• Support production deployments and rollback strategies.
• Ensure operational readiness prior to launch.
• Assist with vulnerability remediation, audits, security events, and compliance requirements.
• Analyze operational metrics to identify areas for improvement.
• Promote Site Reliability Engineering (SRE) best practices across CCaaS teams.
• Engage in architecture reviews focused on scalability and resilience.
• Mentor developers and operations teams on SRE methodologies.
• Collaborate with Development, DevOps, Infrastructure, Security, Business, and Vendor teams.
• Debug production issues across various services and technology stack levels.
• Enhance service health visibility through metrics, logs, and traces.
• Assess the cost and business impact of SLA breaches and downtime.
• Generally, 5–6 years of relevant experience is required.
• A post-secondary degree in a related field or an equivalent combination of education and experience is necessary.
• Extensive education and business experience have provided deep knowledge and technical proficiency.
• Foundational understanding of DevOps principles is expected.
• Foundational knowledge of cybersecurity and privacy concepts, principles, and solutions is required.
• Intermediate proficiency in IT Infrastructure Library (ITIL) is necessary.
• Intermediate proficiency in Robotic Process Automation (RPA) is required.
• Intermediate cloud computing skills are necessary.
• Intermediate knowledge of configuration management is expected.
• Intermediate proficiency in container orchestration is required.
• Intermediate skills in system design and implementation are necessary.
• Intermediate proficiency in incident management is required.
• Advanced skills in alerting and log configuration with OpenSearch and CloudWatch are essential.
• Advanced proficiency in Dynatrace or similar Application Performance Management (APM) tools, including configuration and dashboarding, is required.
• Advanced automation skills and experience with automation pipelines are necessary.
• Advanced proficiency in automated testing is expected.
• Strong verbal and written communication, collaboration, analytical and problem-solving skills, along with data-driven decision-making abilities, are essential.
• Performance-based incentives are available.
• Discretionary bonuses are offered.
• Health insurance coverage is provided.
• Tuition reimbursement is available.
• Accident and life insurance are included.
• Retirement savings plans are offered.
• Comprehensive training and coaching opportunities are provided.
• Managerial support is available.
• Opportunities for network-building are encouraged.
• Accommodations are available upon request during the selection process.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.