
Senior Site Reliability Engineer, Production Support
Posted Aug 1

Posted Aug 1
This is a fully remote position, open to applicants in Canada.
• Assist in the deployment, operation, and ongoing maintenance of a production service on Kubernetes.
• Oversee application health, availability, and performance metrics.
• Analyze and resolve production incidents utilizing logs, monitoring tools, and debugging techniques.
• Conduct log analysis with Splunk to determine root causes and address service-related issues.
• Work collaboratively with software engineers to enhance service reliability and operational effectiveness.
• Engage in incident response and production support tasks.
• Aid in initial debugging of UI-related issues related to Web Components.
• Contribute to the ongoing enhancement of automation, monitoring, and operational workflows.
• Assist in the support of CI/CD pipelines and cloud-native deployment methodologies.
• Over 5 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Production Operations.
• Profound hands-on experience with Kubernetes in production settings.
• Experience in supporting cloud-native applications.
• Background in monitoring production systems and resolving complex incidents.
• In-depth knowledge of Splunk for log analysis and debugging purposes.
• Proficient in working within Linux environments.
• Understanding of networking principles and distributed systems.
• Strong skills in troubleshooting and root cause analysis.
• Exceptional written and spoken English proficiency (B2+).
• Competitive salary and laptop provided.
• Opportunities for professional development and training.
• Work with innovative cloud and container technologies.
• Flexible work arrangements and a collaborative team atmosphere.
• Contribute to organization-wide digital transformation initiatives.
Thermo Fisher Scientific
ASCOR
Get handpicked remote jobs straight to your inbox weekly.