
Site Reliability Engineer (SRE) – UI/UX
Posted Aug 4

Posted Aug 4
This is a fully remote position, open to applicants in Canada.
• Assist in the deployment, operation, and ongoing maintenance of production services utilizing Kubernetes.
• Oversee the health, availability, and performance of services.
• Analyze and troubleshoot production incidents using logs, monitoring, and debugging tools.
• Conduct log analysis and incident debugging with Splunk.
• Identify service-related issues and work together with engineering teams to ensure timely resolutions.
• Engage in incident response and production support activities.
• Provide first-level debugging of UI-related issues involving Web Components.
• Contribute to service reliability and continuous improvement initiatives.
• Assist with CI/CD pipelines and operations of cloud-native applications as required.
• Collaborate effectively within a client-directed backlog and adhere to established priorities.
• Minimum of 4 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Support, or a similar role.
• Hands-on experience in supporting the deployment, operation, and ongoing maintenance of production services on Kubernetes.
• Proven experience in monitoring service health, troubleshooting production issues, and enhancing service reliability.
• Expertise in using Splunk for log analysis and incident debugging.
• Background in participating in production incident response and conducting root-cause analysis.
• Familiarity with Web Components and capability to perform first-level debugging of UI-related issues.
• Strong troubleshooting, analytical, and problem-solving abilities.
• Experience in collaborating with software engineering and cross-functional teams.
• Ability to work independently and effectively manage a client-directed backlog.
• Excellent written and spoken English skills, at least at a B2 level.
• Experience in supporting CI/CD pipelines.
• Knowledge of multi-tenant services.
• Experience with cloud-native application operations.
• Background in supporting high-availability enterprise or SaaS platforms.
• Familiarity with additional monitoring and observability tools.
• Experience with cloud platforms such as AWS, Azure, or GCP.
• Familiarity with container and deployment technologies such as Docker and Helm.
• Competitive salary.
• Laptop provided.
• Opportunities for professional development and training.
• Work with advanced cloud and container technologies.
• Flexible work arrangements and a collaborative team environment.
• Contribute to organization-wide digital transformation initiatives.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.