
Staff Site Reliability Engineer
Posted Aug 5

Posted Aug 5
This is a fully remote position, open to applicants in Canada, +1 more state.
• Lead the advancement of the Password Safe platform, its infrastructure, and the deployment ecosystem.
• Design highly available and resilient systems that operate across both cloud and on-premises environments.
• Take ownership of the reliability, scalability, and efficiency of shared services, CI/CD pipelines, and platform engineering projects.
• Design, scale, and maintain secure, resilient systems on AWS/Azure and in on-premises settings.
• Advocate for platform engineering initiatives that minimize cognitive load and enhance software engineering speed.
• Manage and optimize API gateways, service meshes, caches, configuration management, and secrets management.
• Promote an Everything as Code culture by utilizing declarative, version-controlled infrastructure, pipelines, and configurations deployed via GitOps workflows.
• Standardize, secure, and enhance CI/CD pipelines for safe, repeatable, and rapid deployments.
• Design and implement chaos engineering frameworks along with disaster recovery simulations.
• Architect and enhance telemetry using metrics, logs, traces, Grafana Cloud, Datadog, and OpenTelemetry.
• Define and enforce SLOs and SLIs for critical applications.
• Collaborate with engineering leadership to shape the long-term SRE strategy.
• Mentor and coach engineers, lead design and incident reviews, and create reliability documentation and tools.
• Over 7 years of experience in SRE, DevOps, or Platform Engineering, including a minimum of 2 years in a Senior or Staff role.
• Demonstrated experience in managing automation and infrastructure across both cloud and on-premises environments.
• Proficiency with Docker and Kubernetes, including cluster management, networking, and security features.
• Previous experience with canary or blue-green release orchestration strategies.
• Skilled in at least one systems programming language such as Go, Java, or C#.
• In-depth knowledge of OpenTelemetry and best practices for monitoring and observability.
• Familiarity with GitOps workflows along with Infrastructure as Code and Configuration as Code best practices.
• Nice to have: Knowledge in UI automation testing.
• Nice to have: Experience in designing and testing microservice-based applications.
• Nice to have: Experience with virtual machines and testing environments.
• Nice to have: Background in continuous integration environments.
• Nice to have: Strong understanding of Linux and Windows internals.
• Nice to have: Experience in migrating workloads between on-premises and cloud environments.
• Nice to have: Understanding of modern DevSecOps practices.
• Flexible work culture.
• A focus on trust and continuous learning.
• Recognition for growth and impact.
• Commitment to a diversity and inclusion-oriented culture.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.