
Senior SRE, Managed Gateways
Posted Jul 23

Posted Jul 23
This is a fully remote position, open to applicants in Canada.
• Lead, mentor, and motivate a high-performing team of Site Reliability Engineers focused on Kong's Managed Gateway solutions.
• Design and implement resilient, scalable, and fault-tolerant cloud-native systems utilizing technologies such as Kubernetes, Golang, and prominent cloud service providers.
• Take ownership of the complete operational lifecycle, from proactive monitoring and alerting to incident management and blameless post-mortem analyses, ensuring ongoing service enhancement.
• Foster a culture of developer satisfaction by implementing automation, self-service tools, and efficient workflows for deploying and managing API gateways.
• Establish, monitor, and report on key SLOs and SLIs to guarantee the optimal performance and reliability of Managed Gateways.
• Advocate for the prevention of technical debt and promote architectural best practices that improve system resilience and minimize operational overhead.
• Collaborate across departments with Product, Engineering, and Customer Success to influence roadmap strategies and ensure operational readiness for new features.
• Work directly with enterprise clients—partnering with Product leadership, Professional Services, and Customer Success—to facilitate the end-to-end onboarding and implementation of Cloud Gateways, and to develop repeatable playbooks and platform capabilities from recurring implementation patterns.
• Leverage extensive cross-cloud expertise (AWS, GCP, Azure) to address unique customer configurations and transform complex setups into successful, production-ready deployments.
• Significant experience as a Site Reliability Engineer, particularly with highly available and distributed systems.
• Profound knowledge of Kubernetes and cloud-native architectures, ideally across various public cloud providers (AWS, GCP, Azure).
• Strong skills in Golang or other modern programming languages for automation and tool creation.
• Demonstrated experience in constructing and maintaining CI/CD pipelines and infrastructure as code (Terraform, Ansible).
• Comprehensive understanding of monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, ELK stack, Datadog).
• Experience with managed services, API gateways, or related network infrastructure is highly advantageous.
• Health insurance
• 401(k) plan
• Short and long-term disability benefits
• Basic life and AD&D insurance
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.