
Site Reliability Engineer
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in Canada.
• Serve as the initial point of contact for production alerts and incidents across various services, managing the process from triage to resolution.
• Identify and resolve issues directly in AWS and Kubernetes, addressing problems such as failing pods, resource exhaustion, deployment errors, networking/DNS, database, and caching issues.
• Execute rollbacks, scale, reconfigure, or apply patches to infrastructure to promptly restore services.
• Escalate issues to development teams only when a code modification is required, ensuring a detailed diagnosis is provided.
• Oversee the PagerDuty configuration and incident response during business hours while striving to minimize MTTD and MTTR.
• Conduct blameless post-mortems and lead technical follow-up initiatives.
• Automate runbooks and repetitive operational tasks, utilizing AI-assisted triage, investigation, and remediation.
• Develop and maintain monitoring solutions for Kubernetes workloads and services using Prometheus/Mimir, Loki, Tempo, Grafana, and OpenTelemetry.
• Monitor production releases and detect regressions in latency, errors, or resource consumption.
• Design and manage synthetic checks, smoke tests, health checks, and load/performance tests.
• Collaborate with engineering teams to address performance and reliability challenges, defining SLOs, SLIs, and error budgets.
• Strengthen the platform through Terraform modifications, Kubernetes resource optimization, autoscaling, CI/CD checks, secrets management, and IAM enhancements.
• Maintain comprehensive service documentation and document architectural decisions.
• Minimum of 3 years in a cloud engineering, DevOps, or SRE position providing support for production web applications.
• Practical experience in responding to production incidents, including diagnosing and resolving issues directly.
• Strong hands-on experience with AWS in production environments, including EKS, EC2, RDS, VPC networking, IAM, and CloudWatch.
• Extensive experience with Kubernetes in production, encompassing troubleshooting, debugging, and workload monitoring.
• Proven track record of identifying and resolving performance and reliability issues in collaboration with engineering teams.
• Familiarity with monitoring and observability tools such as Prometheus, Grafana, Loki, Datadog, or CloudWatch.
• Experience with on-call and alerting tools like PagerDuty.
• Capability to develop automated tests or checks for production reliability, including synthetic, smoke, health, or load tests.
• Willingness to partake in a potential future on-call rotation outside of regular hours.
• Proficiency in infrastructure as code using Terraform and experience working within CI/CD pipelines.
• Strong grounding in Linux, networking, and container fundamentals.
• Competency in scripting and automation using Bash, Python, or similar languages.
• Ability to communicate calmly and clearly during incidents and across teams.
• Preferred: Familiarity with AI tools for SRE tasks; experience with Azure and possibly GCP; knowledge of multiple technology stacks; AWS certification; proficiency with LGTM or OpenTelemetry at scale; experience with k6, Locust, or JMeter; MySQL/PostgreSQL management; Redis/Memcached optimization; Cloudflare; and expertise in cloud cost management and capacity planning.
• Remote work opportunity.
• Preferential consideration may be offered to candidates located within a reasonable commuting distance to one of the offices.
• Equal opportunity employer.
• Unique accommodations available during the interview process.
• Potential criminal background check during the final interview phase.
Koniag Government Services
FP Markets (First Prudential Markets)
Modern Campus
InRule
Get handpicked remote jobs straight to your inbox weekly.