
Director of SRE
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Canada.
• Oversee the design, implementation, and management of a scalable, dependable, and highly available cloud-based infrastructure (AWS/Azure).
• Establish Site Reliability Engineering (SRE) best practices, encompassing monitoring, incident response, capacity planning, and performance optimization.
• Enhance observability, monitoring, and alerting systems to ensure prompt detection and resolution of reliability challenges.
• Advocate for automation-first methodologies, minimizing manual processes through Infrastructure-as-Code (IaC) and CI/CD pipelines.
• Guide a team of SREs, embodying Blackpoint Cyber's management principles of Coach, Model, Care, to define critical business outcomes, formulate action plans, and assist the team in achieving their goals.
• Maintain active contributions in an SRE capacity.
• Design, implement, and maintain essential infrastructure, including automated deployment for attack infrastructure, secure identity and productivity environments, and safe data storage solutions.
• Develop and implement security hygiene and monitoring policies that comply with Blackpoint Cyber's security standards.
• Monitor and enhance cloud expenditures, ensuring that resource utilization remains cost-effective without compromising reliability.
• Manage and guide a worldwide team of SREs, DevOps engineers, and cloud infrastructure experts.
• Collaborate with security teams to ensure adherence to compliance, bolster security measures, and prepare for disaster recovery.
• Over 10 years of experience in SRE, DevOps, or Cloud Infrastructure roles.
• More than 5 years of experience in personnel management, leading an SRE team.
• Extensive knowledge of AWS, Azure, or GCP, with a focus on cost management and scaling strategies.
• Proficient in Infrastructure-as-Code (IaC) (e.g., Terraform, CloudFormation, Pulumi).
• Practical experience with CI/CD pipelines, Kubernetes, and container orchestration.
• Expertise in monitoring, logging, and observability tools (e.g., Prometheus, Grafana, Datadog, Splunk).
• Demonstrated ability to optimize cloud costs (COGS) while upholding reliability and performance standards.
• Strong leadership, teamwork, and problem-solving abilities.
• Familiarity with SLA/SLO/SLIs will be advantageous.
• A general understanding of the current AI tooling landscape and its application in increasing velocity and enhancing stability within SRE.
• Health insurance
• Vision insurance
• Dental insurance
• Life insurance
• 401k plan
• Discretionary Time Off
Cribl
Flock Safety
Pear Tree.
GFT Technologies
Get handpicked remote jobs straight to your inbox weekly.