
Staff Platform Engineer
Posted 19 hours ago

Posted 19 hours ago
This is a fully remote position, open to applicants in Canada.
• Develop and implement a DevOps strategy while overseeing infrastructure architecture across diverse multi-environment and multi-region cloud systems.
• Design and manage scalable Kubernetes platforms along with containerized infrastructure at scale.
• Take ownership of the infrastructure as code strategy and establish standards across various environments.
• Spearhead the implementation of DevSecOps, focusing on secrets management, compliance, auditing, IAM, and zero-trust networking.
• Enhance platform reliability, oversee performance SLAs, and optimize costs across production systems.
• Guide complex cloud migrations and drive platform modernization efforts.
• Direct the observability strategy and maintain production reliability practices.
• Oversee the design and operation of AI/ML platform infrastructure, including model serving and deployment, GPU workload orchestration, LLM gateway and observability, vector store infrastructure, and CI/CD for AI/ML systems.
• Utilize modern AI assistants like Claude and Cursor to enhance delivery quality and speed.
• Collaborate with engineering, product, and leadership teams to align platform strategy with business and delivery objectives.
• Articulate infrastructure decisions and trade-offs to both technical and non-technical stakeholders.
• Facilitate design reviews, architecture discussions, and release readiness evaluations.
• Set platform engineering standards and best practices.
• Mentor junior and mid-level engineers.
• Serve as a technical escalation point for intricate infrastructure and platform issues.
• Assess emerging tools and technologies to enhance platform reliability and improve the developer experience.
• A minimum of 7 years of professional experience in DevOps or platform engineering, with a focus on leading complex platform initiatives.
• Proficient scripting and programming abilities in languages such as Python, Go, Java, or Bash.
• In-depth expertise in at least one major cloud platform.
• Advanced skills in Kubernetes and container orchestration.
• Proficient in infrastructure-as-code techniques across various tools.
• Extensive CI/CD architecture experience at scale.
• Strong background in DevSecOps, covering secrets management, compliance, and auditing.
• Knowledge of networking, IAM, cloud security architecture, and zero-trust principles.
• Familiarity with service mesh, distributed systems, and microservices architecture.
• Significant experience in AI/ML platform infrastructure, including model serving and deployment, GPU workload orchestration, LLM gateways and observability, vector store infrastructure, and CI/CD for AI/ML systems.
• Proven leadership and technical mentoring capabilities.
• Excellent communication skills for engaging with stakeholders.
• Demonstrated daily utilization and expertise with AI-forward tools such as Claude and Cursor.
• Strong problem-solving abilities and sound judgment in navigating ambiguous technical and business challenges.
• Cloud certifications or FinOps experience is an advantage.
• Helpful: Experience with HPC cluster infrastructure across AWS, CoreWeave, GCP, and OCI.
• Helpful: Knowledge of HPC job schedulers and workload managers like Slurm or equivalent.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible work hours and remote work options.
• Opportunities for professional development and continuous learning.
• Collaborative and innovative work environment.
Grafana Labs
Grafana Labs
Grafana Labs
Get handpicked remote jobs straight to your inbox weekly.