
Staff DevOps Engineer
Posted Aug 27

Posted Aug 27
This is a fully remote position, open to applicants in Canada.
• Take ownership of and enhance Nexxa's fundamental infrastructure, which encompasses compute, networking, storage, and deployment systems from start to finish.
• Create and manage CI/CD pipelines that facilitate safe iterations across AI, data, and product engineering teams.
• Develop and sustain infrastructure-as-code for reproducible and auditable cloud and on-premises/edge environments.
• Design and oversee Kubernetes platforms for training, inference, and application workloads, incorporating GPU scheduling and autoscaling capabilities.
• Provide support for data warehouses, lakehouse architectures, feature stores, embedding indices, retrieval pipelines, as well as model training, evaluation, and serving infrastructure.
• Establish and promote observability practices that include metrics, logging, tracing, and alerting across distributed systems.
• Implement reliability practices such as SLOs/SLIs, incident response, postmortems, and on-call rotations.
• Design security measures and compliance protocols across cloud infrastructure, secrets management, and access control.
• Make informed trade-offs between cost, latency, reliability, and developer velocity.
• Collaborate with engineering leadership to shape the infrastructure roadmap and platform strategy.
• Mentor engineers on best practices in infrastructure and advocate for operational excellence.
• 6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or other infrastructure-centric software engineering roles.
• Extensive hands-on experience with AWS, GCP, or Azure at a production scale.
• Proven production experience with Kubernetes, including GPU workload scheduling.
• Familiarity with infrastructure-as-code tools such as Terraform, Pulumi, or their equivalents.
• Experience with CI/CD systems, including GitHub Actions, GitLab CI, CircleCI, Jenkins, or ArgoCD.
• Knowledge in designing and managing observability stacks like Prometheus, Grafana, Datadog, or OpenTelemetry.
• Experience in supporting ML/AI infrastructure is highly advantageous.
• Excellent scripting or programming skills in Python, Go, or Bash.
• Demonstrated ability to independently define and lead infrastructure projects from design to production rollout.
• Strong instincts for incident management and the capability to lead during outages and conduct root-cause analysis.
• Experience with cloud and edge/on-premises infrastructure, particularly in industrial or manufacturing environments.
• Familiarity with Snowflake, BigQuery, Redshift, or Databricks.
• Knowledge of service mesh, zero-trust networking, or compliance frameworks for industrial/critical infrastructure such as SOC 2 or IEC 62443.
• A track record of building internal developer platforms or self-service infrastructure tools.
• Experience in scaling infrastructure teams or establishing technical direction at the Staff level.
• Significant opportunities for career growth and advancement.
• Comprehensive salary and equity compensation package.
• Innovative environment centered on transforming heavy industries through AI and automation.
• Collaborative culture that values innovation, discipline, and continuous improvement.
Chalice AI
DaVita Kidney Care
WEX
WEX
Get handpicked remote jobs straight to your inbox weekly.