Staff DevOps Engineer

Posted Aug 27

This is a fully remote position, open to applicants in Canada.

📋 Description

• Take ownership of and enhance Nexxa's fundamental infrastructure, which encompasses compute, networking, storage, and deployment systems from start to finish.

• Create and manage CI/CD pipelines that facilitate safe iterations across AI, data, and product engineering teams.

• Develop and sustain infrastructure-as-code for reproducible and auditable cloud and on-premises/edge environments.

• Design and oversee Kubernetes platforms for training, inference, and application workloads, incorporating GPU scheduling and autoscaling capabilities.

• Provide support for data warehouses, lakehouse architectures, feature stores, embedding indices, retrieval pipelines, as well as model training, evaluation, and serving infrastructure.

• Establish and promote observability practices that include metrics, logging, tracing, and alerting across distributed systems.

• Implement reliability practices such as SLOs/SLIs, incident response, postmortems, and on-call rotations.

• Design security measures and compliance protocols across cloud infrastructure, secrets management, and access control.

• Make informed trade-offs between cost, latency, reliability, and developer velocity.

• Collaborate with engineering leadership to shape the infrastructure roadmap and platform strategy.

• Mentor engineers on best practices in infrastructure and advocate for operational excellence.


⛳️ Requirements

• 6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or other infrastructure-centric software engineering roles.

• Extensive hands-on experience with AWS, GCP, or Azure at a production scale.

• Proven production experience with Kubernetes, including GPU workload scheduling.

• Familiarity with infrastructure-as-code tools such as Terraform, Pulumi, or their equivalents.

• Experience with CI/CD systems, including GitHub Actions, GitLab CI, CircleCI, Jenkins, or ArgoCD.

• Knowledge in designing and managing observability stacks like Prometheus, Grafana, Datadog, or OpenTelemetry.

• Experience in supporting ML/AI infrastructure is highly advantageous.

• Excellent scripting or programming skills in Python, Go, or Bash.

• Demonstrated ability to independently define and lead infrastructure projects from design to production rollout.

• Strong instincts for incident management and the capability to lead during outages and conduct root-cause analysis.

• Experience with cloud and edge/on-premises infrastructure, particularly in industrial or manufacturing environments.

• Familiarity with Snowflake, BigQuery, Redshift, or Databricks.

• Knowledge of service mesh, zero-trust networking, or compliance frameworks for industrial/critical infrastructure such as SOC 2 or IEC 62443.

• A track record of building internal developer platforms or self-service infrastructure tools.

• Experience in scaling infrastructure teams or establishing technical direction at the Staff level.


🏝️ Benefits

• Significant opportunities for career growth and advancement.

• Comprehensive salary and equity compensation package.

• Innovative environment centered on transforming heavy industries through AI and automation.

• Collaborative culture that values innovation, discipline, and continuous improvement.

People also viewed

Chalice AI14 hours ago

Senior Site Reliability Engineer

US flagNew York OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$180k – $190k/year
ApplyView job
DaVita Kidney Care14 hours ago

Senior DevOps Engineer

US flagTennessee OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$115k – $183k/year
ApplyView job
WEX16 hours ago

SRE & Application Services Intern – Graduate/Master's

US flagUnited States OnlyInternshipDevOps & Site Reliability Engineer (SRE)$30 – $45/hour
ApplyView job
WEX16 hours ago

DevOps Engineer, AI Engineering Intern

US flagUnited States OnlyInternshipDevOps & Site Reliability Engineer (SRE)
ApplyView job
Dscout16 hours ago

Senior DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DaVita Kidney Care17 hours ago

Senior Manager, DevOps

US flagFlorida OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$115k – $183k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers