Senior DevOps Engineer, AI Platform

Posted Sep 2

This is a fully remote position, open to applicants in Canada.

📋 Description

• Transform technical designs for applications and platforms into production-ready cloud infrastructure with minimal oversight.

• Create, provision, manage, and troubleshoot Kubernetes environments, focusing on Azure Kubernetes Service and Oracle Kubernetes Engine.

• Assist with AI workloads, including LiteLLM-based gateways, Python agent runtimes, RAG workers, MCP services, background workers, and asynchronous processing pipelines.

• Design and oversee ingress and egress networking, load balancers, DNS, TLS, private connectivity, routing, NAT, firewalls, network policies, and service-to-service communication.

• Construct and maintain infrastructure for web applications and backend services, including APIs, databases, caches, queues, scheduled jobs, and event-driven workloads.

• Develop and sustain CI/CD pipelines utilizing Jenkins and Bitbucket, integrating Docker, Helm, Kubernetes, ArgoCD, and container registries.

• Automate the provisioning and configuration of infrastructure using Terraform, Helm, Kubernetes manifests, Python, Bash, and related tools.

• Implement comprehensive observability through metrics, logs, distributed tracing, dashboards, alerts, health checks, and SLOs.

• Take ownership of production readiness, incident resolution, root cause analysis, scalability, reliability, and infrastructure cost optimization.

• Create reusable infrastructure patterns that enable engineering teams to deploy new services swiftly and consistently.

• Assess required cloud resources, Kubernetes configurations, namespaces, scaling models, and supporting services for new workloads.

• Configure connectivity, identities, and secrets; provision and manage PostgreSQL, Redis, RabbitMQ, storage, and shared services.

• Build CI/CD workflows in Jenkins and Bitbucket for build, testing, container publishing, deployment, validation, and rollback.

• Establish observability, health checks, dashboards, alerts, capacity monitoring, and operational runbooks prior to production launch.

• Manage infrastructure delivery through UAT and production, collaborating with architects and engineers on design decisions.


⛳️ Requirements

• A minimum of 7 years of experience in DevOps, SRE, Platform Engineering, Cloud Infrastructure, or a comparable role.

• Extensive hands-on experience managing production Kubernetes environments and in-depth knowledge of networking, scheduling, storage, autoscaling, security, and troubleshooting.

• Strong experience with Microsoft Azure, including AKS, networking, identity, storage, and monitoring.

• Experience with OCI is preferred, or a proven ability to work across multiple cloud providers.

• In-depth knowledge of cloud networking, including virtual networks, subnets, routing, NAT, load balancers, private networking, DNS, TLS, firewalls, ingress, and egress.

• Significant experience with Jenkins, Bitbucket, Docker, Terraform, Helm, Kubernetes, and Infrastructure as Code.

• Proven track record of supporting production web applications and backend services, including REST APIs, microservices, background workers, and asynchronous architectures.

• Practical experience with databases, caching, and messaging systems such as PostgreSQL, Redis, RabbitMQ, or similar technologies.

• Experience in implementing production observability using OpenTelemetry, Grafana, Prometheus, Sentry, cloud monitoring, or similar tools.

• Strong Linux and systems troubleshooting skills in a production environment.

• Working knowledge of Python, particularly for backend services built with frameworks like FastAPI.

• Familiarity with at least one additional programming language such as C#, Java, Go, JavaScript, or TypeScript.

• Understanding of HTTP, HTTPS, DNS, TCP/IP, proxies, authentication, APIs, connection pooling, caching, concurrency, queues, retries, dead letter queues, and asynchronous processing.

• Ability to analyze application logs and stack traces to diagnose latency, memory, CPU, connection, and dependency issues.


🏝️ Benefits

• Comprehensive health, dental, and vision insurance.

• Generous paid time off and holiday leave.

• Opportunities for professional development and continuous learning.

• Flexible work arrangements and remote work options.

• Collaborative and supportive work environment.

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers