
Senior DevOps Engineer, AI Platform
Posted Sep 2

Posted Sep 2
This is a fully remote position, open to applicants in Canada.
• Transform technical designs for applications and platforms into production-ready cloud infrastructure with minimal oversight.
• Create, provision, manage, and troubleshoot Kubernetes environments, focusing on Azure Kubernetes Service and Oracle Kubernetes Engine.
• Assist with AI workloads, including LiteLLM-based gateways, Python agent runtimes, RAG workers, MCP services, background workers, and asynchronous processing pipelines.
• Design and oversee ingress and egress networking, load balancers, DNS, TLS, private connectivity, routing, NAT, firewalls, network policies, and service-to-service communication.
• Construct and maintain infrastructure for web applications and backend services, including APIs, databases, caches, queues, scheduled jobs, and event-driven workloads.
• Develop and sustain CI/CD pipelines utilizing Jenkins and Bitbucket, integrating Docker, Helm, Kubernetes, ArgoCD, and container registries.
• Automate the provisioning and configuration of infrastructure using Terraform, Helm, Kubernetes manifests, Python, Bash, and related tools.
• Implement comprehensive observability through metrics, logs, distributed tracing, dashboards, alerts, health checks, and SLOs.
• Take ownership of production readiness, incident resolution, root cause analysis, scalability, reliability, and infrastructure cost optimization.
• Create reusable infrastructure patterns that enable engineering teams to deploy new services swiftly and consistently.
• Assess required cloud resources, Kubernetes configurations, namespaces, scaling models, and supporting services for new workloads.
• Configure connectivity, identities, and secrets; provision and manage PostgreSQL, Redis, RabbitMQ, storage, and shared services.
• Build CI/CD workflows in Jenkins and Bitbucket for build, testing, container publishing, deployment, validation, and rollback.
• Establish observability, health checks, dashboards, alerts, capacity monitoring, and operational runbooks prior to production launch.
• Manage infrastructure delivery through UAT and production, collaborating with architects and engineers on design decisions.
• A minimum of 7 years of experience in DevOps, SRE, Platform Engineering, Cloud Infrastructure, or a comparable role.
• Extensive hands-on experience managing production Kubernetes environments and in-depth knowledge of networking, scheduling, storage, autoscaling, security, and troubleshooting.
• Strong experience with Microsoft Azure, including AKS, networking, identity, storage, and monitoring.
• Experience with OCI is preferred, or a proven ability to work across multiple cloud providers.
• In-depth knowledge of cloud networking, including virtual networks, subnets, routing, NAT, load balancers, private networking, DNS, TLS, firewalls, ingress, and egress.
• Significant experience with Jenkins, Bitbucket, Docker, Terraform, Helm, Kubernetes, and Infrastructure as Code.
• Proven track record of supporting production web applications and backend services, including REST APIs, microservices, background workers, and asynchronous architectures.
• Practical experience with databases, caching, and messaging systems such as PostgreSQL, Redis, RabbitMQ, or similar technologies.
• Experience in implementing production observability using OpenTelemetry, Grafana, Prometheus, Sentry, cloud monitoring, or similar tools.
• Strong Linux and systems troubleshooting skills in a production environment.
• Working knowledge of Python, particularly for backend services built with frameworks like FastAPI.
• Familiarity with at least one additional programming language such as C#, Java, Go, JavaScript, or TypeScript.
• Understanding of HTTP, HTTPS, DNS, TCP/IP, proxies, authentication, APIs, connection pooling, caching, concurrency, queues, retries, dead letter queues, and asynchronous processing.
• Ability to analyze application logs and stack traces to diagnose latency, memory, CPU, connection, and dependency issues.
• Comprehensive health, dental, and vision insurance.
• Generous paid time off and holiday leave.
• Opportunities for professional development and continuous learning.
• Flexible work arrangements and remote work options.
• Collaborative and supportive work environment.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.