Remotery

Principal ML Ops Engineer

Posted 10 hours ago

This is a fully remote position, open to applicants in Ukraine.

📋 Description

• Develop and manage production-grade model serving infrastructure utilizing frameworks like vLLM, TGI, Triton, or similar alternatives.

• Create and implement reliable deployment pipelines featuring blue/green and canary rollout strategies for machine learning models.

• Design and uphold auto-scaling systems, multi-model serving architectures, and intelligent request routing layers.

• Enhance GPU utilization, memory efficiency, network throughput, and performance of model artifact storage.

• Create observability systems for monitoring inference latency, throughput, GPU usage, cost metrics, and overall system health.

• Oversee model registries and CI/CD pipelines to facilitate automated and reproducible model deployments.

• Manage the entire lifecycle of ML systems from development to production, including operational support and on-call duties.

• Establish engineering best practices and contribute to platform scalability within a dynamic startup environment.


⛳️ Requirements

• A minimum of 4 years of experience in ML Ops, Platform Engineering, Site Reliability Engineering, or comparable infrastructure roles focused on ML systems.

• Practical experience with model serving frameworks such as vLLM, TGI, Triton, or similar.

• A solid background in container orchestration and managing GPU-based workloads in production settings.

• Familiarity with MLOps tools including model registries, experiment tracking, and automated deployment pipelines.

• Expertise in Python and infrastructure-as-code tools (e.g., Terraform, Helm, or similar).

• Strong comprehension of distributed systems, performance optimization, and production reliability engineering.

• Capability to effectively leverage AI coding assistants to enhance development and debugging processes.

• Possess an ownership mindset with the ability to work independently in a remote-first setting.


🏝️ Benefits

• Take charge of critical infrastructure that supports a rapidly expanding AI-native cloud platform.

• Construct foundational ML inference systems from the ground up in a high-growth, well-funded startup environment.

• Operate at the intersection of distributed systems, GPU computing, and sustainable cloud architecture.

• Acquire profound expertise in next-generation AI infrastructure and large-scale model serving systems.

• Play a key role in shaping core engineering decisions and establishing best practices that will scale with the organization.

People also viewed

TTEC1 hour ago

Member of Technical Staff, Machine Learning, Artificial Intelligence

US flagCalifornia OnlyFull-timeMachine Learning Engineer$160k – $190k/year
ApplyView job
Grafana Labs2 hours ago

Senior Machine Learning Engineer, Developer Advocacy

DE flagGermany OnlyFull-timeMachine Learning Engineer€97k – €116.4k/year
ApplyView job
Forward Financing1 day ago

Senior Software Engineer, MLOps

CA flagCanada OnlyFull-timeMachine Learning Engineer$175k – $220k/year
ApplyView job
Maze1 day ago

Machine Learning Engineer

GB flagUnited Kingdom OnlyFull-timeMachine Learning Engineer£100k – £135k/year
ApplyView job
3Pillar Global2 days ago

Architect / Principal, AI ML Engineer

US flagUnited States OnlyFull-timeMachine Learning Engineer
ApplyView job
Skylum2 days ago

Machine Learning Engineer

UA flagUkraine OnlyFull-timeMachine Learning Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers