Senior MLOps Engineer

Posted 5 days ago

This is a fully remote position, open to applicants in Ukraine, +1 more country.

📋 Description

• Design and oversee a scalable machine learning infrastructure on GCP utilizing Vertex AI, Google Kubernetes Engine (GKE), Google Cloud Storage (GCS), Cloud Run, and GPU/TPU compute instances.

• Manage the complete deployment lifecycle for machine learning models.

• Develop high-throughput, low-latency inference services through containerization and specialized serving frameworks like Triton Inference Server, vLLM, and MLflow.

• Create automated and reproducible pipelines for model training, testing, evaluation, and deployment using Airflow, Vertex AI Pipelines, and GitHub Actions.

• Establish monitoring for system health and ML-specific metrics, such as latency, throughput, uptime, feature drift, prediction accuracy, and shifts in data distribution.

• Deliver scalable training environments, optimized runtime infrastructure, and standardized deployment templates for AI engineers.

• Collaborate with Data Engineers on feature stores, dataset versioning, and workflows for stream and batch data processing.

• Guide the transition of AI prototypes and notebooks into robust, secure, and auto-scaling microservices.

• Work alongside AI Researchers, Data Engineers, and Backend teams to connect experimentation and production systems.


⛳️ Requirements

• A minimum of 5 years of practical experience in designing, deploying, and maintaining production machine learning workloads within cloud environments.

• In-depth, hands-on experience with Google Cloud Platform (GCP), including Vertex AI, Cloud Storage, GKE, Cloud Run, and IAM/VPC configurations.

• Proficiency in containerization (Docker, Kubernetes/GKE) and specialized serving tools (Triton, vLLM, MLflow).

• Demonstrated success with workflow orchestrators (Airflow, Vertex AI Pipelines) and modern CI/CD tools (GitHub Actions, ArgoCD).

• Strong experience in managing cloud resources using Terraform.

• Skilled in Python and SQL for scripting, automation, API development, and data manipulation.

• Practical experience with logging, telemetry, and drift detection tools (Grafana, Prometheus, GCP Cloud Monitoring, or specialized ML observability frameworks).

• Experience with large-scale LLM or Deep Learning inference/training workloads.

• GCP Professional Machine Learning Engineer or GCP Professional Cloud Architect certifications are preferred.

• Familiarity with feature stores such as Feast or Vertex AI Feature Store.


🏝️ Benefits

• Opportunity to learn new technologies, products, and markets in a fast-paced, growth-oriented environment.

• Collaborate with talented individuals at a company that values its people.

• An inclusive community with a commitment to maintaining a workplace free from discrimination and harassment.

People also viewed

Shield AI10 hours ago

Staff Deep Learning Engineer, State Estimation

US flagUnited States OnlyFull-timeMachine Learning Engineer$200k – $300k/year
ApplyView job
Weekday (YC W21)1 day ago

ML Engineer

IN flagIndia OnlyFull-timeMachine Learning Engineer₹2.5M – ₹5M/year
ApplyView job
Roadpass Digital1 day ago

Senior AI/ML Engineer

US flagUnited States OnlyFull-timeMachine Learning Engineer
ApplyView job
MWDN1 day ago

AI/ML Engineer

HR flagCroatia OnlyFull-timeMachine Learning Engineer
ApplyView job
Quora1 day ago

Software Engineer, New Grad – Machine Learning Platform

US flagUnited States, +1 more countryFull-timeMachine Learning Engineer$97.6k – $139k/year
ApplyView job
Amgen1 day ago

Principal Machine Learning Engineer

US flagUnited States OnlyFull-timeMachine Learning Engineer$187.4k – $253.5k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers