
Senior MLOps Engineer
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in Ukraine, +1 more country.
• Design and oversee a scalable machine learning infrastructure on GCP utilizing Vertex AI, Google Kubernetes Engine (GKE), Google Cloud Storage (GCS), Cloud Run, and GPU/TPU compute instances.
• Manage the complete deployment lifecycle for machine learning models.
• Develop high-throughput, low-latency inference services through containerization and specialized serving frameworks like Triton Inference Server, vLLM, and MLflow.
• Create automated and reproducible pipelines for model training, testing, evaluation, and deployment using Airflow, Vertex AI Pipelines, and GitHub Actions.
• Establish monitoring for system health and ML-specific metrics, such as latency, throughput, uptime, feature drift, prediction accuracy, and shifts in data distribution.
• Deliver scalable training environments, optimized runtime infrastructure, and standardized deployment templates for AI engineers.
• Collaborate with Data Engineers on feature stores, dataset versioning, and workflows for stream and batch data processing.
• Guide the transition of AI prototypes and notebooks into robust, secure, and auto-scaling microservices.
• Work alongside AI Researchers, Data Engineers, and Backend teams to connect experimentation and production systems.
• A minimum of 5 years of practical experience in designing, deploying, and maintaining production machine learning workloads within cloud environments.
• In-depth, hands-on experience with Google Cloud Platform (GCP), including Vertex AI, Cloud Storage, GKE, Cloud Run, and IAM/VPC configurations.
• Proficiency in containerization (Docker, Kubernetes/GKE) and specialized serving tools (Triton, vLLM, MLflow).
• Demonstrated success with workflow orchestrators (Airflow, Vertex AI Pipelines) and modern CI/CD tools (GitHub Actions, ArgoCD).
• Strong experience in managing cloud resources using Terraform.
• Skilled in Python and SQL for scripting, automation, API development, and data manipulation.
• Practical experience with logging, telemetry, and drift detection tools (Grafana, Prometheus, GCP Cloud Monitoring, or specialized ML observability frameworks).
• Experience with large-scale LLM or Deep Learning inference/training workloads.
• GCP Professional Machine Learning Engineer or GCP Professional Cloud Architect certifications are preferred.
• Familiarity with feature stores such as Feast or Vertex AI Feature Store.
• Opportunity to learn new technologies, products, and markets in a fast-paced, growth-oriented environment.
• Collaborate with talented individuals at a company that values its people.
• An inclusive community with a commitment to maintaining a workplace free from discrimination and harassment.
Shield AI
Weekday (YC W21)
Roadpass Digital
Get handpicked remote jobs straight to your inbox weekly.