
Senior MLOps Engineer
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in Lithuania.
• Design and oversee scalable machine learning infrastructure on GCP utilizing Vertex AI, Google Kubernetes Engine (GKE), Google Cloud Storage (GCS), Cloud Run, and GPU/TPU computing instances.
• Manage the complete deployment lifecycle for machine learning models.
• Develop high-throughput, low-latency inference services through containerization and specialized serving frameworks like Triton Inference Server, vLLM, and MLflow.
• Create automated and reproducible pipelines for model training, testing, evaluation, and deployment using Airflow, Vertex AI Pipelines, and GitHub Actions.
• Establish monitoring for system health and ML-specific metrics, such as feature drift, prediction accuracy, and shifts in data distribution.
• Provide scalable training environments, optimized runtime infrastructures, and standardized deployment templates for AI engineers.
• Partner with Data Engineers to integrate model pipelines with feature stores, dataset versioning, and stream/batch data processing workflows.
• Lead the conversion of AI prototypes and notebooks into robust, secure, and auto-scaling microservices.
• Collaborate with AI Researchers, Data Engineers, and Backend teams to connect experimentation and production systems.
• Minimum of 5 years of hands-on experience in designing, deploying, and maintaining production machine learning workloads within cloud environments.
• Extensive practical knowledge of Google Cloud Platform (GCP), including Vertex AI, Cloud Storage, GKE, Cloud Run, and IAM/VPC configurations.
• Proficient in containerization (Docker, Kubernetes/GKE) and specialized serving tools (Triton, vLLM, MLflow).
• Demonstrated success with workflow orchestrators (Airflow, Vertex AI Pipelines) and modern CI/CD tools (GitHub Actions, ArgoCD).
• Strong experience in managing cloud resources using Terraform.
• Expertise in Python and SQL for scripting, automation, API development, and data manipulation.
• Practical experience with logging, telemetry, and drift detection tools (Grafana, Prometheus, GCP Cloud Monitoring, or specialized ML observability frameworks).
• Experience in executing large-scale LLM or Deep Learning inference/training workloads.
• Possession of GCP Professional Machine Learning Engineer or GCP Professional Cloud Architect certifications.
• Knowledge of feature stores (e.g., Feast, Vertex AI Feature Store).
• Opportunity to acquire new technologies, products, and markets in a dynamic, growth-focused environment.
• Collaborate with talented individuals in a company that values its people.
• Be part of an inclusive community dedicated to preventing discrimination and harassment.
Quora
Amgen
Get handpicked remote jobs straight to your inbox weekly.