
Senior MLOps Engineer
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in Bulgaria.
• Design and oversee scalable machine learning infrastructure on GCP utilizing Vertex AI, GKE, GCS, Cloud Run, and GPU/TPU compute resources.
• Take full ownership of the deployment lifecycle for machine learning models from start to finish.
• Develop high-throughput, low-latency inference services through containerization and specialized serving frameworks.
• Create automated and reproducible pipelines for model training, testing, evaluation, and deployment processes.
• Set up monitoring systems for both overall system health and ML-specific metrics, including triggers for automated retraining.
• Offer scalable training environments, optimized runtime infrastructure, and standardized deployment templates for AI engineers.
• Collaborate with Data Engineers to manage feature stores, dataset versioning, and data processing workflows for both stream and batch data.
• Guide the transformation of AI prototypes and notebooks into robust, secure, and auto-scaling microservices.
• Work alongside AI Researchers, Data Engineers, and Backend teams to connect experimentation with enterprise-level production systems.
• A minimum of 5 years of practical experience in designing, deploying, and maintaining production ML workloads within cloud environments.
• Extensive, hands-on experience with Google Cloud Platform (GCP), specifically Vertex AI, Cloud Storage, GKE, Cloud Run, and IAM/VPC configurations.
• Proficient in containerization technologies, including Docker and Kubernetes/GKE.
• Expertise in specialized serving tools such as Triton, vLLM, and MLflow.
• Demonstrated experience with Airflow and Vertex AI Pipelines.
• Solid background with contemporary CI/CD tools such as GitHub Actions and ArgoCD.
• Proficient in managing cloud resources using Terraform.
• Strong skills in Python and SQL for scripting, automation, API development, and data manipulation.
• Experience with logging, telemetry, and drift detection tools, including Grafana, Prometheus, GCP Cloud Monitoring, or specialized ML observability frameworks.
• Experience with large-scale LLM or deep learning inference/training workloads is a plus.
• Possession of GCP Professional Machine Learning Engineer or GCP Professional Cloud Architect certifications is a plus.
• Familiarity with feature stores such as Feast or Vertex AI Feature Store is a plus.
• Opportunity to acquire knowledge of new technologies, products, and markets in a dynamic, growth-focused environment.
• Chance to witness personal impact within a small, agile organization.
• Collaborate with skilled individuals in an inclusive community.
• Commitment to non-discrimination and anti-harassment policies.
Quora
Amgen
Get handpicked remote jobs straight to your inbox weekly.