
Senior MLOps Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Estonia, +1 more country.
β’ Design and oversee scalable ML infrastructure on GCP utilizing Vertex AI, Google Kubernetes Engine, Google Cloud Storage, Cloud Run, and GPU/TPU compute instances.
β’ Manage the complete deployment lifecycle for machine learning models.
β’ Develop high-throughput, low-latency inference services through containerization and specialized serving frameworks.
β’ Create automated, reproducible pipelines for model training, testing, evaluation, and deployment.
β’ Establish monitoring for system health and ML-specific metrics, including feature drift, prediction accuracy, and shifts in data distribution.
β’ Deliver scalable training environments, optimized runtime infrastructure, and standardized deployment templates for AI and research engineers.
β’ Collaborate with Data Engineers on feature stores, dataset versioning, and workflows for stream and batch data processing.
β’ Guide the evolution of AI prototypes and notebooks into robust, secure, auto-scaling microservices.
β’ Work alongside AI Researchers, Data Engineers, and Backend teams to connect experimentation with enterprise-grade production systems.
β’ Minimum 5 years of hands-on experience in designing, deploying, and maintaining production ML workloads within cloud environments.
β’ In-depth practical knowledge of Google Cloud Platform, including Vertex AI, Cloud Storage, GKE, Cloud Run, and IAM/VPC configurations.
β’ Proficiency in containerization technologies such as Docker and Kubernetes/GKE.
β’ Expertise in specialized serving tools like Triton, vLLM, and MLflow.
β’ Proven experience with Airflow and Vertex AI Pipelines.
β’ Demonstrated experience with contemporary CI/CD tools, including GitHub Actions and ArgoCD.
β’ Solid experience in managing cloud resources utilizing Terraform.
β’ Proficient in Python and SQL.
β’ Practical experience with logging, telemetry, and drift detection tools such as Grafana, Prometheus, GCP Cloud Monitoring, or specialized ML observability frameworks.
β’ Experience with large-scale LLM or deep learning inference/training workloads is preferred.
β’ GCP Professional Machine Learning Engineer or GCP Professional Cloud Architect certification is preferred.
β’ Familiarity with feature stores like Feast or Vertex AI Feature Store is preferred.
β’ Opportunity to address real customer challenges in cybersecurity.
β’ Chance to observe personal impact within a dynamic, agile organization.
β’ Career advancement and opportunities to explore new technologies, products, and markets.
β’ Collaborate with skilled colleagues in an inclusive environment.
β’ Commitment to non-discrimination and anti-harassment.
β’ Equal opportunity for all applicants regardless of protected characteristics.
Torc Robotics
Bose Corporation
ZoomInfo
Coderio
Get handpicked remote jobs straight to your inbox weekly.