
Senior MLOps Engineer
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in Poland.
• Design and oversee scalable machine learning infrastructure on GCP utilizing Vertex AI, Google Kubernetes Engine (GKE), Google Cloud Storage (GCS), Cloud Run, and GPU/TPU compute resources.
• Take responsibility for the complete deployment lifecycle of machine learning models.
• Develop high-throughput, low-latency inference services using containerization and specialized serving frameworks such as Triton Inference Server, vLLM, and MLflow.
• Create automated, reproducible pipelines for model training, testing, evaluation, and deployment using Airflow, Vertex AI Pipelines, and GitHub Actions.
• Implement monitoring solutions for system health and machine learning-specific metrics to enable automated retraining triggers.
• Provide scalable training environments, optimized runtime infrastructure, and standardized deployment templates for AI engineers.
• Collaborate with Data Engineers to integrate model pipelines with feature stores, dataset versioning, and both stream and batch data processing workflows.
• Lead the transition of AI prototypes and notebooks into robust, secure, auto-scaling microservices.
• Work alongside AI Researchers, Data Engineers, and Backend teams to link experimentation with enterprise-grade production systems.
• Minimum of 5 years of practical experience in designing, deploying, and maintaining production-level ML workloads in cloud environments.
• Extensive hands-on experience with Google Cloud Platform (GCP), including Vertex AI, Cloud Storage, GKE, Cloud Run, and IAM/VPC configurations.
• Proficient in containerization technologies (Docker, Kubernetes/GKE) and specialized serving tools (Triton, vLLM, MLflow).
• Proven experience with workflow orchestrators (Airflow, Vertex AI Pipelines) and modern CI/CD tools (GitHub Actions, ArgoCD).
• Solid background in managing cloud resources using Terraform.
• Proficient in Python and SQL for scripting, automation, API development, and data manipulation.
• Hands-on experience with logging, telemetry, and drift detection tools (Grafana, Prometheus, GCP Cloud Monitoring, or specialized ML observability frameworks).
• Experience with running large-scale LLM or Deep Learning inference/training workloads.
• GCP Professional Machine Learning Engineer or GCP Professional Cloud Architect certifications.
• Familiarity with feature stores like Feast or Vertex AI Feature Store.
• Opportunity to learn about new technologies, products, and markets in a dynamic, growth-focused environment.
• Collaborate with talented individuals in an inclusive company where people are valued.
• Chance to make a tangible impact within a nimble and scrappy organization.
Shield AI
Weekday (YC W21)
Roadpass Digital
Get handpicked remote jobs straight to your inbox weekly.