
Platform Engineer – MLOps
Posted 14 hours ago

Posted 14 hours ago
This is a fully remote position, open to applicants in United States.
• Develop and sustain CI/CD pipelines for machine learning workloads, focusing on the testing, packaging, and promotion of models and pipelines.
• Take responsibility for ML lifecycle tools related to experiment tracking, model registry, versioning, and promotion gates.
• Establish procedures for promoting development to production, incorporating validation gates, rollback mechanisms, and audit trails.
• Execute model deployment and serving strategies for both batch inference and real-time endpoints.
• Create and oversee observability for ML workloads, addressing pipeline health, model and data drift, performance metrics, latency, and cost considerations.
• Collaborate with data scientists to translate notebooks and experiments into well-governed, reliable, and repeatable pipelines.
• Manage and enhance Databricks workspaces, Unity Catalog metastores, and catalogs while implementing suitable access controls.
• Provide support for the AWS infrastructure that underpins the Lakehouse, with a focus on IAM, networking, S3, and associated platform services.
• Implement least-privilege access and Unity Catalog governance across data and ML assets.
• Monitor and optimize Databricks jobs and compute resources for both cost efficiency and performance.
• Maintain architectural documentation and runbooks for deployment, troubleshooting, and production assistance.
• Collaborate with data engineering and data science teams to establish and uphold platform standards, automation, and documentation.
• Convert ML workload requirements into dependable, governed, and reusable platform capabilities.
• Over 5 years of relevant experience in platform engineering, DevOps, MLOps, ML engineering, or similar engineering roles.
• Practical experience with the ML lifecycle, including experiment tracking, model registry, versioning, and deployment (MLflow or equivalent).
• Proven experience in constructing CI/CD pipelines utilizing Git-based workflows for ML or data workloads.
• Familiarity with Databricks platform functionalities, such as workspaces, Unity Catalog, compute, jobs/workflows, permissions, and environment settings.
• Strong Python scripting abilities for automation and pipeline development.
• Capability to collaborate closely with data scientists and data engineers to translate ML requirements into reliable platform functionalities.
• Proficiency in English, both written and spoken.
• Experience with AWS, especially IAM, networking, S3, and related platform services (preferred).
• Background in Infrastructure as Code with Terraform or Terragrunt (preferred).
• Knowledge of Databricks Asset Bundles, Delta Live Tables, or Databricks Workflows (preferred).
• Familiarity with feature stores or feature platforms such as Chalk, and real-time model serving (preferred).
• Experience with notebook-based data science environments like Domino Data Lab (preferred).
• Exposure to on-premises or hybrid Kubernetes environments and Git-based deployment workflows using GitLab (preferred).
• Databricks ML Associate or AWS certifications (preferred).
• Experience with observability tools like CloudWatch and Prometheus/Grafana, as well as cost-optimization techniques (preferred).
• Permanent full-time employment.
• Option for remote work.
• A global and multicultural environment fostering diverse perspectives and collaboration.
• A startup atmosphere characterized by rapid movement and impact-driven initiatives.
• An ownership mindset, allowing engineers to take pride in their creations.
• A collaborative, friendly, open, curious, and supportive organizational culture.
Parallel Partners
Playbypoint
Alectrona
Cloudera
Get handpicked remote jobs straight to your inbox weekly.