
MLOps, LLMOps Engineer – Mid-Level
Posted Sep 10

Posted Sep 10
This is a fully remote position, open to applicants in India.
• Operationalize the entire ML lifecycle, which includes training, evaluation, packaging, deployment, and monitoring, on Databricks.
• Implement ML workflows utilizing Bronze → Silver → Gold medallion architecture with Delta Lake.
• Establish model-management patterns in Unity Catalog for governance, lineage, discovery, and access control.
• Develop reusable templates for ML/LLM jobs, workflows, and deployment processes.
• Create and maintain cluster policies that align with enterprise platform guardrails.
• Productionize ML and GenAI models for use cases related to damage prevention, asset integrity, land management, and stakeholder engagement.
• Design, construct, and maintain production-grade LLM and RAG pipelines.
• Implement vector search and retrieval architectures utilizing technologies such as Databricks Vector Search.
• Deploy and manage model-serving and inference endpoints while optimizing performance, scalability, reliability, and cost.
• Implement batch, streaming, and online inference patterns, establishing service-level expectations and SLAs.
• Integrate data contracts, quality gates, PII controls, data residency policy-as-code, and end-to-end lineage into the ML lifecycle.
• Use Databricks Asset Bundles and GitHub Actions to version, test, and promote ML/LLM assets across DEV, QA, and PROD environments.
• Construct automated unit, integration, regression, data-quality, and model-quality test suites along with deployment gates.
• Instrument pipelines and services for SLOs, monitoring, alerting, Jira ticket creation, model/data/feature drift detection, and LLM observability.
• Develop runbooks, troubleshooting procedures, operational documentation, disaster-recovery procedures, and participate in on-call rotations and DR testing.
• Enforce cost and ownership tags, support showback/chargeback reporting, monitor AI infrastructure and inference costs, and identify optimization opportunities.
• Collaborate with teams from Data Science, Data Engineering, Platform, Product, Security, and domain areas.
• 3–5 years of experience in MLOps, LLMOps, ML Engineering, Data Engineering, or ML engineering focused on platforms.
• Hands-on experience with Databricks Jobs and Workflows, Delta Lake, Unity Catalog, and Databricks SQL Warehouses.
• Experience in building and maintaining CI/CD pipelines for data and ML workloads using GitHub Actions and Databricks Asset Bundles (DABs).
• Familiarity with DEV → QA → PROD environment promotion and parameterized deployments.
• Experience in secure secrets management using Azure Key Vault, AWS KMS/Secrets Manager, or equivalent technologies.
• Strong grasp of data contracts, schema governance, and automated data/feature validation.
• Experience with validation frameworks similar to Great Expectations or other rule-based data-quality solutions.
• Experience in building observable production pipelines with metrics, dashboards, alerting, and SLO monitoring.
• Practical experience with RBAC/ABAC, Unity Catalog security, PII detection and obfuscation, private networking, data-access controls, and policy-as-code for data residency.
• Strong proficiency in Python and SQL.
• Working knowledge of distributed computing and job orchestration within Databricks/Spark environments.
• Ability to troubleshoot production ML/data workloads and participate in operational support and incident resolution.
• Preferred: hands-on experience with LLM/GenAI workflows, prompt engineering, RAG, LLM evaluation, AI safety and guardrails, retrieval/response-quality evaluation, latency optimization, and token/API-cost optimization.
• Preferred: experience with geospatial data and analytics, including PostGIS, spatial joins, spatial indexing and tiling, coordinate systems and projections, and GIS-based feature engineering.
• Preferred: experience integrating Power BI with Databricks SQL Warehouses and semantic layers.
• Preferred: practical knowledge of FinOps, including resource tagging, budget management, cost monitoring, showback/chargeback, and cost anomaly detection.
• Preferred: knowledge of Databricks disaster-recovery patterns, including Delta Lake Deep Clone, Delta Sharing, cross-region recovery, tiered RTO/RPO strategies, and DR testing.
• Preferred: hands-on experience with Microsoft Azure and AWS.
• Preferred: understanding of cloud-native security patterns, including Private Link, VPC/VNet connectivity and peering, egress restrictions, KMS, AWS Secrets Manager, Azure Key Vault, and data-plane isolation.
• Competitive Salary – A competitive compensation package based on experience and qualifications.
• Medical, Dental, and Vision Insurance – Comprehensive insurance coverage to support you and your family.
• 401(k) Plan with Company Match.
• Generous Paid Time Off (PTO) – Time off to support work-life balance and personal needs.
• Company-Paid Holidays – Paid holidays throughout the year.
• Flexible Work Options – Work-from-home opportunities are available, depending on role and business needs.
• On-Call Compensation – Additional pay for eligible on-call shifts.
Workana
Sowelo Consulting sp. z o.o. sp. k.
Sowelo Consulting sp. z o.o. sp. k.
H&R Block
Get handpicked remote jobs straight to your inbox weekly.