Data Science ML/Gen AI Engineer – Mid-Level

atIrth SolutionsRemoteIN flagIndiaFull-timeData ScientistMid-levelSenior

Posted Sep 10

This is a fully remote position, open to applicants in India.

πŸ“‹ Description

β€’ Contribute to medallion architecture pipelines (Bronze β†’ Silver β†’ Gold) utilizing Databricks.

β€’ Implement data quality checks, validation gates, data contracts, column-level lineage, policy-as-code, PII masking, obfuscation, and access controls.

β€’ Collaborate with data engineering and governance teams to enhance data reliability, discoverability, and documentation.

β€’ Explore, prototype, evaluate, and productionize machine learning and GenAI solutions for forecasting, anomaly detection, NLP, RAG, LLM-powered assistants/copilots, and predictive analytics.

β€’ Package and manage models using Unity Catalog model management/registries.

β€’ Design architectures for batch and streaming inference.

β€’ Define success metrics, KPIs, and A/B testing strategies in collaboration with Product and business stakeholders.

β€’ Transition successful experiments from prototype to production with SLAs, monitoring, documentation, and operational runbooks.

β€’ Build production workflows, jobs, and notebooks as infrastructure/assets-as-code utilizing Databricks Asset Bundles (DABs).

β€’ Implement CI/CD pipelines using GitHub Actions.

β€’ Create reliable, observable, scalable, and cost-efficient data and ML workloads.

β€’ Implement proactive monitoring and alerting, automate incident creation and tracking through Jira, and apply FinOps principles.

β€’ Contribute business metrics, definitions, and semantic models to Unity Catalog.

β€’ Support consumption through Power BI and Databricks AI/BI.

β€’ Develop and maintain trusted data products in collaboration with domain teams.

β€’ Implement secure data and ML architectures using RBAC and ABAC within Unity Catalog.

β€’ Manage credentials and secrets using Azure Key Vault or KMS.

β€’ Support compliance requirements for SOC 2, ISO 27001, GDPR, and PIPEDA.

β€’ Produce audit evidence for data lineage, access reviews, data retention, security controls, and disaster recovery testing.

β€’ Participate in governance and security reviews, addressing identified gaps.

β€’ Partner with Product, Data Engineering, Platform, and domain teams to achieve measurable customer and business impact.


⛳️ Requirements

β€’ 3–6 years of experience in Data Science, Machine Learning, or ML Engineering, with a demonstrated ability to transition models from development to production.

β€’ Proficient programming and data skills in Python, SQL, and Spark/PySpark.

β€’ Hands-on experience with Databricks, including Delta Lake, Unity Catalog, Databricks SQL (DBSQL), Jobs and Workflows, and Medallion architecture.

β€’ Strong grasp of ML fundamentals, including feature engineering, model training and selection, model evaluation and validation, model monitoring, data-quality monitoring, and model and data drift detection.

β€’ Practical experience with GenAI/LLM, including prompt engineering, Retrieval-Augmented Generation (RAG), vector databases/vector stores, LLM evaluation, AI safety and guardrails, as well as LLM latency, scalability, and cost tradeoffs.

β€’ Experience in implementing CI/CD for data and ML workloads, including GitHub Actions, Databricks Asset Bundles (DABs), DEV β†’ QA β†’ PROD environment promotion, and secrets and configuration management.

β€’ Familiarity with data contracts and data-quality frameworks, including schema governance, automated expectations/testing, validation, and quarantine/error-handling workflows.

β€’ Strong understanding of data security and compliance, including PII handling and protection, RBAC/ABAC, data residency requirements, and policy-as-code.

β€’ Excellent communication and collaboration skills with Product, Engineering, Data, and domain teams.

β€’ Ability to produce technical documentation, such as Architecture Decision Records (ADRs), runbooks, experiment reports, and operational documentation.

β€’ Preferred: Experience with Microsoft Azure, AWS, geospatial data and analytics, streaming and real-time data, MLflow, Unity Catalog Model Serving, data and ML observability, FinOps, disaster recovery/business continuity planning, resilience practices, and utilities, energy, infrastructure, or public works.

β€’ Nice to have: Expertise in predictive/risk-scoring/failure-prediction models, anomaly detection, time-series forecasting, GIS/geospatial asset data, asset-integrity data, regulatory/compliance/audit reporting, and operational risk indicators.


🏝️ Benefits

β€’ Competitive compensation package based on experience and qualifications.

β€’ Medical, Dental, and Vision Insurance.

β€’ 401(k) Plan with Company Match.

β€’ Generous Paid Time Off (PTO).

β€’ Company-Paid Holidays.

β€’ Flexible Work Options / work-from-home opportunities, depending on role and business needs.

β€’ On-Call Compensation for eligible on-call shifts.

People also viewed

PODS15 hours ago

Data Scientist II

US flagFlorida OnlyFull-timeData Scientist
ApplyView job
USA TODAY Network1 day ago

Data Science Manager

US flagNew York OnlyFull-timeData Scientist$189.3k – $195k/year
ApplyView job
General Dynamics Information Technology1 day ago

Principal Data Scientist, AI Development and Governance

US flagUnited States OnlyFull-timeData Scientist$119k – $161k/year
ApplyView job
v4c.ai1 day ago

Data Scientist – Onshore

US flagUnited States OnlyFull-timeData Scientist
ApplyView job
Maker Lab1 day ago

Data Scientist

IN flagIndia OnlyFreelanceData Scientist
ApplyView job
Map of Ag1 day ago

Senior Data Scientist

GB flagUnited Kingdom OnlyFull-timeData ScientistΒ£55k – Β£70k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers