Data Scientist – Mid-Level

Posted Sep 10

This is a fully remote position, open to applicants in India.

📋 Description

• Contribute to the Databricks medallion architecture pipelines spanning from Bronze to Silver and ultimately to Gold.

• Implement checks for data quality, validation gates, data contracts, lineage, governance, policy-as-code, PII masking, obfuscation, and access controls.

• Explore, prototype, assess, and productionize machine learning and GenAI solutions for forecasting, anomaly detection, NLP, RAG, LLM assistants/copilots, and predictive analytics.

• Package and manage models utilizing Unity Catalog model management and registries.

• Design architectures for batch and streaming inference.

• Define success metrics, KPIs, and A/B testing strategies in collaboration with Product and business stakeholders.

• Transition successful experiments from prototype to production, ensuring SLAs, monitoring, documentation, and operational runbooks are in place.

• Construct production workflows, jobs, and notebooks as infrastructure/assets-as-code using Databricks Asset Bundles.

• Implement CI/CD pipelines utilizing GitHub Actions.

• Design data and ML workloads that are reliable, observable, scalable, and cost-effective.

• Implement monitoring, alerting, automated Jira incident tracking, FinOps practices, resource tagging, workload policies, and cost monitoring.

• Contribute semantic models, business metrics, and definitions to Unity Catalog.

• Support data consumption through Power BI and Databricks AI/BI.

• Develop and maintain trusted data products alongside domain teams.

• Establish secure data and ML architectures using RBAC and ABAC.

• Manage credentials and secrets via Azure Key Vault or KMS.

• Support compliance with SOC 2, ISO 27001, GDPR, and PIPEDA.

• Maintain audit evidence for lineage, access reviews, retention, security controls, and disaster recovery testing.

• Engage in governance and security reviews, addressing any identified gaps.

• Deliver data and AI use cases from conceptualization through prototyping and production to tangible business impact.


⛳️ Requirements

• 3–6 years of experience in Data Science, Machine Learning, or ML Engineering, demonstrating a strong history of transitioning models from development to production.

• Proficient programming and data skills in Python, SQL, and Spark/PySpark.

• Practical experience with Databricks, including Delta Lake, Unity Catalog, Databricks SQL, Jobs and Workflows, and Medallion architecture.

• Solid understanding of feature engineering, model training and selection, model evaluation and validation, model monitoring, data-quality monitoring, and model and data drift detection.

• Hands-on experience with GenAI/LLM, including prompt engineering, RAG, vector databases/vector stores, LLM evaluation, AI safety and guardrails, and considerations of LLM latency, scalability, and cost.

• Experience with CI/CD implementation for data and ML workloads, encompassing GitHub Actions, Databricks Asset Bundles, environment promotion, and secrets/configuration management.

• Familiarity with data contracts and data-quality frameworks, including schema governance, automated expectations/testing, validation, and quarantine/error-handling workflows.

• Strong knowledge of data security and compliance, covering PII handling and protection, RBAC/ABAC, data residency requirements, and policy-as-code.

• Capable of producing technical documentation, such as ADRs, runbooks, experiment reports, and operational documentation.

• Preferred: Experience with Microsoft Azure, AWS, geospatial data and analytics, streaming/real-time data, MLflow, Unity Catalog Model Serving, data and ML observability, FinOps, DR/BCP, resilience, and in utilities/energy/infrastructure industries.

• Nice-to-have: Familiarity with predictive/risk-scoring/failure-prediction models, anomaly detection, time-series forecasting, GIS/geospatial ML features, asset-integrity risk models, regulatory/audit reporting, and operational decision-support tools.


🏝️ Benefits

• Competitive compensation package based on experience and qualifications.

• Medical, Dental, and Vision Insurance.

• 401(k) Plan with Company Match.

• Generous Paid Time Off (PTO).

• Company-Paid Holidays.

• Flexible Work Options / work-from-home opportunities, depending on role and business needs.

• On-Call Compensation for eligible on-call shifts.

People also viewed

PODS15 hours ago

Data Scientist II

US flagFlorida OnlyFull-timeData Scientist
ApplyView job
USA TODAY Network1 day ago

Data Science Manager

US flagNew York OnlyFull-timeData Scientist$189.3k – $195k/year
ApplyView job
General Dynamics Information Technology1 day ago

Principal Data Scientist, AI Development and Governance

US flagUnited States OnlyFull-timeData Scientist$119k – $161k/year
ApplyView job
v4c.ai1 day ago

Data Scientist – Onshore

US flagUnited States OnlyFull-timeData Scientist
ApplyView job
Maker Lab1 day ago

Data Scientist

IN flagIndia OnlyFreelanceData Scientist
ApplyView job
Map of Ag1 day ago

Senior Data Scientist

GB flagUnited Kingdom OnlyFull-timeData Scientist£55k – £70k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers