
Data Science ML/Gen AI Engineer β Mid-Level
Posted Sep 10

Posted Sep 10
This is a fully remote position, open to applicants in India.
β’ Contribute to medallion architecture pipelines (Bronze β Silver β Gold) utilizing Databricks.
β’ Implement data quality checks, validation gates, data contracts, column-level lineage, policy-as-code, PII masking, obfuscation, and access controls.
β’ Collaborate with data engineering and governance teams to enhance data reliability, discoverability, and documentation.
β’ Explore, prototype, evaluate, and productionize machine learning and GenAI solutions for forecasting, anomaly detection, NLP, RAG, LLM-powered assistants/copilots, and predictive analytics.
β’ Package and manage models using Unity Catalog model management/registries.
β’ Design architectures for batch and streaming inference.
β’ Define success metrics, KPIs, and A/B testing strategies in collaboration with Product and business stakeholders.
β’ Transition successful experiments from prototype to production with SLAs, monitoring, documentation, and operational runbooks.
β’ Build production workflows, jobs, and notebooks as infrastructure/assets-as-code utilizing Databricks Asset Bundles (DABs).
β’ Implement CI/CD pipelines using GitHub Actions.
β’ Create reliable, observable, scalable, and cost-efficient data and ML workloads.
β’ Implement proactive monitoring and alerting, automate incident creation and tracking through Jira, and apply FinOps principles.
β’ Contribute business metrics, definitions, and semantic models to Unity Catalog.
β’ Support consumption through Power BI and Databricks AI/BI.
β’ Develop and maintain trusted data products in collaboration with domain teams.
β’ Implement secure data and ML architectures using RBAC and ABAC within Unity Catalog.
β’ Manage credentials and secrets using Azure Key Vault or KMS.
β’ Support compliance requirements for SOC 2, ISO 27001, GDPR, and PIPEDA.
β’ Produce audit evidence for data lineage, access reviews, data retention, security controls, and disaster recovery testing.
β’ Participate in governance and security reviews, addressing identified gaps.
β’ Partner with Product, Data Engineering, Platform, and domain teams to achieve measurable customer and business impact.
β’ 3β6 years of experience in Data Science, Machine Learning, or ML Engineering, with a demonstrated ability to transition models from development to production.
β’ Proficient programming and data skills in Python, SQL, and Spark/PySpark.
β’ Hands-on experience with Databricks, including Delta Lake, Unity Catalog, Databricks SQL (DBSQL), Jobs and Workflows, and Medallion architecture.
β’ Strong grasp of ML fundamentals, including feature engineering, model training and selection, model evaluation and validation, model monitoring, data-quality monitoring, and model and data drift detection.
β’ Practical experience with GenAI/LLM, including prompt engineering, Retrieval-Augmented Generation (RAG), vector databases/vector stores, LLM evaluation, AI safety and guardrails, as well as LLM latency, scalability, and cost tradeoffs.
β’ Experience in implementing CI/CD for data and ML workloads, including GitHub Actions, Databricks Asset Bundles (DABs), DEV β QA β PROD environment promotion, and secrets and configuration management.
β’ Familiarity with data contracts and data-quality frameworks, including schema governance, automated expectations/testing, validation, and quarantine/error-handling workflows.
β’ Strong understanding of data security and compliance, including PII handling and protection, RBAC/ABAC, data residency requirements, and policy-as-code.
β’ Excellent communication and collaboration skills with Product, Engineering, Data, and domain teams.
β’ Ability to produce technical documentation, such as Architecture Decision Records (ADRs), runbooks, experiment reports, and operational documentation.
β’ Preferred: Experience with Microsoft Azure, AWS, geospatial data and analytics, streaming and real-time data, MLflow, Unity Catalog Model Serving, data and ML observability, FinOps, disaster recovery/business continuity planning, resilience practices, and utilities, energy, infrastructure, or public works.
β’ Nice to have: Expertise in predictive/risk-scoring/failure-prediction models, anomaly detection, time-series forecasting, GIS/geospatial asset data, asset-integrity data, regulatory/compliance/audit reporting, and operational risk indicators.
β’ Competitive compensation package based on experience and qualifications.
β’ Medical, Dental, and Vision Insurance.
β’ 401(k) Plan with Company Match.
β’ Generous Paid Time Off (PTO).
β’ Company-Paid Holidays.
β’ Flexible Work Options / work-from-home opportunities, depending on role and business needs.
β’ On-Call Compensation for eligible on-call shifts.
USA TODAY Network
General Dynamics Information Technology
Get handpicked remote jobs straight to your inbox weekly.