
Data Engineer III
Posted Jul 19

Posted Jul 19
This is a fully remote position, open to applicants in California.
• Lead the architecture and design of intricate data pipelines utilizing the Databricks lakehouse architecture (Unity Catalog, Delta Lake, Structured Streaming).
• Establish the technical framework for data engineering projects, mentor junior engineers, and uphold code quality standards through effective leadership and code reviews.
• Design and develop data foundations that facilitate AI/ML capabilities, including feature stores, embedding pipelines, vector search indexes, and model training datasets.
• Ensure alignment of data engineering solutions with business strategies, including support for Agentic AI workloads.
• Take ownership of the health, scalability, and modernization of data infrastructure with Databricks as the primary platform, encompassing workload migration, compute optimization, and Unity Catalog adoption.
• Enhance pipeline performance (Delta Lake table layouts, clustering, Z-ordering) and implement monitoring/alerting best practices with defined SLAs.
• Construct data infrastructure that supports Agentic AI systems, including real-time data access layers, context retrieval pipelines, and data services accessible to agents.
• Collaborate across various functions, including DevOps, Platform Engineering, and MLOps, to integrate data solutions into the wider technology framework and shared AI infrastructure, such as MLflow registries, feature stores, and agent orchestration layers.
• Provide strategic consultation to Senior Leadership on complex projects and drive initiatives for continuous improvement.
• Advocate for data governance at all levels for data, models, and AI assets.
• Implement data quality strategies (master data management, validation rules, Delta Live Tables expectations) to ensure trust in enterprise data.
• Act as a liaison among data engineering, AI engineering, and business teams, promoting data literacy and stewardship.
• Bachelor's degree in Computer Science, Engineering, or a related field (Master's preferred).
• Over 5 years of experience with Python and SQL in data engineering for large-scale ML/analytics workloads.
• More than 5 years of experience in designing, building, and troubleshooting scalable ETL/ELT pipelines for critical production systems.
• At least 3 years of experience with cloud data services (AWS), container orchestration (Docker, Kubernetes), and Infrastructure as Code (Terraform, CloudFormation).
• A minimum of 3 years architecting ML workflows and data platforms with CI/CD, automated testing, and distributed processing (Spark).
• Over 3 years of experience collaborating cross-functionally with Data Science, MLOps, Platform Engineering, and DevOps teams.
• At least 3 years of experience implementing data quality testing and optimizing SQL/Python for cost/performance efficiency in the cloud.
• Comprehensive understanding of the full Data Science SDLC and experience mentoring engineers.
• Strongly preferred: Databricks & AI Platform expertise.
• Minimum of 2 years of hands-on experience with Databricks (Delta Lake, Unity Catalog, Databricks SQL).
• Familiarity with MLflow experiment tracking and model registry workflows.
• Experience in designing pipelines that support AI/ML inference, including real-time feature engineering, embedding generation, and context retrieval for LLM-based systems.
• Understanding of how data engineering underpins Agentic AI: agent-accessible data services, low-latency retrieval, and pipelines that enable autonomous multi-step workflows.
• Familiarity with Databricks Mosaic AI, Vector Search, and/or Feature Store.
• Awareness of FinOps practices, including compute cluster optimization and cost attribution by workload.
• Knowledge of Salesforce/Heroku data infrastructures.
• Experience with data virtualization tools (e.g., Dremio).
• Understanding of Platform Engineering principles and internal developer platforms.
• Experience with migrating from traditional data warehouses/lakes to unified lakehouse architectures.
• Familiarity with Odaseva data security and management solutions.
• Group health insurance benefits (medical, vision, dental).
• FSA and HSA healthcare accounts.
• Life and accident insurance.
• Adoption and fertility assistance.
• Paid parental leave of up to 6 weeks.
• Short/long term disability coverage.
• Paid time off for vacation, personal needs, and sick leave.
• Up to 17 days of Choice Time Off (CTO) per calendar year.
• Up to 11 paid holidays per calendar year.
• Opportunity to participate in the company's 401(k) savings and investment plan or deferred compensation plan, with an employer match of 100% on the first 3% of contributions.
Omada Health
BPO Global Services S.A.S
Get handpicked remote jobs straight to your inbox weekly.