Remotery

Data Engineer III

atDysonRemoteUS flagCaliforniaFull-timeData EngineerMid-levelSenior$104k – $153k/year

Posted Jul 19

This is a fully remote position, open to applicants in California.

📋 Description

• Lead the architecture and design of intricate data pipelines utilizing the Databricks lakehouse architecture (Unity Catalog, Delta Lake, Structured Streaming).

• Establish the technical framework for data engineering projects, mentor junior engineers, and uphold code quality standards through effective leadership and code reviews.

• Design and develop data foundations that facilitate AI/ML capabilities, including feature stores, embedding pipelines, vector search indexes, and model training datasets.

• Ensure alignment of data engineering solutions with business strategies, including support for Agentic AI workloads.

• Take ownership of the health, scalability, and modernization of data infrastructure with Databricks as the primary platform, encompassing workload migration, compute optimization, and Unity Catalog adoption.

• Enhance pipeline performance (Delta Lake table layouts, clustering, Z-ordering) and implement monitoring/alerting best practices with defined SLAs.

• Construct data infrastructure that supports Agentic AI systems, including real-time data access layers, context retrieval pipelines, and data services accessible to agents.

• Collaborate across various functions, including DevOps, Platform Engineering, and MLOps, to integrate data solutions into the wider technology framework and shared AI infrastructure, such as MLflow registries, feature stores, and agent orchestration layers.

• Provide strategic consultation to Senior Leadership on complex projects and drive initiatives for continuous improvement.

• Advocate for data governance at all levels for data, models, and AI assets.

• Implement data quality strategies (master data management, validation rules, Delta Live Tables expectations) to ensure trust in enterprise data.

• Act as a liaison among data engineering, AI engineering, and business teams, promoting data literacy and stewardship.


⛳️ Requirements

• Bachelor's degree in Computer Science, Engineering, or a related field (Master's preferred).

• Over 5 years of experience with Python and SQL in data engineering for large-scale ML/analytics workloads.

• More than 5 years of experience in designing, building, and troubleshooting scalable ETL/ELT pipelines for critical production systems.

• At least 3 years of experience with cloud data services (AWS), container orchestration (Docker, Kubernetes), and Infrastructure as Code (Terraform, CloudFormation).

• A minimum of 3 years architecting ML workflows and data platforms with CI/CD, automated testing, and distributed processing (Spark).

• Over 3 years of experience collaborating cross-functionally with Data Science, MLOps, Platform Engineering, and DevOps teams.

• At least 3 years of experience implementing data quality testing and optimizing SQL/Python for cost/performance efficiency in the cloud.

• Comprehensive understanding of the full Data Science SDLC and experience mentoring engineers.

• Strongly preferred: Databricks & AI Platform expertise.

• Minimum of 2 years of hands-on experience with Databricks (Delta Lake, Unity Catalog, Databricks SQL).

• Familiarity with MLflow experiment tracking and model registry workflows.

• Experience in designing pipelines that support AI/ML inference, including real-time feature engineering, embedding generation, and context retrieval for LLM-based systems.

• Understanding of how data engineering underpins Agentic AI: agent-accessible data services, low-latency retrieval, and pipelines that enable autonomous multi-step workflows.

• Familiarity with Databricks Mosaic AI, Vector Search, and/or Feature Store.

• Awareness of FinOps practices, including compute cluster optimization and cost attribution by workload.

• Knowledge of Salesforce/Heroku data infrastructures.

• Experience with data virtualization tools (e.g., Dremio).

• Understanding of Platform Engineering principles and internal developer platforms.

• Experience with migrating from traditional data warehouses/lakes to unified lakehouse architectures.

• Familiarity with Odaseva data security and management solutions.


🏝️ Benefits

• Group health insurance benefits (medical, vision, dental).

• FSA and HSA healthcare accounts.

• Life and accident insurance.

• Adoption and fertility assistance.

• Paid parental leave of up to 6 weeks.

• Short/long term disability coverage.

• Paid time off for vacation, personal needs, and sick leave.

• Up to 17 days of Choice Time Off (CTO) per calendar year.

• Up to 11 paid holidays per calendar year.

• Opportunity to participate in the company's 401(k) savings and investment plan or deferred compensation plan, with an employer match of 100% on the first 3% of contributions.

People also viewed

Remofirst1 day ago

Senior Data Engineer

EG flagEgypt OnlyFull-timeData Engineer
ApplyView job
Omada Health1 day ago

Staff Software Engineer, Data Products

US flagUnited States OnlyFull-timeData Engineer$202.4k – $253k/year
ApplyView job
Moovx1 day ago

Senior Data Engineer

Latin AmericaFull-timeData Engineer
ApplyView job
BPO Global Services S.A.S1 day ago

Data Engineer

CO flagColombia OnlyFull-timeData Engineer$10/hour
ApplyView job
GSB Solutions1 day ago

Technical Program Manager – Data & Power Platform

MX flagMexico OnlyFull-timeData Engineer$113k/year
ApplyView job
DOMVS iT1 day ago

Data Engineer, Mid/Senior

BR flagBrazil OnlyFull-timeData Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers