
Senior AI Data Engineer
Posted Sep 14

Posted Sep 14
This is a fully remote position, open to applicants in Ohio.
• Gather, design, and transform intricate data into formats that can be understood by Data Scientists and Business Analysts.
• Establish pipelines for transferring raw data into Azure Synapse utilizing Spark, Python, SQL, and C# in line with Data Warehouse and Data Lakehouse architectural guidelines.
• Create machine learning and regression analysis functionalities using Spark-Python-Pandas, OpenAI, and Azure ML.
• Guide junior Data Engineers by providing mentorship and establishing protocols and guardrails.
• Collaborate with Data Analysts and Business Analysts to clarify project requirements and facilitate execution.
• Offer peer review assistance and ensure adherence to relevant coding standards.
• Apply expertise in the Software Development Life Cycle.
• Design, develop, and sustain scalable batch and streaming data pipelines for AI and machine learning tasks.
• Construct and maintain ETL/ELT processes for training, testing, and production datasets.
• Assist in the development and management of feature stores.
• Oversee the maintenance of data lakes, lakehouses, and data warehouse solutions.
• Integrate and manage vector databases and storage solutions for RAG and other AI applications.
• Collaborate with engineers and architects on data architecture, enhancing pipeline performance, and ensuring AI-ready data platforms.
• Implement data quality, validation, observability, and monitoring features.
• Address security, governance, privacy, and regulatory compliance necessities.
• Document data flows, architectures, schemas, and operational processes.
• Support and adjust Model Context Protocol integrations under senior supervision.
• Work in partnership with AI Data Engineering, IT Data Engineering, AI Engineering, Infrastructure, Security, and business stakeholders.
• Over 4 years of experience in Python development related to data engineering (e.g., Spark, pandas).
• More than 4 years of experience in SQL pertaining to data engineering.
• Proven experience in designing and executing complex data pipelines, while ensuring data quality and consistency.
• Strong grasp of data warehouse and Delta Lake design principles.
• Solid background in data analytics.
• Comprehensive understanding of the Software Development Life Cycle.
• Bachelor's Degree or relevant certifications alongside equivalent years of experience.
• Familiarity with AI/ML data workflows.
• Proficient in Python and SQL.
• Understanding of ETL/ELT processes and data modeling.
• Knowledge of relational and NoSQL databases.
• Capability to support AI-focused data pipelines and uphold data quality practices.
• Preferred: Experience with vector databases and RAG.
• Preferred: Familiarity with Spark, Kafka, dbt, dlt, Hadoop, or NiFi.
• Preferred: Experience with cloud data services such as Azure Data Factory, AWS Glue, or GCP Dataflow.
• Preferred: Knowledge of PyTorch or other machine learning frameworks.
• Preferred: Experience with Docker, Kubernetes, and CI/CD.
• Preferred: Experience with MCP integrations.
• Travel requirement is less than 10%.
• Microsoft Cloud Certification available as a bonus opportunity.
• Familiarity with Machine Learning and AI considered a bonus opportunity.
• Web Development experience is viewed as a plus.
CuraLinc Healthcare
VSP Vision Care
Adoreal
Get handpicked remote jobs straight to your inbox weekly.