
Senior AI Data Engineer
Posted Sep 14

Posted Sep 14
This is a fully remote position, open to applicants in Ohio.
• Gather, design, and transform intricate data into formats that can be understood by Data Scientists and Business Analysts.
• Create pipelines to transfer raw data into Azure Synapse utilizing Spark, Python, SQL, and C# in alignment with Data Warehouse and Data Lakehouse architectural guidelines.
• Build capabilities for machine learning and regression analysis using Spark-Python-Pandas, OpenAI, and Azure ML.
• Guide junior Data Engineers and establish protocols and best practices.
• Collaborate with Data Analysts and Business Analysts to clarify project requirements and steer execution.
• Provide support for peer reviews and ensure compliance with coding standards.
• Demonstrate expertise in the Software Development Life Cycle.
• Design, develop, and sustain scalable batch and streaming data pipelines tailored for AI and machine learning applications.
• Construct and maintain ETL/ELT processes for training, testing, and production datasets.
• Assist in the development and management of feature stores.
• Develop and oversee data lakes, lakehouses, and data warehouse solutions.
• Integrate and manage vector databases and vector storage for RAG and other AI-related applications.
• Collaborate with engineers and architects to create data architectures and enhance pipeline performance.
• Implement data quality, validation, observability, and monitoring solutions.
• Uphold security, governance, privacy, and regulatory compliance standards.
• Document data flows, architectures, schemas, and operational procedures.
• Support and modify integrations related to Model Context Protocol.
• Work alongside AI Data Engineering, IT Data Engineering, AI Engineering, Infrastructure, Security, and business stakeholders.
• Minimum of 4 years of experience in Python development pertinent to data engineering (Spark, Pandas, etc.).
• At least 4 years of experience in SQL related to data engineering.
• Experience in designing and executing complex data pipelines while ensuring data quality and consistency.
• Strong understanding of data warehouse and Delta Lake design principles.
• Solid background in data analytics.
• Comprehensive knowledge of the Software Development Life Cycle.
• Bachelor’s Degree or relevant certifications, along with equivalent years of experience.
• Familiarity with AI/ML data workflows.
• Proficient in Python and SQL.
• Understanding of ETL/ELT processes and data modeling.
• Knowledge of relational and NoSQL databases.
• Capability to support data pipelines focused on AI and data quality practices.
• Microsoft Cloud Certification is a plus.
• Familiarity with Machine Learning and AI is advantageous.
• Web Development experience is beneficial.
• Preference for candidates with vector database and RAG experience.
• Preferred experience with Spark, Kafka, dbt, dlt, Hadoop, or NiFi.
• Knowledge of cloud data services like Azure Data Factory, AWS Glue, or GCP Dataflow is preferred.
• Experience with PyTorch or other machine learning frameworks is advantageous.
• Familiarity with Docker, Kubernetes, and CI/CD practices is preferred.
• Experience with MCP integrations is a plus.
• Willingness to travel less than 10%.
• Equal Opportunity Employer/Protected Veterans/Individuals with Disabilities.
• Reasonable accommodations available for applicants with disabilities.
• Remote work arrangement.
CuraLinc Healthcare
VSP Vision Care
Adoreal
Get handpicked remote jobs straight to your inbox weekly.