
Engineer – Gen AI
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in Tennessee.
• Design, construct, and uphold robust ETL/ELT pipelines that integrate on-premise systems, AWS services, and Azure platforms for Snowflake and downstream AI applications.
• Create and sustain feature stores and analytically optimized datasets to support machine learning workflows.
• Develop pipelines for generative AI applications, including the extraction, transformation, chunking, and loading of both structured and unstructured data into vector databases.
• Act as a Snowflake power user and technical lead, executing data modeling, Snowpipe automation, and optimizing compute and storage resources.
• Carry out non-invasive data extraction from legacy systems, ensuring system stability is maintained.
• Design and oversee cross-platform workflows utilizing Airflow, AWS Step Functions, and Azure Data Factory.
• Collaborate with IT, database, infrastructure, and security teams to address connectivity and access issues, securing production approval for integrations.
• Implement automated frameworks for data quality, validation, and observability.
• Enhance storage, compute efficiency, and query performance across Snowflake, AWS, and Azure environments.
• Collaborate with MLOps, Data Science, and AI teams to convert experimental use cases into scalable, production-ready data solutions.
• Perform additional responsibilities as assigned.
• A Master’s degree in Computer Science, Data Engineering, or a related discipline from an accredited institution is preferred.
• A minimum of six (6) years of practical data engineering experience, showcasing a history of building production-grade pipelines for Data Science and AI within multi-cloud environments, or an equivalent combination of education and experience is required.
• Expert-level understanding of Snowflake architecture, including data sharing, performance optimization, and integration with external cloud AI services.
• Advanced, hands-on proficiency with AWS (S3, Glue, Lambda) and Azure (Data Factory, Synapse) data services.
• Mastery of Python, SQL, and PySpark is essential.
• Extensive experience with data orchestration and containerization technologies, such as Docker.
• Demonstrated ability to interface with on-premise SQL, Mainframe extracts, and flat files, transforming them for contemporary cloud usage.
• Strong grasp of Machine Learning data requirements, including feature engineering.
• Solid understanding of Generative AI data requirements, including vectorization and embedding pipelines.
• Ability to thrive in a collaborative team setting.
• Capacity to meet or surpass Performance Competencies.
• Proficient in managing work-related stress and juggling multiple priorities while adhering to deadlines.
• Willingness to travel as necessary.
• Work-life balance.
• Reasonable accommodations provided when applicable and appropriate.
• Equal Opportunity Employer.
• Drug-Free Workplace.
Mercor
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.