
Data Engineer
Posted Jul 25

Posted Jul 25
This is a fully remote position, open to applicants in Texas.
⢠Creates, constructs, and oversees robust ETL/ELT pipelines that gather data from on-premise systems, AWS services (S3, RDS), and Azure platforms (Blob Storage, Azure SQL), centralizing and refining data for use in Snowflake and subsequent AI applications.
⢠Develops and sustains feature stores and analytically optimized datasets that facilitate machine learning processes, ensuring that data is clean, versioned, reproducible, and statistically sound for Data Science teams.
⢠Designs data pipelines that empower generative AI applications, including the automated extraction, transformation, chunking, and loading of both structured and unstructured data into vector databases across AWS and Azure environments.
⢠Serves as a Snowflake power user and technical lead, executing advanced data modeling strategies, automating Snowpipe, and optimizing compute and storage to accommodate high-concurrency analytics and AI tasks.
⢠Implements non-intrusive data extraction techniques to access crucial data from legacy systems that have been in place for decades, while maintaining system stability and preventing disruptions to essential business operations.
⢠Crafts and oversees intricate, cross-platform data workflows utilizing orchestration tools such as Airflow, AWS Step Functions, and Azure Data Factory to ensure dependable, synchronized data movement throughout the organizationās multi-cloud infrastructure.
⢠Collaborates closely with central IT, database administrators, infrastructure, and security teams to address connectivity and access issuesāincluding PrivateLink, IAM, network segmentation, and firewall controlsāwhile securing production approval for new data integrations.
⢠Establishes automated frameworks for data quality, validation, and observability to identify data drift, anomalies, and integrity challenges that could adversely affect production analytics, machine learning, or AI systems.
⢠Enhances efficiency across the data ecosystem by optimizing storage, compute usage, and query performance in Snowflake, AWS, and Azure, ensuring responsible cost management and measurable ROI for Transformation Office initiatives.
⢠Functions as a dedicated engineering partner to MLOps, Data Science, and AI teams, quickly adapting to changing data requirements and transforming experimental use cases into scalable, production-ready data solutions.
⢠Masterās degree in Computer Science, Data Engineering, or a related field from an accredited institution is preferred.
⢠A minimum of six (6) years of practical data engineering experience, demonstrating a history of building production-grade pipelines for Data Science and AI in multi-cloud settings or an equivalent combination of education and experience is required.
⢠Expert-level understanding of Snowflake architecture, including data sharing, performance tuning, and the integration of Snowflake with external cloud AI services.
⢠Advanced, hands-on experience with AWS (S3, Glue, Lambda) and Azure (Data Factory, Synapse) data services.
⢠Mastery of Python, SQL, and PySpark.
⢠Extensive experience with data orchestration and containerization (Docker).
⢠Proven capability to work with legacy technologies (on-premise SQL, Mainframe extracts, flat files) and adapt them for modern cloud use.
⢠A comprehensive understanding of the specific data requirements for Machine Learning (feature engineering) and Generative AI (vectorization and embedding pipelines).
⢠A proactive attitude, able to navigate enterprise bureaucracy and technical debt to deliver code at the pace required by a Transformation Office.
⢠Capacity to work effectively in a team environment.
⢠Ability to meet or surpass Performance Competencies.
⢠Work-life balance.
⢠Opportunities for professional development.
⢠Equity and performance bonus opportunities.
Omada Health
BPO Global Services S.A.S
Get handpicked remote jobs straight to your inbox weekly.