
Data Engineer
Posted Jul 17

Posted Jul 17
This is a fully remote position, open to applicants in India.
• Design, construct, and optimize efficient ETL/ELT pipelines to reliably handle and transform millions of daily data points.
• Engage directly with the open-source technology stack by building and maintaining production data workflows utilizing tools like dbt, ClickHouse, and orchestration frameworks such as Airflow or Dagster.
• Develop and fine-tune databases by writing and optimizing complex SQL across SQL Server and Postgres, focusing on schema design, indexing, performance tuning, query optimization, and root-cause analysis.
• Assist in the migration process from Azure Databricks to a more open and flexible AWS-based architecture, ensuring minimal disruption to high-volume daily data delivery.
• Contribute to the development of proofs-of-concept to evaluate new tools and patterns, providing clear feedback to guide the team’s adoption strategies.
• Leverage AI-assisted tools, such as GitHub Copilot, Claude, and ChatGPT, to enhance pipeline development, SQL generation, debugging, testing, and documentation, while diligently reviewing and validating outputs prior to deployment.
• Ensure data quality and reliability by implementing robust testing, validation, monitoring, observability, and CI/CD practices to maintain accurate data and healthy pipelines at scale.
• Understand and integrate systems both upstream and downstream of your pipelines, from data ingestion to client-facing platforms, to guarantee clean, end-to-end data delivery.
• Collaborate effectively with a distributed team, working alongside engineers in North America and India, participating in code reviews, and gaining insights into the Advertising and Market Research domain.
• Experience: 7–8+ years in data engineering/ETL, showcasing a strong record as a hands-on engineer in building and operating production data pipelines.
• Core databases: Proficient in SQL and RDBMS with substantial hands-on experience in SQL Server and Postgres, including schema design, performance tuning, and complex query optimization.
• Open-source and modern stack: Practical experience with tools such as dbt, ClickHouse, and open-source orchestration tools (e.g., Airflow, Dagster), with the ability to select the appropriate tool for specific tasks.
• Programming: Strong skills in Python (or a similar language) for constructing and automating data pipelines.
• AI-assisted development: Proven experience utilizing AI coding assistants and effective prompting techniques to enhance efficiency, with the capability to verify, test, and refine AI-generated code and queries.
• Cloud: Hands-on experience with a major cloud platform; AWS is preferred as we are transitioning from Azure Databricks to AWS.
• Engineering fundamentals: Familiarity with Git, conducting code reviews, and writing tested, maintainable, and well-documented code.
• Education: Bachelor’s or Master’s degree in Computer Science, Engineering, or a related discipline, or equivalent practical experience.
• Competitive salary and performance-based bonuses.
• Comprehensive health and wellness benefits.
• Opportunities for professional development and career advancement.
• Flexible working hours and remote work options.
Omada Health
BPO Global Services S.A.S
Get handpicked remote jobs straight to your inbox weekly.