
Data Engineer II
Posted Sep 18

Posted Sep 18
This is a fully remote position, open to applicants in India.
• Design, construct, and sustain scalable batch and real-time data pipelines utilizing Maxwell, Kafka, Spark, and dbt.
• Develop and enhance data models adhering to the Medallion Architecture (Bronze, Silver, Gold).
• Create and manage cloud-native data platforms employing S3, Spark, Trino, and BigQuery.
• Design frameworks for data ingestion that leverage CDC (Maxwell), Kafka, and event-driven architectures.
• Build, optimize, and maintain data warehouses and data marts.
• Collaborate with Product Managers, Data Analysts, Backend Engineers, and Business stakeholders to transform requirements into effective data solutions.
• Develop reusable dbt models, testing frameworks, and comprehensive documentation.
• Optimize Spark jobs, Trino queries, and storage configurations.
• Oversee the lifecycle of key data pipelines, ensuring availability, monitoring, SLA compliance, and incident resolution.
• Create reusable frameworks, automation, CI/CD pipelines, and best engineering practices for the core data platform.
• Ensure data integrity through validation, monitoring, lineage tracking, and observability, while implementing security and governance protocols.
• Provide reliable datasets, semantic models, and Metabase dashboards to facilitate analytics and decision-making.
• 4+ years of practical experience in designing and constructing scalable data platforms, data lakes, and data warehouses.
• Strong expertise in Spark (Scala, Python) and SQL.
• Experience in developing production-grade data pipelines and distributed data processing applications.
• Hands-on experience with Apache Spark, including distributed data processing, performance tuning, and optimization.
• Experience in constructing batch and streaming data pipelines utilizing Kafka, CDC/Maxwell, or equivalent event-driven architectures.
• Solid understanding of contemporary data lake architectures, including Medallion Architecture, data modeling, partitioning, and storage optimization.
• Familiarity with cloud-native data platforms and technologies such as Amazon S3, BigQuery, Trino, or similar analytics engines.
• Experience in designing dimensional models, star schemas, and dependable data marts.
• Hands-on experience with dbt, including reusable models, automated testing, and documentation.
• Strong knowledge of data quality, observability, lineage, and engineering best practices.
• Experience in optimizing large-scale data pipelines, SQL queries, and distributed processing tasks.
• Familiarity with CI/CD, Git-based development workflows, infrastructure automation, and modern software engineering best practices.
• Ability to independently manage projects from design through to production.
• Strong communication and stakeholder management abilities.
• Experience working collaboratively across Product, Engineering, Analytics, and Business teams.
• Passion for scalable data platforms, improving developer experience, ensuring platform reliability, and pursuing operational excellence.
• Inclusive and Diverse Environment: We promote an inclusive and diverse workplace that values innovation and supports remote work options.
• Competitive compensation packages.
• Potential share options available for certain roles.
• Regular training opportunities.
• Annual learning stipend.
• High degree of autonomy in your work.
• Mentorship programs.
• Ambitious goals that foster both professional and company growth.
CuraLinc Healthcare
VSP Vision Care
Adoreal
Get handpicked remote jobs straight to your inbox weekly.