
Data Engineer II
Posted Sep 17

Posted Sep 17
This is a fully remote position, open to applicants in India.
• Design, construct, and uphold scalable batch and real-time data pipelines utilizing Maxwell, Kafka, Spark, and dbt.
• Create and enhance Bronze, Silver, and Gold data models employing Medallion Architecture.
• Develop and sustain cloud-native data platforms using S3, Spark, Trino, and BigQuery.
• Design frameworks for data ingestion leveraging CDC, Kafka, and event-driven architectures.
• Establish and manage data warehouses and data marts for reporting and self-service analytics.
• Convert business requirements into scalable data solutions in collaboration with Product, Analytics, Engineering, and Business stakeholders.
• Develop reusable dbt models, testing frameworks, and comprehensive documentation.
• Optimize Spark jobs, Trino queries, and data storage layouts.
• Oversee critical data pipeline lifecycles, including monitoring, SLA compliance, and incident resolution.
• Create reusable frameworks for data platforms, automation, CI/CD pipelines, and engineering practices.
• Ensure data quality through validation, monitoring, lineage, observability, security, and governance.
• Provide trusted datasets, semantic models, and Metabase dashboards to support decision-making.
• 4+ years of practical experience in designing and developing scalable data platforms, data lakes, and data warehouses.
• Strong expertise in Spark (Scala, Python) and SQL.
• Proven experience in building production-grade data pipelines and distributed data processing applications.
• Hands-on experience with Apache Spark, including distributed data processing, performance tuning, and optimization.
• Experience in constructing batch and streaming data pipelines using Kafka, CDC/Maxwell, or similar event-driven architectures.
• Solid understanding of modern data lake architectures, including Medallion Architecture, data modeling, partitioning, and storage optimization.
• Familiarity with cloud-native data platforms and technologies like Amazon S3, BigQuery, Trino, or comparable analytics engines.
• Experience in designing dimensional models, star schemas, and dependable data marts.
• Hands-on experience with dbt, reusable models, automated testing, and documentation.
• Knowledge of data quality, observability, lineage, and engineering best practices.
• Experience in optimizing large-scale data pipelines, SQL queries, and distributed processing jobs.
• Understanding of CI/CD, Git-based development workflows, infrastructure automation, and contemporary software engineering best practices.
• Capability to independently manage projects from design to production.
• Experience collaborating with Product, Engineering, Analytics, and Business teams.
• Inclusive and diverse workplace.
• Remote work environment.
• Competitive compensation.
• Potential share options for certain roles.
• Regular training opportunities.
• Annual learning stipend.
• High degree of autonomy.
• Mentorship programs.
CuraLinc Healthcare
VSP Vision Care
Adoreal
Get handpicked remote jobs straight to your inbox weekly.