
Data Engineer
Posted Sep 17

Posted Sep 17
This is a fully remote position, open to applicants in Brazil.
• Create a scalable data platform that integrates various sources for convenient access.
• Design and enhance data tools for orchestration, governance, Data-Lakehouse, Business Intelligence, and related functions.
• Ensure seamless operation of data systems for analysts, scientists, and engineers.
• Optimize data pipelines for ingestion, processing, and output in a microservices architecture.
• Build, maintain, and monitor ETL/ELT processes, orchestrating workflows using Temporal.
• Diagnose and enhance the performance, scalability, and reliability of data infrastructure, including S3, Apache Iceberg, and ClickHouse.
• Collaborate with data scientists, analysts, and backend engineers to understand their data requirements and provide effective solutions.
• Implement and advocate for best practices in data quality, governance, and security across the platform.
• A minimum of 3 years' experience as a Data Engineer or in a comparable data infrastructure position.
• Strong command of SQL and practical experience with data modeling.
• Familiarity with data lake/lakehouse architectures, such as Apache Iceberg and S3.
• Experience with analytical/columnar databases, including ClickHouse.
• Proven experience in constructing and orchestrating ETL/ELT pipelines, utilizing tools such as Temporal or Airflow.
• Proficient programming skills in Python and/or Scala/Java.
• Background working within a microservices architecture and cloud environments, preferably AWS.
• Self-driven, with strong multitasking abilities and a proven record as a team player.
• Excellent communication skills, capable of working both independently and collaboratively.
• Practical experience with Apache Spark or similar technologies for large-scale data processing.
• Professional proficiency in both written and spoken English.
• This role is focused on batch data processing rather than real-time streaming.
• Nice to have: experience with open-source data platforms and tools.
• Nice to have: familiarity with BI and visualization tools like Superset, Looker, Tableau, or Metabase.
• Nice to have: experience with Docker and Kubernetes.
• Nice to have: knowledge of infrastructure-as-code and CI/CD practices.
• Nice to have: experience with AWS EMR and cloud-based Apache Spark workloads.
• Nice to have: familiarity with AI-assisted development tools such as GitHub Copilot or Cursor.
• Remote work flexibility.
• Full-time employment opportunity.
• Chance to contribute to an open-source data platform.
CuraLinc Healthcare
VSP Vision Care
Adoreal
Get handpicked remote jobs straight to your inbox weekly.