
Data Engineer
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Brazil.
• Develop, sustain, and enhance S3 → Databricks/Airflow → dbt pipelines.
• Expand the data platform to accommodate the increasing data volumes of the organization.
• Oversee production pipelines, pinpointing failures, delays, or unusual data occurrences.
• Diagnose and modify workflows when issues arise.
• Establish metrics and monitor systems with an emphasis on operating costs, data quality, and consistency.
• Manage infrastructure via code utilizing Terraform.
• Deploy the data platform utilizing Kubernetes (EKS) and Argo.
• Ensure data governance with OpenMetadata, focusing on availability, scalability, and access control.
• Recognize PII and implement masking or anonymization before data is utilized for analytics or product purposes.
• Define and refine self-service ELT/ETL tools.
• Prepare and supply data for predictive models and LLMs, including RAG and embeddings when applicable.
• Assess new technologies and ensure the production stack operates reliably.
• Assist internal teams using the platform by addressing technical inquiries and advocating best practices.
• Proficiency in software engineering (Python, Java, or Scala) and familiarity with software development best practices (e.g., Clean Architecture and TDD).
• Understanding of Lambda data architecture and analytical data modeling.
• Experience with AWS, primarily focusing on S3 for ingestion.
• Familiarity with Databricks for distributed processing tasks.
• Knowledge of Airflow for orchestrating pipelines.
• Experience with Docker and Kubernetes, including both stateless and stateful workloads.
• Proficient in Terraform for infrastructure as code.
• Experience with Argo CD and declarative deployment via Kubernetes/EKS.
• Competence in dbt for analytical modeling and transformation.
• Familiarity with OpenMetadata for data governance, lineage, and access control.
• Knowledge of language model APIs, such as OpenAI, Cohere, and Hugging Face.
• Understanding of security protocols and sensitive data management, including PII identification, masking, and anonymization.
• Experience in DataOps, monitoring, alerting, and resolving failures in production pipelines.
• Familiarity with Kappa architecture as a complementary asset to Lambda architecture.
• Awareness of alternative open-source solutions for data ingestion, processing, and delivery.
• Understanding of data security and governance best practices, referencing DAMA conceptually.
• Interest in or knowledge of the fundamentals of LLMs and generative AI applied to data.
• Experience with streaming data processing technologies such as Kafka, Flink, and Druid.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible working hours and remote work options.
• Opportunities for professional development and career advancement.
• Collaborative and innovative work environment.
Katapult Labs
Magna Legal Services
Huron
Strategic Systems International
Get handpicked remote jobs straight to your inbox weekly.