
Data Engineer
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in Argentina, +5 more countries.
• Develop batch and streaming ingestion and transformation pipelines utilizing Spark, Kafka, dbt, and Airflow.
• Strategically address idempotency, backfills, schema evolution, and late-arriving data.
• Design data warehouses and lakehouses on platforms such as Snowflake, BigQuery, Redshift, or Databricks.
• Take ownership of partitioning, file layout, query cost, and performance considerations.
• Construct chunking, embedding, and indexing pipelines for RAG retrieval systems.
• Manage retrieval infrastructure using pgvector, Pinecone, Qdrant, or Azure AI Search.
• Set up data quality tests, expectations, lineage, and alert notifications.
• Implement PII classification, masking, and manage row- and column-level access, retention, deletion, and audit trails.
• Deploy and maintain containerized systems on Azure or AWS.
• Oversee CI/CD, orchestration, observability, cost management, and runtime budgets.
• Engage within client environments, repositories, standups, and occasionally participate in customer calls.
• Utilize AI-assisted development tools and automated codebase audits to uphold security, cost efficiency, and architectural integrity.
• Collaborate with client teams in alignment with their working hours.
• A minimum of 5 years of experience in building and operating production data pipelines.
• Proficiency in Python and SQL as primary programming languages.
• Experience in testing, code review, CI/CD, Git, containers, and orchestration.
• Strong expertise in designing and constructing data warehouses or lakehouses.
• Knowledge of dimensional modeling and incremental processing.
• Familiarity with distributed processing at production scale using Spark, Kafka, Flink, or similar technologies.
• Experience with production orchestration tools like Airflow, Dagster, or Prefect.
• Understanding of retries, idempotency, and backfill strategies.
• Ability to manage version-controlled transformations with tests and lineage using dbt or an equivalent tool.
• Experience with cloud deployment, preferably Azure, with AWS being acceptable.
• Proficiency with Docker, CI/CD pipelines, and infrastructure as code using GitHub Actions, Terraform, or Bicep.
• Capability to articulate pipeline cost, runtime, and throughput decisions.
• Active utilization of AI-assisted coding tools such as Claude Code, Cursor, or GitHub Copilot in project delivery.
• Excellent written and spoken English skills at C1 level or above.
• Confidence in explaining technical trade-offs directly to clients.
• A Bachelor's degree in Computer Science, Data Science, or a related field, or equivalent professional experience.
• Paid time off (PTO).
• Observance of U.S. Holidays.
• AI Training opportunities.
• Mentored career development programs.
• Profit-sharing initiatives.
• Competitive remuneration in USD.
CuraLinc Healthcare
VSP Vision Care
Adoreal
Get handpicked remote jobs straight to your inbox weekly.