
Senior Data Engineer
Posted Sep 8

Posted Sep 8
This is a fully remote position, open to applicants in Colombia.
• Design, construct, and enhance PySpark ETL/ELT pipelines utilizing Amazon EMR and AWS Glue.
• Process nationwide, county-partitioned property data on both daily and monthly schedules.
• Operate and advance the Bronze → Silver → Gold data lake on S3 using Apache Hudi and AWS Glue Data Catalog.
• Oversee schema contracts, partitioning, compaction, and performance optimization.
• Develop orchestration using AWS Step Functions, EventBridge, and Lambda.
• Implement reruns, backfills, failure recovery procedures, and maintain documentation.
• Ensure data quality through contracts, write-audit-publish gating, quarantine processes, drift monitoring, and Slack notifications.
• Manage workloads in Aurora PostgreSQL and DynamoDB.
• Deploy data-serving APIs and asynchronous workers to meet production SLAs, complemented with dashboards and runbooks.
• Monitor, report on, and work to decrease AWS data-platform expenses.
• Define infrastructure using Terraform and CloudFormation, integrated with GitHub Actions CI/CD.
• Create feature pipelines, training datasets, and serving paths for machine learning scoring and valuation models.
• Maintain runbooks, architectural documentation, and Confluence data dictionaries.
• Make architectural decisions and collaborate directly with the Data Science team.
• Minimum of 4 years of practical data engineering experience with large-scale distributed data systems and a proven record of production ownership.
• Proficient in advanced PySpark, including performance tuning, partitioning strategies, and cost-sensitive cluster sizing on EMR or similar platforms.
• Strong proficiency in Python, producing clean, tested, production-quality code.
• Extensive AWS expertise with EMR, Glue, Lambda, S3, Athena, Step Functions, EventBridge, IAM, and VPC networking.
• Advanced SQL skills, including the ability to perform complex analytical queries, query optimization, and data modeling on Athena/Presto and PostgreSQL.
• Production experience with at least one open table format; Apache Hudi is highly preferred, with Iceberg or Delta Lake also valued.
• Demonstrated experience in implementing data validation, quality gates, monitoring, and incident response for production data.
• Familiarity with Terraform and/or CloudFormation within a CI/CD workflow.
• Professional working proficiency in both English and Spanish (B2+ level).
• Bachelor’s degree in Computer Science, Systems Engineering, Data Engineering, or equivalent practical experience.
• Experience in real estate, property, or geospatial data is advantageous.
• Experience in building or managing public/internal data APIs is a plus.
• Familiarity with observability tooling is beneficial.
• Experience in ML-adjacent engineering is an asset.
• Proficiency in modern Python tools, Agile/SCRUM methodologies, and documentation practices is a plus.
• Performance Share Bonus.
• Flex PTO (up to 26 days per year).
• Home Office Upgrade Bonus.
• HMO Bonus.
• Full-time Remote Work.
• Impact Moments.
• Opportunities for growth and team-building.
• Ongoing support and budget for skill development.
CuraLinc Healthcare
VSP Vision Care
Adoreal
Get handpicked remote jobs straight to your inbox weekly.