
Data Engineer, Python, PySpark, Databricks
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in Colombia, +1 more country.
• Design, develop, and maintain scalable data pipelines utilizing Python, Apache Spark, PySpark, and Databricks.
• Modernize and transition legacy data processing tasks to secure, cloud-native environments.
• Construct and oversee batch data ingestion pipelines from both structured and unstructured data sources.
• Integrate data from REST APIs, SharePoint, document repositories, enterprise applications, and cloud services.
• Implement data quality measures, monitoring, observability, and operational controls.
• Enhance data workloads for improved performance, scalability, reliability, and cost-effectiveness.
• Develop pipelines for document extraction, classification, metadata enrichment, and automation.
• Employ software engineering methodologies including Git, CI/CD, automated testing, and code reviews.
• Construct, deploy, troubleshoot, and manage production ETL/ELT pipelines.
• Work collaboratively with architects, developers, tax subject matter experts, and platform teams.
• Assess existing codebases to pinpoint opportunities for enhancing maintainability, security, and reliability.
• Utilize AI-assisted development tools to create, review, test, and maintain data pipeline code.
• Over 3 years of experience in data engineering.
• Proficient in Python development.
• Extensive experience with Apache Spark and PySpark.
• Practical knowledge of Databricks.
• Strong SQL and data modeling capabilities.
• Experience in building, testing, deploying, and troubleshooting production ETL/ELT pipelines.
• Familiarity with Git-based source control and CI/CD processes.
• Experience in consuming REST APIs for data ingestion, including aspects like authentication, pagination, and error handling.
• Background in writing automated tests for data pipelines.
• Comprehensive understanding of data quality, reliability, and data engineering best practices.
• Excellent problem-solving, communication, and collaboration abilities.
• Experience with Azure cloud services and the modernization of legacy data pipelines.
• Exposure to document processing, intelligent document extraction, or handling unstructured data.
• Familiarity with AI-assisted development tools like Codex, GitHub Copilot, or similar technologies.
• Experience with streaming data pipelines or large-scale data repositories.
• Capability to analyze unfamiliar codebases, identify gaps and trade-offs, and suggest practical enhancements.
• A stable, long-term contract with opportunities for professional growth.
• Private health insurance coverage.
• A remote-friendly culture that encourages work-life balance.
• Ongoing training, mentorship, and learning programs.
• Complimentary access to AI training resources and cutting-edge AI tools.
• Flexible Paid Time Off (PTO) policy along with paid holidays.
• Engaging, world-class software projects for clients in the US and Latin America.
• Collaboration with skilled software engineers across Latin America and the US.
• A diverse working environment.
Händlerbund
ASRC Federal
SPD Technology
Get handpicked remote jobs straight to your inbox weekly.