
Lead Data Engineer β PySpark, Palantir Foundry
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in Washington.
β’ Provide exceptional client value and maintain high levels of client satisfaction.
β’ Design, improve, and sustain production-quality data pipelines that facilitate model output aggregation and downstream risk analysis.
β’ Develop and enhance governance practices for repositories, encompassing branching strategies, pull request standards, merge policies, release tagging, and version control workflows.
β’ Manage release engineering practices to ensure reproducible, traceable, and auditable production releases.
β’ Refactor and enhance pipeline code to improve modularity, maintainability, scalability, and documentation quality.
β’ Create configuration-driven pipeline patterns that span across various environments and releases.
β’ Assist in testing, validation, benchmarking, and change management for essential data pipelines.
β’ Collaborate with data scientists, machine learning engineers, data engineers, product stakeholders, and other technical teams.
β’ Ensure alignment among teams regarding schemas, interfaces, inputs, and delivery expectations.
β’ Convey complex technical concepts into straightforward updates for both technical and non-technical stakeholders.
β’ Enhance structure and governance in codebases, repositories, and engineering workflows.
β’ Contribute to engineering best practices within a regulated, audit-sensitive delivery environment.
β’ 10-15+ years of experience in data engineering, data science, machine learning engineering, or related fields, utilizing Python.
β’ Proven experience in leading technical teams and managing enterprise-scale data initiatives.
β’ Strong proficiency in PySpark, SQL, and cloud services.
β’ Capability to enhance, refactor, or stabilize existing codebases and pipeline environments.
β’ Familiarity with cloud-optimized datasets, efficient partitioning strategies, and large-scale spatial operations.
β’ Understanding of machine learning model outputs that integrate into downstream data pipelines, platforms, or production systems.
β’ Experience in highly regulated sectors such as utilities, financial services, healthcare, insurance, or similar fields.
β’ Track record of designing maintainable, scalable, and well-documented cloud-based data infrastructure or modern data platform environments.
β’ Expertise in supporting reproducibility, dataset versioning, release traceability, and audit readiness.
β’ Ability to define expected inputs, outputs, schemas, and interfaces across technical teams.
β’ Excellent communication skills.
β’ Detail-oriented and governance-focused approach to engineering.
β’ Hands-on experience with Git-based workflows, code reviews, branching strategies, and release management.
β’ Familiarity with Palantir Foundry is highly preferred.
β’ Experience with GIS technologies and geospatial data platforms.
β’ Competitive base salary.
β’ Performance-based bonuses.
β’ Additional incentives.
β’ Opportunities for training and mentorship.
β’ Project opportunities aimed at career development.
β’ A supportive, globally connected work environment.
QAVION GROUP
ShippyPro
Blood Cancer United
RevoData
Get handpicked remote jobs straight to your inbox weekly.