
Senior Data Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Develop and manage PySpark ETL processes
• Assist in ML-driven scoring models that evaluate the quality and validity of contact data
• Create and refine SQL queries for extensive silver/gold lake-table joins
• Collaborate on the architecture of scoring systems, including rule-based gating and the integration of ML
• Write unit and integration tests for data transformations and model inference processes
• Operate within a bronze/silver/gold medallion lake architecture
• Identify potential for AI and automated workflows in data gathering, identity resolution, and profile aggregation
• Examine data quality issues and assess the outcomes of scoring rules
• Take ownership of architecture and projects from analysis and planning to implementation, testing, and production
• Track data lineage and troubleshoot execution across distributed AWS services
• Record findings related to data flow and document architectural decisions
• Over 5 years of professional experience in data engineering
• Comprehensive project ownership from design to production
• Capability to make and justify architectural decisions with minimal supervision
• Proficient in Python, including experience with PySpark DataFrames
• Extensive hands-on experience with Apache Spark/PySpark, covering partitioning, shuffles, skew, broadcast joins, predicate/partition pruning, caching, and physical query plan analysis
• Advanced SQL skills: multi-way joins, window functions, CTEs, and aggregations at a billion-row scale
• Familiarity with AWS Glue, EMR, S3, Step Functions, Lambda, EventBridge, and CloudWatch
• Experience in writing automated data-pipeline tests using pytest or similar frameworks
• Ability to trace data lineage throughout bronze/silver/gold pipeline stages
• Experience in orchestration with Step Functions, Airflow, or equivalent DAG-style workflows
• Proficient in modeling and deduplicating diverse upstream data into canonical schemas
• Skill in debugging distributed AWS services using CloudWatch logs
• Conduct exploratory data analysis at scale
• Understanding of noisy ground-truth proxies and their implications for score quality
• Knowledge of classification statistics, including precision/recall, error-cost tradeoffs, and calibration
• Experience in backtesting and historical validation
• Conduct root-cause analysis of data quality issues
• Ability to convert vague objectives into measurable criteria
• Discipline in data validation encompassing row counts, cardinality, output differences, data profiling, and silent failure modes
• Experience with automated data-quality verification
• Must possess authorization to work in the U.S.
• Preferred: experience in classical ML model development, scikit-learn, contact data quality, identity resolution, marketing/sales enrichment, CI/CD, EMR, blended rule/ML scoring systems, and LLM/agentic workflows
• Engagement with artificial intelligence and advanced technology
• Commitment to equal opportunity employment
• Visa sponsorship is not provided
Progress Partners
FYUL
CarringtonCrisp
Get handpicked remote jobs straight to your inbox weekly.