
Senior Data Engineer
Posted Aug 19

Posted Aug 19
This is a fully remote position, open to applicants in United States.
• Develop and sustain PySpark ETL processes.
• Assist in machine learning-based scoring models that predict the quality and validity of contact data.
• Create and enhance SQL queries for extensive silver/gold lake-table joins.
• Collaborate on the architecture of scoring systems, incorporating rule-based gating and machine learning positioning.
• Write unit and integration tests for data transformations and model inference.
• Operate within a bronze/silver/gold medallion lake framework.
• Identify potential applications for AI and agentic workflows in data acquisition, identity resolution, and profile aggregation.
• Examine data-quality challenges and evaluate scoring-rule results.
• Take full ownership of architecture and projects from analysis and planning to implementation, testing, and production.
• Trace data lineage and troubleshoot execution across distributed AWS services.
• Document findings related to data flow and architectural decisions.
• Over 5 years of experience in professional data engineering.
• Comprehensive project ownership from design to production.
• Capability to make and justify architectural decisions with minimal supervision.
• Proficient in Python, particularly with PySpark DataFrames.
• Extensive hands-on experience with Apache Spark/PySpark, including partitioning, shuffles, skew, broadcast joins, predicate/partition pruning, caching, and physical query plan analysis.
• Proficient in complex SQL, including multi-way joins, window functions, CTEs, and aggregations at the billion-row scale.
• Familiarity with AWS Glue, EMR, S3, Step Functions, Lambda, EventBridge, and CloudWatch.
• Experience in writing automated tests for data pipelines using pytest or similar tools.
• Ability to trace data lineage across bronze/silver/gold pipeline stages.
• Experience with orchestration tools such as Step Functions, Airflow, or other similar DAG-style workflows.
• Knowledge of modeling and deduplicating diverse upstream data into canonical schemas.
• Proficient in debugging distributed AWS services using CloudWatch logs.
• Conduct exploratory data analysis at scale.
• Understanding of noisy ground-truth proxies and their implications for score quality.
• Familiarity with classification statistics, including precision/recall, error-cost tradeoffs, and calibration.
• Experience with backtesting and historical validation.
• Conduct root-cause analysis of data quality issues.
• Ability to convert ambiguous objectives into testable criteria.
• Maintain a discipline of data validation involving row counts, cardinality, output differences, data profiling, and silent failure modes.
• Experience with automated data quality checks.
• Authorization to work in the U.S. is required.
• Preferred: experience in classical ML model development, scikit-learn, contact data quality, identity resolution, marketing/sales enrichment, CI/CD, EMR, blended rule/ML scoring systems, and LLM/agentic workflows.
• Work with artificial intelligence and state-of-the-art technology.
• Equal opportunity employer.
• Visa sponsorship is not available.
Providence
Gorilla Logic
Archera
Kard
Get handpicked remote jobs straight to your inbox weekly.