
Senior Data Engineer
Posted Jul 23

Posted Jul 23
This is a fully remote position, open to applicants in United States.
• Design, construct, and take ownership of scalable data pipelines and dimensional models utilizing Databricks (PySpark, SQL, medallion architecture) — ensuring timely delivery of reliable data products while adhering to SLAs within your designated scope.
• Ingest data from operational and SaaS sources, such as Salesforce, into the lakehouse, prioritizing managed connectors like Lakeflow Connect when applicable.
• Create and sustain Kimball-style dimensional models — including facts, conformed dimensions, and slowly changing dimensions — serving as the analytics layer of record.
• Develop, test, and document transformations in dbt (models, sources, snapshots, tests, exposures) while maintaining a strong CI discipline.
• Oversee data assets in Unity Catalog, managing catalogs, schemas, permissions, and lineage.
• Enhance performance and cost-effectiveness through cluster and warehouse sizing, Spark tuning, partitioning, and tagging for cost attribution.
• Operationalize machine learning workflows using MLflow for experiment tracking, model registry, and deployment, adhering to MLOps best practices.
• Assist in coordinating daily operations with offshore vendor engineering teams — establishing priorities, sequencing deliverables, and ensuring alignment with sprint commitments and the platform roadmap.
• Convert business and technical requirements into precise specifications, acceptance criteria, and design guidance that offshore teams can follow with minimal ambiguity.
• Conduct quality assurance on offshore deliverables through code reviews, testing, and validation against data standards, performance benchmarks, and definition-of-done before changes are promoted to production.
• Collaborate with data architects, analysts, and business stakeholders to maintain data accuracy and governance, fostering alignment within the team and with immediate cross-functional partners on delivery.
• Maintain engineering standards, code review practices, and documentation conventions across both onshore and offshore contributors.
• Contribute to the team's development by training and mentoring engineers on tools, standards, and best practices as the platform evolves.
• Over 7 years of experience in data engineering on big data and cloud platforms, including more than 3 years of hands-on experience with Databricks (Spark/PySpark, Delta Lake, jobs).
• Demonstrated success in delivering Kimball / dimensional data models within a modern warehouse or lakehouse, with proficiency in SQL and Python (PySpark).
• Production experience with dbt (models, tests, snapshots) and with Unity Catalog for governance, access control, and lineage management.
• Familiarity with the ML lifecycle and MLOps, including using MLflow for experiment tracking, model registry, and deployment.
• Proven experience in translating business and technical requirements into clear specifications and coordinating or supervising offshore and vendor engineering resources, including evaluating their deliverables for quality.
• Excellent communication and stakeholder management skills, with a history of mentoring engineers and establishing technical standards.
• Risepoint is an equal-opportunity employer and advocates for a diverse and inclusive workforce.
Railroad19
GFT Technologies
Get handpicked remote jobs straight to your inbox weekly.