
Staff Software Engineer, Data Products
Posted 49 min ago

Posted 49 min ago
This is a fully remote position, open to applicants in United States.
• Design, develop, and sustain reusable feature datasets that facilitate machine learning applications, including personalization, engagement, risk assessment, churn prediction, recommendation systems, and experimentation.
• Create self-service frameworks that enhance and democratize dataset creation across the data organization.
• Collaborate with Data Scientists to convert modeling requirements into production-ready feature pipelines, supporting the complete model lifecycle from exploration through deployment.
• Identify source data, necessary transformations, and historical timeframes required for feature engineering. Assist in defining and constructing shared, reusable feature definitions across models instead of creating one-off datasets.
• Balance the freshness, accuracy, latency, and computational efficiency of features when designing data pipelines.
• Construct datasets that facilitate both historical model training and future production inference.
• Design and implement batch and streaming data pipelines that convert raw healthcare, behavioral, product, and operational data into reliable ML-ready datasets.
• Develop dependable data processing systems utilizing Python, SQL, Spark, and contemporary cloud data platforms.
• Optimize large-scale distributed processing for enhanced performance, scalability, and cost-efficiency.
• Create data pipelines that are modular, testable, observable, and adaptable as product requirements evolve.
• Ensure data quality through testing, anomaly detection, schema validation, and pipeline monitoring.
• Collaborate with platform teams to support near real-time feature generation when applicable.
• Enhance reproducibility by standardizing feature computation across experimentation and production environments.
• Facilitate rapid experimentation while maintaining long-term maintainability.
• Establish engineering standards for accuracy, documentation, and maintainability.
• Familiarity with feature stores or feature management platforms.
• Understanding of model training pipelines and MLOps workflows.
• 8+ years of experience in building large-scale production data platforms and distributed data pipelines.
• Proven experience in designing reusable datasets that support machine learning, experimentation, or advanced analytics.
• Demonstrated capability in closely partnering with Data Scientists to productionize feature engineering workflows.
• Experience leading technical initiatives across teams and influencing engineering direction.
• Strong track record of working with cloud-native data platforms such as AWS.
• Experience in building production data systems using Databricks, Iceberg, Spark, Redshift, Snowflake, or similar technologies.
• Proficient in developing reliable batch and streaming data pipelines.
• Experience with healthcare, behavioral, or other large-scale event data is a plus.
• Competitive salary with a generous annual cash bonus.
• Equity grants.
• Remote-first work-from-home culture.
• Flexible Time Off to help you rest, recharge, and connect with loved ones.
• Generous parental leave.
• Health, dental, and vision insurance (with above-market employer contributions).
• 401k retirement savings plan.
• Lifestyle Spending Account (LSA).
• Mental Health Support Solutions.
• ...and more!
BPO Global Services S.A.S
GSB Solutions
Get handpicked remote jobs straight to your inbox weekly.