
Senior Software Engineer – ML Data Delivery
Posted 2 hours ago

Posted 2 hours ago
This is a fully remote position, open to applicants in Michigan.
• Design, develop, test, and deploy tools and pipelines for internal quality assurance and identification of issues in pseudo-labeled data.
• Offer statistical support while reporting on the quality, coverage, and health of pseudo-labels and pipelines to internal stakeholders and leadership.
• Assist in secondary data selection (virtual packaging) to curate and provide on-demand data for online model training.
• Contribute to the development of the ML data delivery system for on-demand online perception model training.
• Enhance and optimize pipelines throughout the pseudo-labeling and data delivery framework.
• Act as project lead, guiding less experienced team members through project execution.
• Stay updated on advancements in offline perception, data pipeline engineering, and ML data infrastructure related to autonomous driving.
• Create tools, services, and algorithms utilizing structured software development processes, version control, and documentation.
• Define and execute data ingestion, preparation, curation, and governance for large datasets that support analytics and ML training workflows.
• Evaluate existing capabilities, identify areas for improvement, and propose strategic solutions aligned with operations.
• Generate informational products that support data visualization and accessibility.
• Assess technical advancements to boost productivity and quality, minimize flow times, and enhance operational reliability.
• Establish guidelines and standards for data quality control, data delivery systems, deployment, and related processes.
• Offer technical guidance, leadership, coaching, and mentoring to team members.
• Bachelor’s Degree in Computer Science, Robotics, Electrical Engineering, or a related technical field, along with competencies typically gained through 6+ years of experience, OR a Master’s Degree in a related technical field with competencies typically acquired through 3+ years of experience.
• Highly skilled and proficient in the discipline; capable of conducting complex and significant work with minimal supervision and independent judgment.
• Ability to drive alignment across team interfaces, own technical solutions, achieve consensus, and mentor fellow engineers.
• Familiarity with the offline perception stack and best practices for pseudo-label data production.
• Strong software engineering background with experience in building and operating data pipelines and services at scale.
• Proficient in statistical analysis and reporting, translating data quality findings into actionable insights.
• Experience with scaled MLOps and tooling, ML frameworks, experiment tracking, model registries, MLflow, Weights and Biases, and ML metrics/evaluation/quality.
• Experience in model data curation and Parquet data processing using tools such as PyArrow, Daft, Pandas, or similar.
• Proficiency in Python software development.
• Familiarity with VDI and cloud-based development environments, CI systems like GitHub Actions, and Docker.
• Experience with distributed data processing and/or ML frameworks such as PyTorch, Lightning, Ray, or similar.
• Experience with large-scale data delivery systems and associated quality control is advantageous.
• Pseudo-labeling experience is a plus.
• A competitive compensation package that includes a bonus component and stock options.
• 100% paid medical, dental, and vision premiums for full-time employees.
• 401K plan featuring a 6% employer match.
• Flexible scheduling options.
• Generous paid vacation available immediately upon starting.
• Company-wide holiday office closures.
• AD+D and Life Insurance.
Torc Robotics
Workiva
StackAdapt
Get handpicked remote jobs straight to your inbox weekly.