
Software Engineer, II – Autonomy Data
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants in Virginia.
• Contribute to the design and structuring of the autonomy program’s data lake, encompassing schema definitions, partitioning strategies, and metadata indexing.
• Develop and sustain dependable end-to-end pipelines that ingest high-bandwidth vehicle sensor logs into cloud storage.
• Execute data validation and integrity checks to identify corrupted information, missing sensors, and calibration inconsistencies.
• Establish data retention, tiering, and lifecycle management policies.
• Create tools for querying raw logs and generating curated training and evaluation datasets.
• Automate cost-effective pseudo-labeling workflows at scale during ingestion.
• Implement metrics for data quality and model performance to guide labeling efforts.
• Deploy and maintain visualization tools for log reviews, annotation quality assurance, and autonomy debugging.
• Integrate visualization tools with the data lake to link dataset entries or model failures back to source logs.
• Collaborate with autonomy engineers to define visualization panels and metrics.
• Develop dashboards that illustrate data coverage across various terrains, operating environments, and geographic regions.
• Establish and document data agreements between data services and model-training consumers.
• Collaborate with perception, planning, and embedded engineers throughout the data lifecycle.
• Assist in evolving data engineering standards, best practices, and tool selections.
• Contribute to the data roadmap and communicate insights to senior technical leadership.
• Bachelor’s degree in Computer Science, Computer Engineering, Software Engineering, Electrical Engineering, or a related field with 4+ years of data engineering experience, or a Master’s degree with 2+ years.
• Strong expertise in Python and SQL.
• Proficient in building production-quality data pipelines.
• Experience with cloud data infrastructure, preferably AWS S3, Glue, Athena, Redshift, or equivalent.
• Familiarity with infrastructure-as-code tools such as Terraform or CloudFormation.
• Understanding of data partitioning strategies and columnar storage formats like Parquet and ORC.
• Experience in developing and managing pipelines that process time-series and binary data.
• Capability to evaluate and integrate open-source tools.
• Experience in implementing monitoring, validation, data quality, and lineage tracking.
• Only U.S. citizens are eligible for this position.
• Bonus: experience with autonomous vehicles, robotics, or sensor-driven autonomous systems.
• Bonus: extensive experience with Foxglove or Rerun.
• Bonus: familiarity with MCAP CLI or Python library and MCAP-to-columnar conversion.
• Bonus: experience in ML data curation, diversity sampling, pseudo-labeling, and dataset versioning.
• A competitive compensation package that includes both a bonus component and stock options.
• 100% coverage of medical, dental, and vision premiums for full-time employees.
• 401K plan featuring a 6% employer match.
• Flexible scheduling options and generous paid vacation (available immediately upon starting).
• Company-wide holiday office closures.
• AD+D and Life Insurance coverage.
• Potential sign-on bonuses, relocation assistance, and other forms of compensation as part of the total compensation package.
Cloudera
Stellar Cyber
Pragmatike
Pragmatike
Get handpicked remote jobs straight to your inbox weekly.