
Software Engineer L5/L6 β Model Evaluations, Data Curation
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
β’ Create and construct a shared data curation framework, which includes reusable components, libraries, and workflows.
β’ Develop scalable LLM-driven data transformation pipelines utilizing raw sources such as the Netflix catalog and its metadata.
β’ Implement large-scale batch inference while focusing on quality and token cost efficiency.
β’ Formulate sampling strategies that ensure coverage, diversity, difficulty, and balance across various content and member segments.
β’ Create filtering and quality control techniques, which include LLM-as-judge, evaluation-model-based scoring, deduplication, and validation processes.
β’ Collaborate with researchers to evaluate the impact of curation decisions on model performance.
β’ Transform curated datasets into discoverable artifacts with versioning, explicit lineage, and reproducibility.
β’ Promote the adoption of shared curation practices within MEDC and partner modeling teams.
β’ Proficient in software engineering with Python, particularly in building reusable infrastructure, libraries, or frameworks utilized by fellow engineers and researchers.
β’ Experienced in developing LLM-driven data generation or transformation pipelines, including synthetic data, structured outputs, or large-scale batch inference.
β’ Practical experience with data quality methodologies, including sampling strategies, filtering, deduplication, and model-based quality scoring like LLM-as-judge.
β’ Comprehension of how data choices influence model behavior and the capability to design experiments that measure this impact.
β’ Familiarity with distributed data processing tools such as Spark, Ray, or their equivalents.
β’ Exceptional collaborative skills, especially when working with researchers, data scientists, and platform teams.
β’ Experience with LLM evaluation systems is essential for L6 candidates.
β’ Proven technical leadership in data and evaluation infrastructure; ability to establish technical direction for a multi-engineer initiative is crucial for L6 roles.
β’ Knowledge of dataset versioning, lineage, and artifact management practices.
β’ Skill in optimizing cost and throughput for large-scale LLM inference processes.
β’ Experience with human annotation workflows and aligning LLM judges with human raters.
β’ Familiarity with pipeline orchestration frameworks such as Metaflow, Airflow, or similar.
β’ Background in recommendation systems, personalization, search, or working with content catalogs and metadata.
β’ Health Plans
β’ Mental Health support
β’ 401(k) Retirement Plan with employer match
β’ Stock Option Program
β’ Disability Programs
β’ Health Savings and Flexible Spending Accounts
β’ Family-forming benefits
β’ Life and Serious Injury Benefits
β’ Paid leave of absence programs
β’ Full-time hourly employees accumulate 35 days annually for paid time off, applicable for vacation, holidays, and sick leave.
β’ Full-time salaried employees are granted immediate access to flexible time off.
Atomic - Remote Jobs
Affidea
GoFasti
StructureIt
Get handpicked remote jobs straight to your inbox weekly.