Remotery

Data Scientist

atScienceLogicRemoteUS flagUnited StatesFull-timeData ScientistMid-levelSenior$140k – $165k/year

Posted Jul 29

This is a fully remote position, open to applicants in United States.

📋 Description

• Design and manage evaluation harnesses for LLM and agentic outputs, including golden sets, regression suites, and rubric-based scoring.

• Create and calibrate LLM-as-judge pipelines; validate judges against human annotations while controlling for bias and variance.

• Establish and monitor response-quality metrics such as faithfulness/groundedness, hallucination rate, answer relevance and completeness, instruction adherence, and persona compliance.

• Curate, version, and enhance evaluation datasets as the product and its interfaces progress.

• Benchmark models within the suite against one another to determine which model is best suited for specific tasks and quantify the quality cost associated with using smaller, local models compared to larger alternatives.

• Conduct red-team assessments of the system: prompt injection, jailbreaks, tool misuse, and edge-case identification.

• Develop chaos and stress tests that assess model and agent reliability under adverse or challenging conditions.

• Identify failure modes and integrate findings into guardrails and regression coverage.

• Evaluate retrieval quality across the document corpus, including recall@k, MRR/nDCG, context precision, and recall; run experiments on chunking, indexing, and hybrid retrieval methods.

• Analyze multi-step agent trajectories: tool-call accuracy, trajectory efficiency, replayable state inspection, and guardrail breach behavior.

• Evaluate intent classification and routing quality as measurable elements rather than black boxes.

• Establish ongoing evaluations that capture quality and behavioral regressions when a model in the suite is replaced, upgraded, or re-quantized, or when prompts and pipelines are modified.

• Monitor output distribution and quality drift in production; differentiate between actual regressions and noise in stochastic outputs.

• Suggest and validate corrections using the available levers with local models—prompt adjustments, retrieval and grounding modifications, routing changes, or model selection.

• Develop, deploy, and manage production models that forecast and highlight trends based on operational telemetry, including capacity/resource forecasting, anomaly detection, and early-warning signals on metrics and logs.

• Transition these models from prototype to production and ensure they remain effective through deployment, monitoring, recalibration, and retraining as data and behaviors evolve.

• Define operationally relevant accuracy and lead-time metrics—precision/recall on predicted incidents, forecasting error, and the advance notice of signals—not just offline scores.

• Integrate predictive signals into the LLM and agentic layer to ensure that forecasts and trends inform reasoning, advisories, and recommendations for operators.


⛳️ Requirements

• Bachelor’s or Master’s degree in Data Science, Computer Science, Statistics, Mathematics, or a related field, or equivalent experience.

• Over 3 years of experience in data science, machine learning, or applied quantitative analysis.

• Strong foundation in applied statistics, with the ability to design robust experiments and significance tests on noisy, non-deterministic outputs (beyond just clean A/B testing).

• Experience in building, deploying, and monitoring predictive or time-series models in production settings, including forecasting, anomaly detection, or trend analysis, with recalibration as data changes.

• Proven experience in evaluating, analyzing, or enhancing LLM or NLP systems, including evaluation design, quality measurement, retrieval assessment, or agent analysis.

• Proficiency in Python programming.

• Strong SQL skills and comfort in querying extensive analytical datasets.

• Familiarity with foundational models and practical experience with modern LLM evaluation and tooling layers, including evaluation/harness frameworks, judge pipelines, and libraries used for serving, prompting, and testing models.

• Capability to create analysis and visualizations through coding.


🏝️ Benefits

• Comprehensive medical, dental, and vision insurance plans.

• 401(k) plan with employer matching.

• Flexible Paid Time Off (FTO) for recharging when needed.

• Volunteer Time Off (VTO) - take two days off each calendar year to volunteer with your chosen charitable organization.

• 5-year Service Milestone Sabbatical.

• Paid parental leave.

• Generous employee referral bonus program.

• Pet insurance available.

• Headquarters office conveniently located in Reston Town Center, featuring a well-stocked kitchen with rotating snacks and beverages, plus catered lunch every Thursday.

• Regular virtual company-wide events, including cooking classes, yoga, meditation, and more.

• Opportunity to learn and grow alongside some of the industry's best and brightest professionals!

People also viewed

Valtech7 hours ago

Senior Data Scientist

AR flagArgentina OnlyFull-timeData Scientist
ApplyView job
Paramount7 hours ago

Senior Data Scientist

US flagNew York OnlyFull-timeData Scientist$124k – $186k/year
ApplyView job
DMS International7 hours ago

Data Scientist – Machine Learning Engineer Intern

US flagMaryland OnlyPart-timeData Scientist
ApplyView job
DocPlanner7 hours ago

Global CS Analytics Lead

ES flagSpain OnlyFull-timeData Scientist
ApplyView job
Oscilar8 hours ago

Senior Data Scientist

CA flagCanada OnlyFull-timeData Scientist$175k – $250k/year
ApplyView job
The Groove9 hours ago

Manager, Data Services

US flagUnited States OnlyFull-timeData Scientist
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers