Remotery

Software Engineer – Benchmarking

Posted 3 days ago

This is a fully remote position, open to applicants in California, +1 more state.

📋 Description

• Prepare and manage benchmark datasets, which includes cleaning, preparation, conversion into executable formats, validation, and ongoing upkeep.

• Develop and sustain infrastructure and evaluation pipelines for model APIs and terminal agents.

• Create efficient, containerized environments and viewers for tasking and tool usage evaluation.

• Assist in the fine-tuning of small open-source LLMs and evaluate baseline performance against post-training results.

• Generate scoreboards and leaderboards that display performance metrics at the model, benchmark, task, domain, and rubric levels.

• Create analytical tools to identify failure modes, compare models and agent frameworks, and monitor capability evolution over time.

• Work collaboratively with researchers and engineers to ensure that evaluation data and outputs are precise, consistent, and incorporated into published research.

• Develop HTML viewers and internal tools for task creation, review, quality assurance, and structured data gathering.


⛳️ Requirements

• A minimum of 4 years of professional experience in building and maintaining complex systems.

• Proficiency in Python programming.

• Capable of writing robust, maintainable code and working extensively within existing codebases and infrastructure.

• Experience in preparing, cleaning, and maintaining datasets.

• Familiarity with Docker and reproducible execution environments.

• Ability to collaborate with researchers and scientists, translating methodologies into functional systems.

• Hands-on experience with AI evaluation or familiarity with Harbor, Terminal-Bench, or Inspect is highly advantageous.

• Knowledge of Python, model APIs, agent/evaluation frameworks, and custom evaluation tools.

• Experience with terminal agents and open-source models within the Hugging Face ecosystem and PyTorch.

• Proficiency with Docker.

• Familiarity with React, Next.js, and Tailwind.

• Experience using GitHub, Slack, Notion, and Linear.

• Experience in fine-tuning or post-training of open-source LLMs, or other hands-on machine learning activities, is a bonus.

• Background in agentic, multi-turn, long-context, or tool-use evaluation is a plus.

• Experience validating LLM-as-judge or rubric-based grading systems is an added advantage.

• A background or strong interest in a scientific or technical field is a bonus.

• Experience in building data-heavy dashboards, leaderboards, or visualizations is beneficial.

• Contributions to open-source projects or published work related to benchmarks and measurements are desirable.


🏝️ Benefits

• Competitive salary and equity.

• Medical, dental, and vision insurance coverage.

• 401(k) plan.

• Monthly wellness and fitness stipend.

• Paid time off policy, along with company holidays.

• Annual company off-sites (Tahoe, Mendocino, Mexico City, San Diego, Park City).

• Family-friendly policies.

• Flexible remote work options.

• Paid family leave.

People also viewed

Sigma Software Group10 hours ago

Principal Software Engineer

UA flagUkraine OnlyFull-timeFull-stack Engineer
ApplyView job
Collectly10 hours ago

Senior Software Engineer – Pre Service Team

RS flagSerbia OnlyFull-timeFull-stack Engineer
ApplyView job
ioet10 hours ago

Senior Product Engineer

Latin AmericaFull-timeFull-stack Engineer
ApplyView job
Allata10 hours ago

FullStack Developer, Java, Vue

US flagTexas OnlyFull-timeFull-stack Engineer
ApplyView job
hatch I.T.10 hours ago

Senior Full Stack Developer

US flagDistrict of Columbia, +1 more stateFull-timeFull-stack Engineer
ApplyView job
CVS Health10 hours ago

Senior Manager – Epic EHR Development Lead, Software Development

US flagMassachusetts OnlyFull-timeFull-stack Engineer$130.3k – $260.6k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers