
Software Engineer – Benchmarking
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in California, +1 more state.
• Prepare and manage benchmark datasets, which includes cleaning, preparation, conversion into executable formats, validation, and ongoing upkeep.
• Develop and sustain infrastructure and evaluation pipelines for model APIs and terminal agents.
• Create efficient, containerized environments and viewers for tasking and tool usage evaluation.
• Assist in the fine-tuning of small open-source LLMs and evaluate baseline performance against post-training results.
• Generate scoreboards and leaderboards that display performance metrics at the model, benchmark, task, domain, and rubric levels.
• Create analytical tools to identify failure modes, compare models and agent frameworks, and monitor capability evolution over time.
• Work collaboratively with researchers and engineers to ensure that evaluation data and outputs are precise, consistent, and incorporated into published research.
• Develop HTML viewers and internal tools for task creation, review, quality assurance, and structured data gathering.
• A minimum of 4 years of professional experience in building and maintaining complex systems.
• Proficiency in Python programming.
• Capable of writing robust, maintainable code and working extensively within existing codebases and infrastructure.
• Experience in preparing, cleaning, and maintaining datasets.
• Familiarity with Docker and reproducible execution environments.
• Ability to collaborate with researchers and scientists, translating methodologies into functional systems.
• Hands-on experience with AI evaluation or familiarity with Harbor, Terminal-Bench, or Inspect is highly advantageous.
• Knowledge of Python, model APIs, agent/evaluation frameworks, and custom evaluation tools.
• Experience with terminal agents and open-source models within the Hugging Face ecosystem and PyTorch.
• Proficiency with Docker.
• Familiarity with React, Next.js, and Tailwind.
• Experience using GitHub, Slack, Notion, and Linear.
• Experience in fine-tuning or post-training of open-source LLMs, or other hands-on machine learning activities, is a bonus.
• Background in agentic, multi-turn, long-context, or tool-use evaluation is a plus.
• Experience validating LLM-as-judge or rubric-based grading systems is an added advantage.
• A background or strong interest in a scientific or technical field is a bonus.
• Experience in building data-heavy dashboards, leaderboards, or visualizations is beneficial.
• Contributions to open-source projects or published work related to benchmarks and measurements are desirable.
• Competitive salary and equity.
• Medical, dental, and vision insurance coverage.
• 401(k) plan.
• Monthly wellness and fitness stipend.
• Paid time off policy, along with company holidays.
• Annual company off-sites (Tahoe, Mendocino, Mexico City, San Diego, Park City).
• Family-friendly policies.
• Flexible remote work options.
• Paid family leave.
Sigma Software Group
Collectly
Allata
Get handpicked remote jobs straight to your inbox weekly.