
Senior/Staff AI Model Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Take responsibility for the reliability and quality standards of an AI copilot integrated into a trading platform utilized by sophisticated investors.
• Create and implement evaluation systems that assess correctness, safety, latency, and regression risk across market analysis, portfolio/risk reasoning, and trading workflows, including order placement.
• Design and uphold benchmarks that consist of curated golden sets, scenario suites, stress/adversarial cases, and updated market/regime-based test corpora.
• Establish automated quality gates and regression workflows that prevent releases when critical metrics decline.
• Collaborate with engineering and product teams to define safe tool/action contracts, ensuring deterministic previews, confirmations, and audit trails.
• Manage model improvement loops linked to evaluations, incorporating data collection/labeling strategies, error taxonomy, prompt/tooling modifications, and relevant fine-tuning or preference optimization.
• Create and oversee AI monitoring and incident response systems, including telemetry, alerting, root-cause analysis, and fix-forward methodologies.
• Acquire a comprehensive understanding of trading concepts such as margin, shorting, portfolio margin, risk, and execution, and communicate them accurately to users.
• Engage with the technology stack, including Rust, TypeScript, Postgres, React, React Native, observability/telemetry tools, LLM APIs, model serving, and evaluation/training pipelines.
• A minimum of Eight (8) years of experience in delivering production software.
• Strong proficiency in any programming language.
• Solid understanding of computer science principles, testing methodologies, and systems design.
• Experience in developing evaluation frameworks, test harnesses, and benchmark suites for intricate systems (LLMs/agents/search/retrieval/ranking/recommenders).
• Proven track record of managing model improvement cycles, including dataset curation, labeling/QA, offline experimentation, and implementing changes that yield measurable impacts on benchmarks.
• Capability to define metrics, construct measurement pipelines, and influence engineering/product decisions based on data.
• Comfort in working across the stack, debugging model/tooling failures, instrumenting services, and collaborating with frontend/product teams on UX patterns that enhance safety and trust.
• High level of self-motivation and readiness to tackle unfamiliar challenges to resolve issues.
• Bonus: Experience in fine-tuning, preference optimization, distillation, or prompt/compiler-style techniques to enhance tool-use reliability.
• Bonus: Experience in developing domain-specific benchmarks and adversarial suites for high-stakes applications.
• Bonus: Extensive experience in trading across various asset classes, margin types, etc.
• Bonus: Experience with Rust and performance-sensitive services.
• Bonus: Expertise in designing incident response strategies and SLOs for ML/AI systems.
• Company equity.
• 401k matching.
• Gender-neutral parental leave.
• Comprehensive medical, dental, and vision insurance.
• Lunch stipends.
• Fully stocked kitchens.
• Happy hours.
• Prime location.
• Stunning views.
• Equal opportunity workplace.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.