
AI Benchmark Engineer, Native Language Specialist – French
Posted Sep 8

Posted Sep 8
This is a fully remote position, open to applicants in Canada.
• Design, construct, and validate Terminal-Bench tasks for multilingual software challenges.
• Assess coding agents.
• Develop realistic task environments utilizing datasets and files in French.
• Generate prompts and translations in the native language to pinpoint AI failure points.
• Create robust reference implementations.
• Compose reliable, deterministic verifier scripts.
• Examine execution logs and adjust task difficulty ranging from Easy to Very Hard.
• Execute standard Terminal-Bench configurations against the Haiku, Sonnet, and Opus model tiers.
• Engage in a four-layer human quality-control process: creation, human review, calibration review, and audit.
• Collaborate with automated LLM-based checks to guarantee fairness, grammatical precision, and benchmark integrity.
• Over 5 years of industry experience in software engineering.
• Demonstrated success at leading technology firms and/or graduation from prestigious engineering universities.
• Native or near-native proficiency in French, with a thorough understanding of grammar, register, and phrasing rules.
• High proficiency in English.
• Strong expertise in Python.
• Strong skills in standard shell scripting.
• Strong capabilities in data processing.
• Extensive experience with Terminal/CLI-based development workflows.
• Familiarity with coding agents.
• In-depth technical knowledge of multilingual text processing challenges.
• Understanding of encoding/decoding robustness and Unicode normalization.
• Awareness of locale-dependent conventions, including collation, casing, and non-Gregorian dates.
• Knowledge of text I/O, toolchain interoperability, and safe string operations.
• For specific languages, familiarity with bidirectional/RTL handling, font fallbacks, and rendering/typography in UI or artifacts.
• Reliable availability and commitment to timely delivery.
• Updated CV in English.
• Successful completion of a GenAI assessment.
• Flexible schedule with no fixed hours, check-ins, or micromanagement.
• Competitive compensation rates.
• Prompt payment.
• Access to a variety of innovative projects.
• Opportunities for portfolio and skills development.
• Global community of linguists, subject matter experts, and language professionals.
• No health insurance, paid time off, or retirement contributions are offered.
• Hours are not guaranteed.
• Most tasks require a minimum of 2 hours per day or 10 hours per week.
Mercor
Mercor
Triple Whale 🐳
Get handpicked remote jobs straight to your inbox weekly.