AI Benchmark Engineer, Native Language Specialist – Arabic

Posted 2 days ago

This is a fully remote position, open to applicants in United Arab Emirates (UAE).

📋 Description

• Design, construct, and validate Terminal-Bench tasks for multilingual software challenges.

• Assess coding agents.

• Create realistic task environments utilizing datasets and files in Arabic, ensuring assets are in the target language.

• Identify AI failure points through prompting and translation efforts in Arabic.

• Develop strong reference implementations.

• Write dependable, deterministic verifier scripts.

• Analyze execution logs and adjust task difficulty from Easy to Very Hard.

• Execute standard Terminal-Bench configurations against Haiku, Sonnet, and Opus model tiers.

• Engage in a four-layer human quality control process: creation, human review, calibration review, and audit.

• Collaborate with automated LLM-based checks to ensure fairness, grammatical accuracy, and benchmark integrity.

• Provide support for multilingual AI and human-verified language services for enterprises, governments, and AI developers.


⛳️ Requirements

• Over 5 years of industry experience in software engineering.

• Demonstrated success at leading technology firms and/or graduation from prestigious engineering universities.

• Native or near-native proficiency in Arabic with a thorough understanding of grammar, register, and phrasing rules.

• High proficiency in English.

• Strong skills in Python, standard shell scripting, and data processing.

• Extensive experience with Terminal/CLI-based development workflows.

• Familiarity with coding agents.

• Deep technical knowledge of multilingual text-processing challenges.

• Experience with encoding/decoding robustness and Unicode normalization.

• Understanding of locale-dependent conventions, including collation, casing, and non-Gregorian dates.

• Experience with text I/O, toolchain interoperability, and safe string operations.

• For Arabic: expertise in bidirectional/RTL handling, font fallbacks, and rendering/typography in UI or artifacts.

• Reliable availability and commitment to timely delivery.

• Most tasks require availability of at least 2 hours per day or 15–20 hours per week.

• Ability to provide an updated CV in English.

• Must complete a GenAI assessment.

• Contractors must not operate in regions subject to international embargoes or sanctions.

• Responsible for personal tax obligations.


🏝️ Benefits

• Flexible schedule; work when it suits you, as much or as little as you prefer.

• No fixed hours, check-ins, or micromanagement.

• Competitive compensation rates.

• Timely payments.

• Access to a variety of innovative projects.

• Opportunities for portfolio enhancement and professional skill development across different industries and domains.

• Global network of linguists, subject matter experts, and language professionals.

• Work remotely from any location, at any time you desire.

• No health insurance, paid time off, or retirement benefits are provided.

• Hours are not guaranteed.

People also viewed

Mercor5 hours ago

AI Safety Red Teamer

US flagUnited States OnlyFreelanceArtificial Intelligence$70 – $84/hour
ApplyView job
Mercor6 hours ago

AI Safety Expert – English, Telugu

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor6 hours ago

AI Safety Experts, English, Punjabi

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor15 hours ago

AI Safety Expert, English, Gujarati

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor15 hours ago

AI Safety Experts – English, Punjabi

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor15 hours ago

AI Safety Expert – English, Gujarati

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers