AI Benchmark Engineer, Native Language Specialist – Portuguese (Brazil)

atLILT AIRemoteBR flagBrazilFreelanceArtificial IntelligenceJunior$50 – $75/hour

Posted Sep 9

This is a fully remote position, open to applicants in Brazil.

📋 Description

• Design, construct, and validate Terminal-Bench tasks that assess coding agents.

• Create realistic task environments utilizing datasets and files in Portuguese, ensuring that assets remain in the target language.

• Identify AI failure points through prompting and translation in Portuguese.

• Develop strong reference implementations.

• Write dependable, deterministic verifier scripts, employing rubric-based evaluation only when absolutely necessary.

• Examine execution logs and adjust task difficulty from Easy to Very Hard.

• Execute standard Terminal-Bench configurations against Haiku, Sonnet, and Opus model tiers.

• Engage in a four-layer human quality-control process that includes creation, human review, calibration review, and audit.

• Collaborate with automated LLM-based checks to guarantee fairness, grammatical correctness, and benchmark integrity.


⛳️ Requirements

• A minimum of 1 year of industry experience in software or prompt engineering.

• Demonstrated success at leading technology firms and/or graduation from prestigious engineering universities.

• Native or near-native fluency in Portuguese (Brazil) with a profound understanding of grammar, register, and phrasing rules.

• High proficiency in English.

• Strong expertise in Python, standard shell scripting, and data processing.

• Significant experience with Terminal/CLI-based development workflows.

• Familiarity with coding agents.

• In-depth technical knowledge of multilingual text processing challenges.

• Experience with encoding/decoding robustness and Unicode normalization.

• Understanding of locale-dependent conventions, including collation, casing, and non-Gregorian dates.

• Knowledge of text I/O, toolchain interoperability, and safe string operations.

• An updated CV in English.

• Completion of a GenAI assessment is required.

• Reliable availability and commitment; most tasks necessitate at least 2 hours per day or 10 hours per week.

• Contractors must not be engaged in regions subject to international embargoes or sanctions.

• Contractors are accountable for their own tax obligations.


🏝️ Benefits

• Flexible schedule without fixed hours, check-ins, or micromanagement.

• Competitive rates with prompt payments.

• Access to a variety of innovative projects that enhance your portfolio and develop your skills.

• A global community of linguists, subject matter experts, and language professionals.

• No benefits such as health insurance, paid time off, or retirement contributions are provided.

• Work availability may vary based on project demand, and hours are not guaranteed.

People also viewed

LGC1 day ago

AI Delivery Lead

HU flagHungary OnlyFull-timeArtificial Intelligence
ApplyView job
Mercor1 day ago

AI Safety Red Teamer

US flagUnited States OnlyFreelanceArtificial Intelligence$70 – $84/hour
ApplyView job
Mercor1 day ago

AI Safety Red Teamer

US flagUnited States OnlyFreelanceArtificial Intelligence$70 – $84/hour
ApplyView job
Triple Whale 🐳1 day ago

AI Marketing Strategist

US flagUnited States OnlyFull-timeArtificial Intelligence
ApplyView job
The Cigna Group1 day ago

Data Measurement & Reporting, Digital Data & AI

US flagNew Jersey OnlyFull-timeArtificial Intelligence$96.7k – $161.1k/year
ApplyView job
U.S. Digital Response1 day ago

Chief Technology, AI Officer

US flagUnited States OnlyFull-timeArtificial Intelligence$250k – $277.5k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers