
Senior Python Developer, AI Evaluation & Benchmarking
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in Brazil.
• Design and create coding benchmarks that are utilized to assess advanced AI models.
• Evaluate AI-generated code for accuracy, dependability, efficiency, and exceptional scenarios.
• Develop and sustain scalable data pipelines that facilitate AI evaluation processes.
• Construct structured programming scenarios to assess reasoning, debugging capabilities, and code quality.
• Engage with extensive codebases and multi-language software environments.
• Collaborate with teams dedicated to enhancing the ways AI models comprehend, generate, and assess software.
• Produce clean, maintainable, and thoroughly tested Python code in accordance with software engineering best practices.
• A minimum of 4 years of professional software engineering experience (mandatory).
• Expert-level skills in Python.
• Background in a high-growth technology firm or a prestigious software organization.
• Proficiency in at least one additional programming language such as JavaScript, Go, C++, or a comparable language.
• Familiarity with CI/CD pipelines and automated testing frameworks like pytest, Mocha, or JUnit.
• Strong grasp of software engineering best practices, debugging methodologies, and code quality standards.
• Exceptional analytical and problem-solving abilities.
• Fully remote contract position.
• Weekly payments for work completed and approved during the preceding week.
Creative Chaos
WCG
Get handpicked remote jobs straight to your inbox weekly.