
Senior Python Developer, AI Evaluation & Benchmarking
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in Argentina.
• Design and create coding benchmarks utilized for evaluating cutting-edge AI models.
• Examine AI-generated code for accuracy, reliability, efficiency, and edge cases.
• Develop and sustain scalable data pipelines that facilitate AI evaluation workflows.
• Construct structured programming scenarios to assess reasoning, debugging, and code quality.
• Work with substantial codebases and multi-language software environments.
• Collaborate with teams dedicated to enhancing how AI models comprehend, generate, and assess software.
• Write clean, maintainable, and well-tested Python code adhering to software engineering best practices.
• Minimum of 4 years of professional software engineering experience (mandatory).
• Expert-level knowledge of Python.
• Experience in a rapidly growing technology company or a leading software organization.
• Proficiency in at least one additional programming language, such as JavaScript, Go, C++, or similar.
• Familiarity with CI/CD pipelines and automated testing frameworks like pytest, Mocha, or JUnit.
• Strong grasp of software engineering best practices, debugging, and code quality.
• Outstanding analytical and problem-solving abilities.
• Fully remote contract opportunity.
• Weekly payments for approved work completed in the preceding week.
Creative Chaos
WCG
Get handpicked remote jobs straight to your inbox weekly.