
Senior Python Developer, AI Evaluation, Benchmarking
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in Texas.
• Design and create coding benchmarks that are utilized to assess cutting-edge AI models.
• Examine AI-generated code for accuracy, dependability, efficiency, and edge cases.
• Develop and sustain scalable data pipelines that facilitate AI evaluation processes.
• Construct structured programming scenarios to evaluate reasoning, debugging, and code quality.
• Manage extensive codebases and software environments that use multiple programming languages.
• Collaborate with teams dedicated to enhancing the way AI models comprehend, generate, and assess software.
• Produce clean, maintainable, and thoroughly tested Python code in accordance with software engineering best practices.
• A minimum of 4 years of professional software engineering experience (mandatory).
• Advanced proficiency in Python.
• Background in a fast-growing technology company or a leading software organization.
• Proficient in at least one additional programming language such as JavaScript, Go, C++, or similar.
• Familiarity with CI/CD pipelines and automated testing frameworks like pytest, Mocha, or JUnit.
• Thorough understanding of software engineering best practices, debugging techniques, and code quality standards.
• Exceptional analytical and problem-solving abilities.
• Fully remote contract opportunity.
• Weekly payments for approved work completed in the preceding week.
• Work volume may vary throughout the duration of the engagement.
Get handpicked remote jobs straight to your inbox weekly.