
Senior Software Engineer β Open Source, SWE-Bench Evaluation
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in Argentina.
β’ Analyze software engineering assignments derived from actual GitHub issues and pull requests.
β’ Evaluate genuine bug fixes and feature implementations.
β’ Review open-source repositories, pull requests, unit tests, and test coverage.
β’ Assess repository configurations and dependency management.
β’ Analyze reproducibility and environment settings.
β’ Evaluate task complexity, difficulty, and code modifications that span multiple files or modules.
β’ Determine if tasks are well-defined, technically feasible, sufficiently tested, and reflective of professional software engineering challenges.
β’ Identify flaky tests, absent dependencies, unclear requirements, version conflicts, and environment-specific behaviors.
β’ Offer clear recommendations on whether tasks should be accepted, enhanced, or disregarded.
β’ Provide technical insights and quality evaluations.
β’ Minimum of 3 years of professional software engineering experience.
β’ Strong familiarity with extensive, multi-file codebases.
β’ Proven experience in reviewing pull requests, troubleshooting issues, and maintaining production code.
β’ Solid understanding of unit testing and test coverage principles.
β’ Capability to assess if tests adequately validate a solution without unduly limiting implementation strategies.
β’ Experience in dependency management, environment preparation, and reproducibility.
β’ Comprehensive understanding of Git and GitHub development workflows.
β’ Ability to dissect intricate technical issues and provide clear written feedback.
β’ Contributions to or management of open-source projects (preferable).
β’ Familiarity with SWE-Bench, SWE-Bench Verified, or comparable coding benchmarks (preferable).
β’ Experience with significant Python open-source projects such as Django, Flask, scikit-learn, SymPy, matplotlib, requests, or pytest (preferable).
β’ Experience with Docker, CI/CD, pip, conda, or dependency pinning (preferable).
β’ Knowledge of test fixtures, test isolation, or property-based testing (preferable).
β’ Experience in crafting technical assessments or reviewing coding challenges (preferable).
β’ Background in AI/ML evaluation, data curation, RLHF, or benchmark development (preferable).
β’ Compensation of $65 per hour.
β’ Opportunity for remote work.
β’ Part-time, project-based consulting role.
Kryptos Technologies UK Limited
CBAM-Estimator GmbH
NewRocket
T-Rex Solutions, LLC
Get handpicked remote jobs straight to your inbox weekly.