
Senior Software Engineer β Open Source, SWE-Bench Evaluation
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in Argentina.
β’ Analyze machine learning challenges related to experiment design, model selection, datasets, metrics, preprocessing, distribution shifts, contamination, label noise, feature leakage, hyperparameter tuning, and train/validation/test methodologies, as well as reproducibility and statistical significance.
β’ Assess whether the challenges are technically valid, reproducible, appropriately challenging, and require robust machine learning reasoning.
β’ Evaluate if datasets contain significant and learnable signals.
β’ Identify unintended shortcuts or artifacts present in synthetic datasets.
β’ Ascertain whether tasks necessitate a genuine diagnosis of underlying machine learning issues instead of relying solely on brute-force model selection or extensive hyperparameter searches.
β’ Review the evaluation metrics and improvement thresholds.
β’ Detect instances of metric manipulation, data leakage, and evaluation flaws.
β’ Confirm reproducibility throughout the entire data-to-model-to-evaluation pipeline.
β’ Evaluate whether the difficulty of challenges is appropriately calibrated.
β’ Offer recommendations for enhancing, recalibrating, or omitting problematic tasks.
β’ Analyze machine learning experiments, datasets, metrics, and pipelines related to applied machine learning model training and evaluation challenges.
β’ Over 3 years of practical experience in applied machine learning.
β’ Extensive experience in ML experiment design, model selection, hyperparameter tuning, model evaluation, data preprocessing, and validation.
β’ Deep understanding of train, validation, and test data splits.
β’ Proficiency in identifying data leakage, label noise, distribution shifts, spurious correlations, feature leakage, and data contamination.
β’ Experience in evaluating whether performance enhancements are statistically significant rather than mere random variations.
β’ Strong grasp of machine learning evaluation metrics and the appropriate contexts for their application.
β’ Proven experience in debugging machine learning workloads in both CPU and GPU environments.
β’ Ability to analyze technical issues and provide clear, written feedback.
β’ Experience in participating or contributing to ML competitions such as Kaggle or DrivenData is a plus.
β’ Experience in designing benchmark datasets or ML challenges is advantageous.
β’ Background in data-centric AI or dataset quality is a plus.
β’ Familiarity with synthetic data generation and validation is beneficial.
β’ Knowledge of statistical testing, confidence intervals, and effect sizes is a plus.
β’ Experience with ML evaluation pipelines, RLHF, or AI model evaluation is desirable.
β’ Experience in developing ML curricula or technical assessments is a bonus.
β’ Understanding concepts like shortcut learning, spurious correlations, Goodhartβs Law, Simpsonβs paradox, and metric manipulation is advantageous.
β’ Applicants are required to choose a programming language or library for the interview and provide a location.
β’ Opportunity for remote work.
β’ Part-time, project-based consulting engagement.
Kryptos Technologies UK Limited
CBAM-Estimator GmbH
NewRocket
T-Rex Solutions, LLC
Get handpicked remote jobs straight to your inbox weekly.