
Applied ML Engineer
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Singapore, +3 more countries.
• Replicate and assess research methodologies utilizing open-weight and API-accessible models.
• Develop evaluation datasets, probes, scoring techniques, baselines, calibration assessments, and experimental harnesses.
• Engage with model weights, logits, hidden states, activations, model APIs, and inference infrastructure.
• Construct and enhance evaluation infrastructure, which includes runners, judges, persistence, experiment orchestration, and reporting.
• Transform research workflows into product experiences, comprising experiment configurations, runs, traces, comparisons, reports, and review processes.
• Explore verification techniques under fine-tuning, merging, quantization, distillation, safety removal, and intentional evasion.
• Design controlled experiments that differentiate significant signals from artifacts or confounding factors.
• Author technical reports that clarify measured evidence, interpretations, and hypotheses.
• Deliver production-quality systems equipped with APIs, background jobs, observability, testing, and documentation.
• Reproduce a published model-provenance or verification technique within a six-month timeframe.
• Create a repeatable model-verification runner that includes versioned inputs, artifacts, metrics, and reports.
• Integrate a verification workflow into Construct and ensure its accessibility through the Eldros UI.
• Conduct controlled experiments involving base, fine-tuned, merged, quantized, and distilled models.
• Enhance the understanding of when verification methods are effective, ineffective, and the reasons behind their performance.
• Proficient Python engineering skills with hands-on experience in PyTorch and Hugging Face Transformers.
• Solid comprehension of ML evaluation, encompassing dataset design, baselines, metrics, calibration, false positives, false negatives, statistical uncertainty, and reproducibility.
• Capability to read ML research papers and implement methodologies from foundational principles.
• Experience in developing production software beyond notebook environments, including APIs, asynchronous jobs, databases, logging, testing, and deployment.
• Comfort with open-weight models and an understanding of contemporary LLM inference systems.
• Ability to navigate both backend and frontend domains; experience working with React/TypeScript product interfaces.
• Strong technical judgment regarding experimental evidence.
• High level of initiative and a strong sense of ownership.
• Ability to thrive in a dynamic startup atmosphere.
• Valuable experience in model provenance, fingerprinting, watermarking, distillation detection, red-teaming, safety evaluations, interpretability, activation and representation analysis, DSPy, LiteLLM, Temporal, Ray, vLLM, PostgreSQL/pgvector, Next.js, React, TypeScript, data visualization, experiment dashboards, GPU model serving, and adversarial evaluations.
• Competitive salary and equity options.
• Flexible work hours and remote work opportunities.
• Comprehensive health, dental, and vision insurance.
• Opportunities for professional development and continuous learning.
• A collaborative and innovative work environment.
Shield AI
Weekday (YC W21)
Roadpass Digital
Get handpicked remote jobs straight to your inbox weekly.