
Software Engineer β Infrastructure & Platform
Posted Aug 18

Posted Aug 18
This is a fully remote position, open to applicants in United States.
β’ Design and construct sandboxed evaluation environments that enable AI models to execute code securely, utilize tools, engage with services, and perform intricate tasks.
β’ Develop backend services and infrastructure that facilitate large-scale, repeatable evaluations of AI and agentic systems.
β’ Create agent scaffolding and evaluation harnesses, incorporating tool-use loops, context management, retries, state management, token budgets, and workflows involving multiple agents or subagents.
β’ Build systems for provisioning and orchestrating isolated environments utilizing Docker, Kubernetes, virtual machines, and cloud infrastructure.
β’ Devise secure strategies for networking, permissions, secrets, credentials, and resource isolation within model-driven environments.
β’ Create APIs, internal tools, and automation that empower researchers, engineers, and subject-matter experts to conduct evaluations effectively.
β’ Enhance evaluation reliability and reproducibility through logging, observability, snapshotting, debugging tools, and automated testing.
β’ Develop systems capable of executing thousands of evaluation tasks consistently while capturing artifacts and telemetry necessary for understanding model behavior.
β’ Collaborate with analysts, red team members, and domain experts to convert complex evaluation concepts into robust technical frameworks.
β’ Examine failures throughout the evaluation stack to differentiate between model limitations and failures in infrastructure, harnesses, or environments.
β’ 3β5+ years of professional software engineering experience, especially in backend development, infrastructure, platform, SRE, or distributed systems engineering.
β’ Proficient programming skills in Python and experience in developing production-quality software.
β’ Experience in designing and managing backend services, APIs, or distributed systems.
β’ Practical experience with Docker, Kubernetes, virtual machines, or other container/orchestration technologies.
β’ Familiarity with AWS, GCP, or comparable cloud infrastructure.
β’ Strong comprehension of Linux systems, networking, authentication, permissions, and infrastructure security.
β’ Experience with infrastructure-as-code or automation tools like Terraform.
β’ Excellent debugging skills across application, infrastructure, and networking layers, particularly in agentic loops.
β’ Capability to construct systems that are reproducible, observable, scalable, and secure.
β’ Comfort in addressing ambiguous technical challenges where architecture and requirements may rapidly change.
β’ Interest in AI systems, agentic workflows, AI security, or model evaluations; previous professional AI experience is advantageous but not mandatory.
β’ Performance-based annual bonus.
β’ Support for conferences, continuing education, or leadership development.
β’ Fully remote work environment.
β’ Comprehensive health, dental, and vision insurance.
β’ Generous paid time off and holiday schedule.
β’ 401(k) plan.
Second Nature
Headway
SYNCREON
Rentokil Pest Control North America
Get handpicked remote jobs straight to your inbox weekly.