
Senior Software Test Engineer
Posted Sep 12

Posted Sep 12
This is a fully remote position, open to applicants in Texas.
• Construct and sustain the evaluation platform, which encompasses harnesses, datasets, graders, and CI integration.
• Assess agent performance through metrics such as task success, tool usage accuracy, drift, latency, and cost.
• Conduct hands-on exploratory end-to-end testing reflecting how members, operations teams, and providers engage with the product.
• Collaborate with business stakeholders to comprehend genuine workflows and confirm that the appropriate product has been developed.
• Cultivate domain knowledge in Curative’s operational workflows.
• Utilize traces, failure taxonomies, and dashboards to convert agent challenges into actionable insights.
• Replicate production failures, determine root causes, remedy issues, or provide detailed diagnostics.
• Offer guidance on quality concerns and make decisions regarding product behavior or requirements when necessary.
• Prioritize testing and quality efforts based on risk and business impact.
• Evaluate AI-generated code modifications and uphold quality standards.
• Over 5 years of experience in software quality, testing, or engineering across a broad spectrum.
• Proficient in writing automation scripts and conducting thorough exploratory testing.
• Expertise in LLM evaluation, including golden datasets, LLM-as-judge, programmatic graders, and regression suites for prompts and agent loops.
• Hands-on experience with AI agents and their failure modes, such as silent drift, tool misuse, and compounding errors.
• In-depth debugging skills with logs, traces, queries, and production forensics.
• Capability to resolve bugs in Python or TypeScript.
• Ability to engage with non-technical business stakeholders and operations leaders.
• Comfort with uncertainty, quick-paced environments, and working without established processes or sprint boundaries.
• Practical quality judgment.
• Excellent written communication skills.
• Proficient in using Claude Code, Cursor, or similar as a primary development tool.
• Employs AI to enhance quality efforts while acknowledging the limitations of AI judgment.
• Reviews every AI-generated difference and refrains from merging based solely on instinct.
• Highly preferred: experience establishing evaluation or observability infrastructure from the ground up.
• Highly preferred: familiarity with LLM observability tools such as LangSmith, Braintrust, Arize, or custom-built solutions.
• Highly preferred: experience defining unsupervised agent capabilities and setting up guardrails.
• Highly preferred: background in product requirements and behavior definition.
• Highly preferred: developing substantial domain knowledge in a complex operational business.
• Curative Health Plan (100% employer-covered medical premiums for you and 50% coverage for dependents on the base plan.)
• $0 copays and $0 deductibles (with completion of our Baseline Visit).
• Preventive and primary care included.
• Mental health support available.
• One-on-one care navigation provided.
• Chronic condition programs (diabetes, weight, hypertension) offered.
• Maternity and family planning assistance available.
• 24/7/365 Curative Telehealth services.
• Pharmacy benefits included.
• Comprehensive dental and vision coverage offered.
• Employer-provided life and disability coverage with additional supplemental options.
• Flexible spending accounts available.
• Generous PTO policy along with 11 paid annual company holidays.
• 401K plan for full-time employees.
• Generous paid parental leave of up to 8–12 weeks, based on role eligibility.
IndieKidz GmbH
CloudMargin
Quanata
CARE
Get handpicked remote jobs straight to your inbox weekly.