
QA Engineer, AI Products
Posted Jul 14

Posted Jul 14
This is a fully remote position, open to applicants in United States.
• Develop and implement testing strategies for features powered by LLM, which encompass prompt regression testing, output assessment, and detection of hallucinations.
• Create and sustain automated evaluation pipelines (including evaluation sets, golden datasets, and LLM-as-judge frameworks) to identify quality regressions in non-deterministic outputs.
• Conduct black-box and exploratory testing of MDCalc's AI functionalities on both web and mobile platforms, focusing on clinical accuracy, safety, and edge cases.
• Establish quality metrics for AI outputs (including accuracy, faithfulness, relevance, safety, latency, and cost) and set thresholds for determining release readiness.
• Collaborate across various functions with engineers, product managers, ML/AI engineers, and clinical reviewers to define the criteria for "good" AI responses.
• Investigate and categorize AI failure modes, differentiating between model issues, prompt issues, retrieval issues, and integration bugs.
• Engage in team discussions, providing insights on testability, risks, prompt design, and necessary guardrails.
• Assist in developing QA strategies to enhance future testing capabilities, automation, and evaluation coverage as the AI product landscape evolves.
• A minimum of 5 years of experience in software QA, with at least 1 year focused on testing LLM-based or AI/ML features.
• Comprehensive understanding of QA principles, along with skills in test case creation/documentation and best practices applicable to both deterministic and non-deterministic systems.
• Practical experience with LLM tooling and concepts, including prompt engineering, RAG systems, and evaluation frameworks (such as Promptfoo, Braintrust, LangSmith, DeepEval, Ragas, OpenAI Evals), as well as LLM APIs (like OpenAI, Anthropic, etc.).
• Experience in designing automated qualitative evaluation methodologies, including LLM-as-judge, rubric-based scoring, semantic similarity evaluations, and regression testing with golden datasets.
• Proficient in test automation tools, particularly Playwright.
• Strong SQL capabilities for data validation, test data generation, and ensuring data integrity across various systems.
• Familiarity with token usage, latency profiling, and cost monitoring as indicators of quality.
• A keen desire to learn quickly, combined with a positive and solutions-focused mindset.
• Effective communicator, capable of identifying and articulating issues, blockers, and risks when addressing ambiguous or probabilistic failures.
• Self-driven, proactive, and adept at managing time and priorities autonomously.
• Opportunity to make a significant impact in the field of medicine: MDCalc is the most widely used medical reference by physicians, utilized by over 65% of US attending doctors on a weekly basis.
• Comprehensive Medical, Dental, & Vision Coverage, with options available for dependents.
• Company-sponsored short-term insurance.
• Fully-paid parental leave for 8 weeks, following 6 months of employment.
• Company-sponsored 401k plan after 3 months of employment.
• Unlimited vacation policy for salaried positions - we trust you to take the time you need.
• Bi-annual company offsites to foster connection, reflection, and collaborative planning.
• Monthly stipend for remote work expenses.
• A culture characterized by fun and motivated team members who are committed to a greater mission at MDCalc.
HITCONTRACT
Gainwell Technologies
Fresenius Medical Care
Girls For Girls Africa Mental Health Foundation
Get handpicked remote jobs straight to your inbox weekly.