
Quality Engineering Consultant – AI/ML Testing
Posted Sep 2

Posted Sep 2
This is a fully remote position, open to applicants in United States.
• Develop and implement testing strategies for AI and GenAI applications, including solutions powered by LLMs, AI agents, copilots, and conversational systems.
• Assess AI outputs for precision, relevance, groundedness, consistency, detection of hallucinations, and compliance with toxicity and safety standards.
• Create and manage prompt test suites and benchmark datasets.
• Conduct functional, regression, performance, and reliability testing for AI-driven applications.
• Define and implement AI evaluation metrics and quality gates.
• Build and maintain Playwright test automation frameworks.
• Automate UI, API, and comprehensive business workflows.
• Develop reusable automation components and testing accelerators.
• Integrate automated tests into CI/CD pipelines.
• Promote shift-left testing and quality engineering methodologies.
• Design and implement API test cases utilizing Postman, REST Assured, and Playwright API Testing.
• Verify API contracts, authentication, authorization, error handling, and performance metrics.
• Execute integration testing across distributed systems and third-party services.
• Establish AI evaluation frameworks and quality scorecards.
• Measure model performance through precision/recall, relevancy scoring, semantic similarity, groundedness validation, and human-in-the-loop evaluation.
• Analyze trends in AI quality and suggest enhancements.
• Support Responsible AI initiatives and model governance standards.
• Collaborate with engineering, product, and AI teams to proactively identify quality risks.
• Engage in requirement reviews and design discussions.
• Report testing progress, risks, and quality metrics to stakeholders.
• Contribute to quality engineering best practices, standards, and reusable playbooks.
• 4 - 6 years of experience in Quality Engineering, Software Testing, or Test Automation.
• Hands-on experience in testing AI and Generative AI applications.
• Proficiency in modern web application testing and API testing.
• Strong expertise in Playwright automation.
• Practical experience with API testing and automation.
• Experience validating AI/LLM-based applications and AI agents.
• Solid understanding of AI evaluation methodologies and testing approaches.
• Familiarity with M365 Copilot, OpenAI/Azure OpenAI, Claude, Gemini, and Agentic AI frameworks.
• Knowledge of test management tools like Azure DevOps (ADO), Jira, or equivalent platforms.
• Experience with CI/CD pipelines and DevOps practices.
• Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field.
• Remote work arrangement.
Shield AI
CI&T
Symphonic Distribution
Techdinamics Integrations Inc.
Get handpicked remote jobs straight to your inbox weekly.