
Quality Assurance Engineer
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in India.
β’ Develop and sustain automated testing for AI voice and chat agents, encompassing everything from individual conversational exchanges to comprehensive roleplay scenarios.
β’ Assess the quality of LLM outputs - including accuracy, consistency, structured responses, language fidelity, and regression of prompts - utilizing evaluation frameworks (LLM-as-judge, golden datasets, and tolerance-based assertions for non-deterministic outputs).
β’ Verify RAG pipelines: ensuring relevance in retrieval, grounding/faithfulness, and overall answer quality.
β’ Evaluate voice pipelines: checking STT/TTS accuracy and real-time, low-latency performance across multiple languages.
β’ Implement automation across UI/E2E, API, and AI output layers, while constructing CI/CD processes from the ground up - including linting, type-checking, testing, evaluations, coverage gates, and deployment verifications.
β’ Take responsibility for release quality: develop regression strategies, detect issues early, and make clear go/no-go decisions.
β’ Conduct database and load/performance testing; enhance automation capabilities for both desktop and mobile applications.
β’ Assume control of and expand existing QA automation efforts, while reinforcing shared frameworks, tools, and QA processes/standards.
β’ Over 5 years of experience as a QA Automation Engineer, demonstrating a successful history of testing AI systems, beyond just traditional software.
β’ A builder mindset: have created testing frameworks, standards, and CI/CD-integrated automation from the ground up.
β’ Practical experience in testing conversational AI - either voice or chat agents - on actual, delivered projects.
β’ Strong understanding of LLMs, prompt engineering, and RAG - with the ability to design tests for non-deterministic outputs and assess retrieval/generation quality.
β’ Knowledge of voice pipelines (speech-to-text and text-to-speech) and methodologies for testing them through automation.
β’ Hands-on experience with LLM evaluation and observability tools, such as Langfuse, LangSmith, DeepEval, and RAGAS.
β’ Proficient in UI/E2E testing (e.g., using Playwright) and API test automation; experienced in constructing CI/CD pipelines from scratch.
β’ Comfortable with database testing, load/performance assessments, and automation for desktop and mobile testing environments.
β’ Strong discipline with Git, foundational cloud knowledge, and familiarity with the agile sprint process.
β’ Proven success in high-paced, fast-delivery settings - quickly adapts to unfamiliar systems and maintains a practical approach to process.
β’ Flexible work arrangements
Jumio Corporation
Moniepoint Inc. (Formerly TeamApt Inc.)
IUNA AI
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.