Work location - Hyderabad
About the Role
We are looking for an AI QA Engineer with hands-on experience in testing Generative AI, LLM-based applications, AI chatbots, and RAG systems.
This role is not suitable for candidates with only manual testing or generic Selenium automation experience. The ideal candidate should have practical exposure to AI response validation, hallucination detection, prompt testing, grounding verification, RAG retrieval validation, and LLM evaluation frameworks.
Key Responsibilities
- Test AI chatbots, LLM-powered applications, and GenAI workflows across functional, regression, API, and end-to-end scenarios.
- Validate RAG retrieval results, including context relevance, retrieval accuracy, source grounding, and response completeness.
- Evaluate AI-generated responses for hallucinations, factual accuracy, answer relevancy, consistency, and faithfulness.
- Verify grounding and citations by checking whether model responses are supported by retrieved source documents.
- Perform prompt validation, prompt robustness testing, jailbreak testing, and prompt injection attack testing.
- Validate model responses against expected outcomes, golden datasets, business rules, and acceptance criteria.
- Build or execute AI evaluation suites using tools/frameworks such as RAGAS, DeepEval, LangChain evaluation, LangSmith, prompt evaluation frameworks, or custom Python-based metrics.
- Test multi-turn chatbot conversations for context retention, fallback handling, safety behaviour, and response consistency.
- Collaborate with product, engineering, data science, and QA teams to define AI testing strategy and quality metrics.
- Log AI defects clearly with prompt, context, source data, actual response, expected response, screenshots/logs, and evaluation reason.
- Support automation using Python, Playwright, Selenium, REST API testing, Postman, GitHub Actions, Jenkins, or similar tools.
Must-Have Skills
- 5+ years of QA / SDET / automation testing experience
- Hands-on experience in AI testing, GenAI testing, LLM testing, chatbot testing, or RAG testing
- Strong understanding of:
- Hallucination detection
- Grounding validation
- Prompt testing
- Prompt injection / jailbreak testing
- RAG retrieval validation
- AI model response evaluation
- Experience validating model outputs against expected outcomes, source documents, golden datasets, or evaluation metrics
- Working knowledge of Python, JavaScript, or TypeScript
- Experience with automation tools such as Playwright, Selenium, PyTest, or REST Assured
- API testing experience using Postman / REST APIs
- Familiarity with CI/CD pipelines such as GitHub Actions, Jenkins, GitLab CI, or Azure DevOps
- Experience with Jira, Agile/Scrum, test planning, defect reporting, and regression testing
Good-to-Have Skills
- Experience with RAGAS, DeepEval, LangSmith, Langfuse, LlamaIndex, LangChain, or LLM-as-a-Judge
- Exposure to OpenAI, Azure OpenAI, Claude, Gemini, or other LLM APIs
- Knowledge of vector databases such as FAISS, Pinecone, ChromaDB, or Weaviate
- Understanding of embeddings, chunking, retrieval, context precision/recall, and faithfulness metrics
- Experience testing agentic AI workflows, AI agents, MCP-based workflows, or multi-agent systems
- Performance testing exposure for AI systems, including latency, TTFC, streaming consistency, and response quality.