Skip to content

Module 4 · Evaluation

RAG precision and hallucination analysis

Paste a system prompt, the retrieved context and the expected and actual answers. We score retrieval precision, grounding and data readiness, and flag every claim the context doesn't support.

  • About 2 minutes
  • Hallucination risk
  • 5 quality scores
  • Unsupported claims

Test case

One question, what was retrieved and what the model said.

Don't paste personal data, credentials or confidential information. Use synthetic or anonymized samples. See the privacy notice.

3 chunks detected · separate with a line containing only ---

or press Ctrl + Enter

Your diagnostic will appear here

We've loaded a sample with retrieval noise, a duplicate chunk and one invented claim. Press Analyze to see what we catch.

  • Five scores covering retrieval, grounding and data hygiene
  • Every claim in the answer that the context doesn't support
  • Concrete remediation notes for your pipeline
Methodology and official sources

Metric definitions follow the Ragas and DeepEval documentation linked below. By default this page computes deterministic approximations of them using word overlap, with no LLM judge: context precision is rank-weighted average precision over the retrieved chunks, and faithfulness is the share of answer sentences supported by the context. These are screening signals, not a certified evaluation. When the platform is connected to its DeepEval service, the scores are LLM-judged and the page says so.

The hallucination risk bands (Low ≥ 80% faithfulness, Medium ≥ 50%, High below that) and the data readiness score are this platform's own heuristics. They aren't defined by NIST or OWASP. Those sources describe the underlying risks (confabulation, misinformation).

Sources

  1. Ragas: Context Precision · Ragas
  2. Ragas: Faithfulness · Ragas
  3. DeepEval: Contextual Precision metric · Confident AI (DeepEval)
  4. DeepEval: Contextual Recall metric · Confident AI (DeepEval)
  5. DeepEval: Faithfulness metric · Confident AI (DeepEval)
  6. DeepEval: Answer Relevancy metric · Confident AI (DeepEval)
  7. NIST AI 600-1, §2.2 Confabulation · NIST
  8. LLM09:2025 Misinformation · OWASP GenAI Security Project
↑↓ navigate↵ openEsc close