SAG / ARCHITECTURE NOTE
RAGAS and RAG Evaluation: Why Retrieval and Answers Must Be Measured Separately
Explains the definition and need for RAG evaluation, how it works, criteria for applying it to SAG architecture, and a practical checklist, drawing on research and official documentation.
Definition in one sentence
RAG evaluation is a method for diagnosing a RAG pipeline along dimensions such as faithfulness, answer relevance, and context relevance when ground-truth labels are scarce.
Key point: A single score for the final answer cannot tell you whether retrieval failed or generation ignored the evidence. The evaluation dimensions must be separated to identify which layer to improve.
Why is this technology needed?
A single score for the final answer cannot tell you whether retrieval failed or generation ignored the evidence. The evaluation dimensions must be separated to identify which layer to improve.
How it works
It evaluates faithfulness to evidence and relevance based on the question, retrieved context, and answer, and compares them against a human-created reference set where possible. Samples are also reviewed for bias in automated evaluation models.
When designing the system, do not look only at accuracy. Latency, cost, data boundaries, refresh cycles, and behavior on failure must also be defined to make results reproducible in production. It is safer to leave values that automation cannot determine with confidence as unmeasured or requiring review, rather than changing them to zero or success.
Connection to SAG technology
SAG preserves unmeasured states instead of combining rule-based results with actual AI observations. When connecting RAG evaluation, the evaluation model, prompt, and dataset versions should also be recorded in provenance.
Practical checklist
- Separate retrieval and generation metrics
- Have people review samples from automated evaluations
- Pin the evaluation model and prompt versions
- Distinguish the states for failures, empty results, and permission errors from success
- Revalidate before and after changes under the same conditions
Research and official documentation
Reference documents provide evidence for the principles and recommendations. They do not guarantee search visibility, AI mentions, rankings, or revenue; the actual effects of implementation must be verified through observations of service data under the same conditions.
Selection criteria and a concrete application example
If relevant documents were found but the answer added specifications not present in those documents, retrieval relevance and faithfulness to evidence would produce different results. An automated evaluator is not the same as a human judgment of the correct answer. Metrics should be calibrated using a fixed set of questions and error cases verified by people.
Scope of application at SAG
This article covers research principles and extended designs for search AI. Read it in connection with SAG's page collection, evidence recording, and report verification structure, but do not interpret it to mean that all the paper's search algorithms are deployed in the production pipeline. Whether they are applied should be verified using the search module, evaluation data, and execution records.
How to continue reading about this technology
Compare RAG, GraphRAG, and Self-RAG papers and their conditions for application.
SAG / KNOWLEDGE LINKS
