SAG / ARCHITECTURE NOTE

Hybrid Search: Why Combine BM25 and Dense Retrieval

Explains what hybrid search is, why it is needed, how it works, and how to apply it within SAG architecture, with practical checklists and support from research and official documentation.

Download Markdown

Definition in one sentence

Hybrid search combines sparse retrieval based on keyword matching with dense retrieval based on semantic similarity.

Key point: Exact matches matter for product names, codes, and legal terms, while semantic similarity matters for natural-language questions. Using only one approach can miss other kinds of relevance.

Why is this technology needed?

Exact matches matter for product names, codes, and legal terms, while semantic similarity matters for natural-language questions. Using only one approach can miss other kinds of relevance.

How it works

The candidates and scores from the two search systems are combined after score normalization or through rank fusion, and then a reranker determines the final order. Their contributions should be evaluated for each dataset.

Design should account for more than accuracy. Latency, cost, data boundaries, update cycles, and behavior in the event of failure must also be defined to produce reproducible results in operation. It is safer to leave values that automation cannot determine confidently as unmeasured or requiring review, rather than changing them to 0 or marking them as successful.

Connection to SAG technology

When connecting a search layer to SAG’s question-and-evidence structure, a suitable design is to use sparse search to complement unique brand names and dense search to complement context-based questions. Current fixture results should be distinguished from results produced by future search engines.

Practical checklist

  • Collect failure cases for exact-match search and semantic search
  • Evaluate candidate recall separately from final precision
  • Set combination weights using validation data
  • Distinguish the states for failures, empty results, and permission errors from success
  • Revalidate before and after changes under the same conditions

Research and official documentation

Reference documents provide evidence for principles and recommendations. They do not guarantee search visibility, AI mentions, rankings, or revenue. The actual impact of an implementation must be verified through observations using service data under the same conditions.

Selection criteria and a concrete application example

For example, an exact-match search could find product code AX-210, while a semantic search could find candidates for “a product that can be used for a long time while traveling on business.” Because the two scores use different scales, rank fusion or validated normalization is needed rather than simply adding them as-is. The candidate-merging stage and the reranking stage are also separate.

Scope of application at SAG

This article covers the research principles of search AI and approaches to extending it. Read it in connection with SAG’s page collection, evidence recording, and report validation structure, but do not interpret it to mean that all the search algorithms in the papers have been deployed in the production pipeline. Whether they have been applied should be verified through the search module, evaluation data, and execution records.

How to explore related technologies

Compare the papers on RAG, GraphRAG, and Self-RAG and the conditions for applying them.

Articles