SAG / ARCHITECTURE NOTE

ColBERT Late Interaction: Search Design Between Speed and Precision

Explains what ColBERT is and why it is needed, how it works, criteria for applying it to SAG architecture, and a practical checklist, drawing on research and official documentation.

Download Markdown

One-sentence definition

ColBERT is an approach that computes question and document token representations separately, then performs token-level interaction at search time.

Key answer: Compressing an entire document into a single vector can lose fine-grained terms, while jointly encoding every pair is costly. Late interaction offers an option between these two extremes.

Why is this technology needed?

Compressing an entire document into a single vector can lose fine-grained terms, while jointly encoding every pair is costly. Late interaction offers an option between these two extremes.

How it works

Document token vectors are stored in advance, and relevance is calculated by aggregating the scores of the closest document token for each question token. The trade-off between storage space and latency must be measured.

When designing a system, accuracy is not the only consideration. Latency, cost, data boundaries, update frequency, and failure behavior must also be defined to produce reproducible results in production. It is safer to leave values that automation cannot determine with confidence in an unmeasured or needs-review state rather than converting them to zero or treating them as successful.

Connection to SAG technology

It may be a candidate technology for evidence retrieval involving both short key terms and long explanations, such as SAG’s technical and policy documents. Korean-language data and costs must be validated before actual adoption.

Practical checklist

  • Evaluate searches for Korean compound words and English abbreviations
  • Measure index size and response time together
  • Compare against a single-vector approach using the same data
  • Distinguish the states for failures, empty results, and permission errors from success
  • Revalidate before and after changes under the same conditions

Research and official documentation

Reference documents support the principles and recommendations. They do not guarantee search visibility, AI mentions, rankings, or revenue; the effects of actual implementation must be verified through observations using service data under the same conditions.

Selection criteria and a specific application example

Instead of a single document vector, retain vectors for each token. The MaxSim method, which finds the maximum similarity between document tokens and each question token, can preserve fine-grained expressions, but it increases index storage requirements. The paper’s speed comparison reflects results for its dataset and baseline model; it does not guarantee performance for every service.

Scope of application at SAG

This article covers research principles and extensible designs for search AI. Read it in connection with SAG’s page collection, evidence logging, and report verification structure, but do not interpret it to mean that all of the paper’s search algorithms have been integrated into the production pipeline. Whether they have been applied should be verified through the search module, evaluation data, and execution records.

How to continue reading about this technology

Compare RAG, GraphRAG, and Self-RAG papers and their conditions for application.

Articles