---
title: "URL Identity and Deduplication: Does Removing the Hash Make It the Same Page?"
slug: "url-identity-deduplication"
language: "en"
tags: ["url 정체성","아키텍처 노트","sag 기술","수집 아키텍처"]
created: "2026-10-06T08:00:00.000Z"
published: "2026-10-08T10:17:55.048Z"
updated: "2026-10-08T10:18:00.185Z"
sample: false
---

# URL Identity and Deduplication: Does Removing the Hash Make It the Same Page?

## What Is URL Identity?

**It is a way to manage document identity by distinguishing the original link from the address used for comparison.** This note treats URL identity in terms of the responsibilities of inputs, transformations, and outputs, rather than as the name of a feature. To trust an analysis, it must be possible to trace what material was included, what was checked, and how far the conclusions can go.

## Why Is This Technology Needed?

If citations with different fragments are counted as new sources each time, the number of evidence pages is inflated. Conversely, deleting important query parameters can merge different products.

## Design Principles and Data Flow

Set normalization rules according to their purpose. Use hash removal from citation URLs and deduplication for source aggregation, while keeping the original URL traceable.

> **Original URL** → **Normalization by purpose** → **Deduplicated source aggregation**

Each step must not rebrand the success of the previous step as the achievement of the next. Keeping the material's identifier, time period, and verification status linked in the record helps locate omissions and errors and determine what needs to be checked again.

## Connection to the SAG Architecture

SAG citation reports validate safe URLs and do not count repeated instances of the same source more than once based on hash differences. This rule is not extended to removing the meaning of all search URLs.

SAG's operational value lies in connecting this relationship to pages and questions, comparison results, and improvement tasks. Rather than reading only numbers, customers can review both what needs to be improved and the basis for the assessment. Patterns that require additional application should be interpreted according to the scope of the relevant paragraph.

## Illustrative Example and Criteria for Assessment

For illustration, if `product#spec` and `product#service` are both cited in one answer, the number of cited answers for the same page is one. Different URLs such as `product?id=1` and `product?id=2` cannot be merged in the same way.

The example above is provided to explain the structure and calculation; it is not a measured result for a particular customer. In an actual report, the selected period, scope, observation conditions, and original records must be linked so that the same assessment can be checked again.

## Practical Verification Checklist

| Flow stage | What to check |
| --- | --- |
| Original URL | Check for fragment duplicates |
| Normalization by purpose | Preserve meaningful query parameters |
| Deduplicated source aggregation | Whether the original link can be restored |

Check that the same meaning is preserved not only for normal inputs, but also for missing material, duplicate material, and material with different conditions. Linking verification items to completion criteria can reduce the gap between feature descriptions and actual operations.

## Limitations and Points to Consider When Applying

Canonical hints, actual content identity, and normalization for aggregation are separate judgments. Deduplication rules must be managed by version to keep comparison history stable.

## Research and Official Documentation

- [IETF HTTP Semantics RFC 9110](https://www.rfc-editor.org/rfc/rfc9110.html) — A standard for understanding the meaning of HTTP requests, responses, and statuses.

External sources provide background on the design topic above; they do not certify every SAG implementation or customer result. The interpretation of this note's application and its illustrative examples are organized according to SAG's operational structure. Material checked: 2026-10-06.

## Further Reading and Feature Details

- [Related architecture note](/ko/blog/rag-provenance-citation-traceability)
- [Try a service connected to URL identity](/ko/preview/geo?scenario=cream)
- [Feature-specific FAQ](/en/faq)
- [Discuss implementation scope](/ko#inquiry)


## How to Continue Reading About This Technology

Follow the path: HTML ZIP·sitemap → normalization → page version → evidence record.

- [How data becomes evidence](/ko/blog?tag=%EC%88%98%EC%A7%91%20%EC%95%84%ED%82%A4%ED%85%8D%EC%B2%98)
- [Feature guide FAQ](/en/faq)
