SAG / ARCHITECTURE NOTE

ZIP Bombs and Path Traversal: Where Is a Data Collector’s Trust Boundary?

A design that keeps processing resources predictable by limiting the permitted size, entries, and paths of compressed inputs. Even a small upload can require substantial memory after decompression. Using filenames directly as paths can also risk affecting files outside the analysis system.

Download Markdown

What Is a Compressed-Input Boundary?

It is a design that keeps processing resources predictable by limiting the permitted size, entries, and paths of compressed inputs. This note treats compressed-input boundaries in terms of the responsibilities of input, transformation, and output rather than as a feature name. To trust analysis results, it must be possible to trace what data was received, what was checked, and how far conclusions can be drawn.

Why Is This Technology Needed?

Even a small upload can require substantial memory after decompression. Using filenames directly as paths can also risk affecting files outside the analysis system.

Design Principles and Data Flow

Limit compressed size and decompressed size separately, and reject abnormal entries, duplicate paths, and unsupported formats. Record input errors and system failures as distinct states.

Compressed input → Allowlist and resource limits → Validated document

Each stage must not reframe the success of the preceding stage as an outcome of its own. Recording the data identifier, period, and validation status together makes it possible to locate omissions and errors and determine what needs to be checked again.

Connection to the SAG Architecture

The SAG ZIP path limits size and reads permitted HTML without extracting it to disk. It does not include execution of customer data or automatic downloads of external assets in its analysis.

SAG’s operational value lies in connecting this relationship to pages and questions, comparison results, and improvement work. Rather than reading numbers alone, customers can review both what needs attention and the basis for the judgment. Patterns requiring further application should be interpreted according to the scope of the relevant paragraph.

Illustrative Example and Evaluation Criteria

If normal HTML and unsupported compressed entries are mixed in illustrative input, the system should not mark everything as successful. It should distinguish which conditions were rejected and provide errors that allow the data to be prepared again.

The example above is for explaining structure and calculations; it does not represent measured results from a specific customer. Actual reports must link the selected period, target, observation conditions, and original records so that the same judgment can be reviewed again.

Practical Validation Checklist

Flow stageWhat to check
Compressed inputCheck decompressed size and number of entries
Allowlist and resource limitsReject absolute paths and parent paths
Validated documentCheck the policy for partial saves after errors

Check that the same meaning is preserved not only for valid input, but also for empty data, duplicate data, and data with different conditions. Connecting validation items to task completion criteria can reduce the gap between the feature description and actual operations.

Limitations and Points to Note in Application

Processing limits help improve safety, but they do not prove that every malicious document is harmless. Allowlists, library updates, and regression testing for errors must be maintained together.

Research and Official Documentation

External materials provide background on the design topic above; they do not certify every SAG implementation or customer outcome. The interpretations and illustrative examples in this note are organized around SAG’s operational structure. Materials checked: 2026-10-06.

Further Reading and Feature Information

How to Continue Reading About This Technology

Follow HTML ZIP and sitemaps → normalization → page versions → evidence records.

Articles