Per source, the archived documents its syncs brought in — read from its
validators (SourceState.http_cache), so a document two sources share counts
for both. Keys sorted. What ka sync --dry-run averages for its estimate.
The catalog, the inverted index and any embeddings, index/.
The canonical records, records/.
The archived documents,
blobs/.