@maschinenlesbar.org/openka-cli
    Preparing search index...

    Interface CorpusDiskUsage

    interface CorpusDiskUsage {
        blobs: DiskUsage;
        by_source: Record<string, DiskUsage>;
        index: DiskUsage;
        records: DiskUsage;
    }
    Index
    blobs: DiskUsage

    The archived documents, blobs/.

    by_source: Record<string, DiskUsage>

    Per source, the archived documents its syncs brought in — read from its validators (SourceState.http_cache), so a document two sources share counts for both. Keys sorted. What ka sync --dry-run averages for its estimate.

    index: DiskUsage

    The catalog, the inverted index and any embeddings, index/.

    records: DiskUsage

    The canonical records, records/.