OptionalassertThrow StoreError when the blobs cannot be reached at all — a blob directory
on a drive that is not mounted. Optional: a store whose blobs are always there
(the in-memory double) has nothing to check.
Run work with catalog writes deferred, and persist the catalog once when it
returns — also when it throws, so an interrupted run keeps what it indexed.
putCatalogEntry persists on every call, which is the right default for a
one-off caller and the wrong shape for a sync: indexing a record at a time
re-read and rewrote the whole catalog per record, quadratic in the corpus.
The pipeline wraps its record loop in this; nested batches flush once, at the
outermost, and concurrent ones each at their own end.
The digest of every blob held, sorted.
Filesystem path of a stored blob — what ka open hands to the OS.
Every catalog row, ordered by id.
Remove a stored blob; nothing happens when there is none. ka rm --blobs only (issue #28).
Canonical bytes of a stored record, exactly as they sit on disk.
A frozen artifact the factory built and the line consumes — see CONCEPT.md §0. Named, JSON, and written once by a build-time job rather than by a sync.
Frozen embeddings produced by the factory, if any were shipped.
OptionallockTake the corpus for writing, or throw CorpusLockedError when another run
holds it. Returns the release. Re-entrant within one store object, so a
writer may call another writer; withCorpusLock is the usual way in.
Store bytes under their own digest; returns the digest. Idempotent.
Insert or replace many rows and persist once.
putCatalogEntry rewrites the whole catalog on every call, so building an
index a record at a time wrote it N times — 5.6 MiB to land a 153 KiB file
for 200 records, quadratic in the corpus.
OptionalqueueInside batchCatalog: keep a change to shard's postings and apply it when the
batch ends, with every other change to that shard, before the catalog is written.
Returns false outside a batch, and the caller applies the change at once. Optional:
a store without it applies every change at once (issue #30).
Every record id in the corpus, sorted.
Replace every catalog row with entries and persist, without reading what is
there. That is what a rebuild needs: ka reindex read the old catalog in order
to clear it, so a corrupt catalog was the one thing it could not repair.
Every index shard present, sorted.
Every source that has state in this corpus, sorted — including one that synced and stored nothing. The health report needs it: a source is otherwise only visible through the records it produced, so the very case worth flagging (discovery returned nothing) is the case that leaves no trace.
The corpus, as the roles that make it up.
Storeis their intersection and nothing in the codebase has to change because of the split — but a consumer that only reads the catalog can now say so, and a test double for it does not have to implement blob addressing, index sharding and artifact storage to compile.DiscoverOptionsalready took this shape by hand (Pick<Store, "loadArtifact">); these are the seams it was reaching for.The file store draws exactly these lines with section separators, which is the class saying out loud that it is nine concerns wearing one name.