Open the corpus at root, or create it on the first write. For a writer
(sync, a test fixture): a directory that is not there yet is where the corpus
will be. A reader wants FileStore.open instead.
Optionaloptions: FileStoreOptionsReadonlyblobsWhere the archived documents are: <root>/blobs, or the blobs option.
ReadonlyrootAbsolute path of the corpus root, for messages and ka open.
Whether macOS writes an AppleDouble companion (._<name>) beside every file
this corpus holds — a FAT32 or exFAT volume. Known once lock() has run; false
before that, and on every other system.
Throw StoreError when the archived documents cannot be reached
(blobStoreProblem). Whatever reads or writes them — a sync, ka open,
ka verify — asks first, so an unplugged drive is named as such rather than
surfacing as a raw ENOENT or as every document "missing".
Nested scopes — one call inside another's work — flush once, at the
outermost. Concurrent ones, two syncs of one ka sync --source a --source b,
flush each at its own end: counting them together meant the count was rarely
back at zero while both ran, and a run killed then lost every catalog row since
it started, where the checkpoints promise at most one batch.
Every blob held, by digest, sorted: the <sha[0:2]>/<sha>.bin files of the blob
store. Anything else there — a platform file, a stray name — is not a blob.
Where the blob for digest lives — built, not checked: the file may be
missing or corrupt. getBlob checks the bytes, and archivedDocument hands
out a path only after checking them.
Why the archived documents cannot be reached now, or undefined when they can.
Only a blob directory named apart from the corpus can be missing — its drive
unplugged; <root>/blobs is created on the first write.
Every catalog row, ordered by id.
Remove a stored blob; nothing happens when there is none. ka rm --blobs only (issue #28).
Write the catalog, merging this process's changes onto what is on disk now.
A corpus is a directory and nothing locks it, so two ka sync runs — or a
sync racing a reindex — both cached the catalog at startup and then wrote it
whole. The second write dropped the first run's rows: the record stayed on
disk and vanished from search, stats, export, feed and health, silently, with
only ka reindex to recover it.
Re-reading here means a concurrent writer's rows survive. It does not make the
corpus transactional — two writes can still interleave between this read and
the rename — but it removes the failure that needed no race at all, just two
runs that started before either finished. The writers that go through the
library (sync, reindexAll, markHumanVerified) now also hold lock(),
which is what keeps the index shards, written without a merge, intact.
Apply the queued posting changes: each shard read and written once, in shard order.
The archived bytes, checked against their own name. The store is
content-addressed, and a blob that no longer hashes to its name is not the
document the record was built from: read unchecked, ka verify blamed the
extractor for "different bytes with the same version" and ka open handed
the altered file out without comment.
A plan's open round (queueProgressKey), or undefined when none is open.
Canonical bytes of a stored record, exactly as they sit on disk.
The status file a sync keeps (RunStatusRecorder), or undefined when none was written.
Platform files (isPlatformFile) the listings have skipped so far, as paths, sorted.
A frozen artifact the factory built and the line consumes — see CONCEPT.md §0. Named, JSON, and written once by a build-time job rather than by a sync.
Frozen embeddings produced by the factory, if any were shipped.
Take the corpus for writing. The lock is a file created exclusively (wx),
holding the pid, the host and the purpose of the run that took it. A second
writer — another ka sync, a ka reindex, another FileStore in the same
process — gets CorpusLockedError instead of interleaving its writes.
A lock left behind by a run that was killed is taken over when it names this host and a process that no longer exists. A lock from another host (a corpus on a network share) cannot be checked, so it is never taken over: the error names the file to delete.
Who holds the corpus, read without taking it: undefined when nobody does. A
stale lock names this host and a process that is gone — the next writer takes
it over; a lock from another host is never stale here, since it cannot be checked.
Store bytes under their own digest; returns the digest. Idempotent.
Insert or replace many rows and persist once.
putCatalogEntry rewrites the whole catalog on every call, so building an
index a record at a time wrote it N times — 5.6 MiB to land a 153 KiB file
for 200 records, quadratic in the corpus.
Insert or replace one catalog row and persist the catalog.
Record a plan's open round; undefined closes it (the file is removed). Writers hold the lock.
Replace the status file, atomically: a reader never sees half of one.
Inside a batchCatalog scope a change is kept until the batch ends, and each shard
is then read and written once for all of them (issue #30). Reads inside the batch
see the shards as they were; nothing in a sync reads them.
Every record id in the corpus, sorted — the basis for a full re-index. A record
file whose name is not a record id was not written by the store: it is a
StoreError here, so it does not reach getRecord as a caller's usage error —
except a platform file (isPlatformFile), which is skipped.
Replace every catalog row with entries and persist, without reading what is
there. That is what a rebuild needs: ka reindex read the old catalog in order
to clear it, so a corrupt catalog was the one thing it could not repair.
Every shard file present, sorted — used by a full index rebuild and by stats.
Every source that has state in this corpus, sorted — including one that synced and stored nothing. The health report needs it: a source is otherwise only visible through the records it produced, so the very case worth flagging (discovery returned nothing) is the case that leaves no trace.
StaticopenOpen the corpus at root for reading: it must already exist. A missing
directory throws MissingCorpusError, a path that is not a directory
StoreError (both exit 3 in ka). new FileStore on the same path answers
an empty corpus — no records, no matches — which for a mistyped path is the
wrong answer, indistinguishable from a corpus with nothing in it. Nothing is
created.
Optionaloptions: FileStoreOptions
The corpus, as the roles that make it up.
Storeis their intersection and nothing in the codebase has to change because of the split — but a consumer that only reads the catalog can now say so, and a test double for it does not have to implement blob addressing, index sharding and artifact storage to compile.DiscoverOptionsalready took this shape by hand (Pick<Store, "loadArtifact">); these are the seams it was reaching for.The file store draws exactly these lines with section separators, which is the class saying out loud that it is nine concerns wearing one name.