Commands

Every command, argument and option of ka 0.10.0, read from the CLI itself.

Global options

These options work with every command.

OptionDescriptionDefault
-v, --version output the version number
--corpus <dir> corpus directory (default: $OPENKA_CORPUS, else $XDG_DATA_HOME/openka, else ~/.local/share/openka)
--blobs <dir> keep the archived documents here instead of <corpus>/blobs — an existing directory, e.g. on an external drive (default: $OPENKA_BLOBS)
--timeout <ms> timeout per request attempt in milliseconds (a timed-out request is retried once)
--user-agent <ua> User-Agent sent to upstreams
--max-retries <n> retries for a transient 429/503 or a dropped connection (never an over-size response)
--max-response-bytes <n> hard cap on a single response body (default: 134217728, 128 MiB); a document over it is left out of its record, with a warning
--min-host-interval <ms> minimum delay between requests to one host (default: 500); a source's own floor (ka sources show <key>) is never lowered by it
--max-redirects <n> redirects to follow (0 = surface a 3xx as an error)
--compact compact JSON output
--quiet suppress progress output on stderr
--log-format <format> how errors, warnings, notes and progress are written to stderr: text (log4j style: time, level, [topic], message) or jsonl (one JSON object per line: ts, level, topic, msg; `ka sync` adds its event fields); default text

ka sync

fetch, extract and store Anfragen from one or more sources

ka sync [options]

Options

OptionDescriptionDefault
--source <key[@window]> source to sync, repeatable: several run side by side under one corpus lock. A window of its own after @: berlin@2025-01-01..2025-12-31, bund@period=21, bund@2026-01-01..,limit=50 — the shared --since/--until/--period/--limit fill in what it leaves out (baden-wuerttemberg, bayern, berlin, brandenburg, bremen, bund, hamburg, hessen, mecklenburg-vorpommern, niedersachsen, nordrhein-westfalen, parlamentsspiegel, rheinland-pfalz, saarland, sachsen, sachsen-anhalt, schleswig-holstein, thueringen)
--all every source with an adapter of its own (not the parlamentsspiegel aggregator)
--plan <file> run the jobs of a plan file ([[job]] tables: source, since, until, period, limit, ref, retry_failed, only_new, log; see Usage.md)
--restart with --plan, run every job again, also those done in the plan's unfinished round
--wait wait while another run holds the corpus, instead of exiting 3
--dry-run discover only: count the Anfragen and estimate the download, fetching no document and writing nothing
--since <date> only Anfragen dated on or after this date (YYYY-MM-DD)
--until <date> only Anfragen dated on or before this date (YYYY-MM-DD)
--period <n> restrict to one legislative period
--limit <n> stop after this many Anfragen (per source)
--ref <reference> handle only this Anfrage of the window (repeatable), or one it was filed under before; the rest are not asked for
--retry-failed handle only the Anfragen of the window whose last attempt failed (with --ref: those as well)
--only-new skip every Anfrage the corpus holds with all its documents, without a request; handle the rest
--api-key <key> credential for sources that need one (overrides the env var)
--metadata-only download no documents: a new record abstains on qa, a stored one is rebuilt from its archived documents
--force re-extract even when inputs and extractor version are unchanged
--ignore-robots fetch documents from a server whose robots.txt disallows it — your decision; the run warns once per host (the records do not record it)
--ocr <mode> OCR engine for the ocr tier Choices: off, tesseract, tesseract-js
--ocr-language <lang> traineddata language for OCR
--ocr-version <version> require exactly this OCR engine version
--ocr-traineddata <path> traineddata file to hash into the provenance record
--json print the sync report as JSON (an array of reports for several jobs, --all or --plan)
--log-file <path> append the run's events and its other log records to this file as JSON Lines, whatever --log-format is
--allow-fs <type> write to a corpus on this filesystem anyway (repeatable: fat32, exfat); both are refused by default
--min-free <size> free space to keep on the corpus and blob volumes, e.g. 500M or 20G; 0 checks none (default: 1.0 GB)

ka get

print one record in a machine-readable format

ka get [options] <id>

Arguments

<id> required
record id, e.g. berlin-19-10006

Options

OptionDescriptionDefault
--format <format> output format Choices: json, jsonld, csv, md, text
-o, --out <file> write to this file instead of stdout (- = stdout; an existing file needs --force)
--force with --out, replace an existing file

ka show

render one record for reading

ka show [options] <id>

Arguments

<id> required
record id

ka open

print the path of a record's archived source document

ka open [options] <id>

Arguments

<id> required
record id

Options

OptionDescriptionDefault
--role <role> which document to open when there are several (question_pdf, answer_pdf, combined_pdf, metadata)

ka verify

re-run an extraction from the archived bytes and assert identical output

ka verify [options] [id]

Arguments

[id] optional
record id; omit to verify an evenly spaced sample of the corpus

Options

OptionDescriptionDefault
--all verify every record
--limit <n> how many records to verify when no id is given (default: 25)
--ocr <mode> OCR engine to use for records produced with one Choices: off, tesseract, tesseract-js
--json print results as JSON

ka review

work the abstention queue: records the extractor refused to complete

ka review [options]

Options

OptionDescriptionDefault
--parliament <key> restrict to one parliament (as in search, export and stats)
--source <key> the same as --parliament
--include-known-gaps also the records whose only holes are fields their parliament never provides (ka sources show <key>)
--limit <n> how many records to list (default: 20)
--mark-verified <id> record that a human checked this record against its source
--group-by <what> summarise the queue per source by the kind of field abstained on, with example ids Choices: field
--json print the queue as JSON

ka reindex

rebuild the search index and catalog from the stored records

ka reindex [options]

ka sources

the source map and its health

ka sources list

every parliament, its adapter status and its last sync

ka sources list [options]

Options

OptionDescriptionDefault
--json print as JSON

ka sources count

how many Anfragen each upstream holds, beside how many the corpus has — a request or two per source, no download

ka sources count [options]

Options

OptionDescriptionDefault
--source <key> count only this source (repeatable; default: every parliament)
--period <n> count one legislative period (DIP can; the Parlamentsspiegel cannot)
--api-key <key> credential for sources that need one (overrides the env var)
--json print as JSON

ka sources show

what one source does and what is specific about it

ka sources show [options] <key>

Arguments

<key> required
source key

ka reextract

re-extract stored records from their archived bytes with this build's extractor — no network; then rebuild the index

ka reextract [options] [ids...]

Arguments

[ids…] optional
record ids; or select with --all or the filters

Options

OptionDescriptionDefault
--all every record in the corpus
--force also the records this build's extractor already stamped
--dry-run say what would change, and write nothing
--ocr <mode> OCR engine for records produced with one Choices: off, tesseract, tesseract-js
--json print the report as JSON
--parliament <key> restrict to a parliament (repeatable; 17 known, see `ka sources list`)
--party <name> restrict to Anfragen asked by this party (repeatable)
--year <yyyy> restrict to the year the Anfrage was asked (repeatable)
--period <n> restrict to a legislative period (repeatable)
--from <date> asked on or after this date (a record whose question date is unknown is left out)
--to <date> asked on or before this date (a record whose question date is unknown is left out)

ka rm

remove records from the corpus, with their catalog rows and index postings (takes the corpus lock)

ka rm [options] [ids...]

Arguments

[ids…] optional
record ids; or select with the filters

Options

OptionDescriptionDefault
--documents also remove their archived documents that no remaining record refers to
--orphaned-documents remove every archived document no record refers to, and no record
--move-to <dir> move the files under this directory (records/, blobs/) instead of deleting them
--dry-run say what would be removed, and change nothing
--json print the report as JSON
--parliament <key> restrict to a parliament (repeatable; 17 known, see `ka sources list`)
--party <name> restrict to Anfragen asked by this party (repeatable)
--year <yyyy> restrict to the year the Anfrage was asked (repeatable)
--period <n> restrict to a legislative period (repeatable)
--from <date> asked on or after this date (a record whose question date is unknown is left out)
--to <date> asked on or before this date (a record whose question date is unknown is left out)

ka export

export the corpus (or a selection of it) in bulk

ka export [options]

Options

OptionDescriptionDefault
--format <format> output format Choices: csv, jsonl, jsonld
--limit <n> maximum records to export (most relevant first with --query)
--parliament <key> restrict to a parliament (repeatable; 17 known, see `ka sources list`)
--party <name> restrict to Anfragen asked by this party (repeatable)
--year <yyyy> restrict to the year the Anfrage was asked (repeatable)
--period <n> restrict to a legislative period (repeatable)
--from <date> asked on or after this date (a record whose question date is unknown is left out)
--to <date> asked on or before this date (a record whose question date is unknown is left out)
--query <terms> restrict to records matching these search terms
-o, --out <file> write to this file instead of stdout (- = stdout; an existing file needs --force)
--force with --out, replace an existing file

ka feed

an Atom feed of the newest matching Anfragen

ka feed [options]

Options

OptionDescriptionDefault
--limit <n> entries in the feed
--title <text> feed title (default: OpenKA — Kleine Anfragen)
--id <url> feed id / self link (default: urn:openka:feed)
--parliament <key> restrict to a parliament (repeatable; 17 known, see `ka sources list`)
--party <name> restrict to Anfragen asked by this party (repeatable)
--year <yyyy> restrict to the year the Anfrage was asked (repeatable)
--period <n> restrict to a legislative period (repeatable)
--from <date> asked on or after this date (a record whose question date is unknown is left out)
--to <date> asked on or before this date (a record whose question date is unknown is left out)
--query <terms> restrict to records matching these search terms
-o, --out <file> write to this file instead of stdout (- = stdout; an existing file needs --force)
--force with --out, replace an existing file

ka schema

print the JSON Schema of the canonical record

ka schema [options]

ka stats

what is in this corpus: coverage, completeness, extractor versions, disk use, and breakdowns with --by

ka stats [options]

Options

OptionDescriptionDefault
--by <dimension> break the records down by party, ministry, month, year, period, parliament; twice for a cross-tab (--by party --by year)
--no-disk leave out what the corpus takes on disk (one stat per file)
--json print as JSON
--parliament <key> restrict to a parliament (repeatable; 17 known, see `ka sources list`)
--party <name> restrict to Anfragen asked by this party (repeatable)
--year <yyyy> restrict to the year the Anfrage was asked (repeatable)
--period <n> restrict to a legislative period (repeatable)
--from <date> asked on or after this date (a record whose question date is unknown is left out)
--to <date> asked on or before this date (a record whose question date is unknown is left out)

ka doctor

check the corpus: filesystem, free space, lock, catalog against records, macOS ._* files (exit 3 on a problem)

ka doctor [options]

Options

OptionDescriptionDefault
--fix remove the macOS ._* and .DS_Store files from the corpus (takes the corpus lock)
--orphaned-documents also count the archived documents no record refers to (reads every record)
--json print the diagnosis as JSON
--allow-fs <type> write to a corpus on this filesystem anyway (repeatable: fat32, exfat); both are refused by default
--min-free <size> free space to keep on the corpus and blob volumes, e.g. 500M or 20G; 0 checks none (default: 1.0 GB)

ka status

what a running sync is doing — progress, rate, time left — or how the last one ended

ka status [options]

Options

OptionDescriptionDefault
--json print the status as JSON
--watch look again every 5 s until no sync runs
--stalled-after <duration> exit 1 when a running sync has not moved for this long (90s, 10m, 2h), or its process is gone

ka config

credentials kept apart from the corpus, in $XDG_CONFIG_HOME/openka/credentials (bund.api-key)

ka config set

store a credential: typed at a prompt without echo, or piped in — never given as an argument

ka config set [options] <name>

Arguments

<name> required
bund.api-key

ka config get

show a stored credential, masked (abcd…wxyz) unless --reveal

ka config get [options] <name>

Arguments

<name> required
bund.api-key

Options

OptionDescriptionDefault
--reveal print the whole value, for a script that passes it on — it then is on your screen or in its log

ka config unset

remove a stored credential

ka config unset [options] <name>

Arguments

<name> required
bund.api-key

ka config list

every stored credential, masked, and where the file is

ka config list [options]