ka sync
fetch, extract and store Anfragen from one or more sources
ka sync [options]Optionen
| Option | Beschreibung | Standard |
|---|---|---|
--source <key[@window]> |
source to sync, repeatable: several run side by side under one corpus lock. A window of its own after @: berlin@2025-01-01..2025-12-31, bund@period=21, bund@2026-01-01..,limit=50 — the shared --since/--until/--period/--limit fill in what it leaves out (baden-wuerttemberg, bayern, berlin, brandenburg, bremen, bund, hamburg, hessen, mecklenburg-vorpommern, niedersachsen, nordrhein-westfalen, parlamentsspiegel, rheinland-pfalz, saarland, sachsen, sachsen-anhalt, schleswig-holstein, thueringen) | |
--all |
every source with an adapter of its own (not the parlamentsspiegel aggregator) | |
--plan <file> |
run the jobs of a plan file ([[job]] tables: source, since, until, period, limit, ref, retry_failed, only_new, log; see Usage.md) | |
--restart |
with --plan, run every job again, also those done in the plan's unfinished round | |
--wait |
wait while another run holds the corpus, instead of exiting 3 | |
--dry-run |
discover only: count the Anfragen and estimate the download, fetching no document and writing nothing | |
--since <date> |
only Anfragen dated on or after this date (YYYY-MM-DD) | |
--until <date> |
only Anfragen dated on or before this date (YYYY-MM-DD) | |
--period <n> |
restrict to one legislative period | |
--limit <n> |
stop after this many Anfragen (per source) | |
--ref <reference> |
handle only this Anfrage of the window (repeatable), or one it was filed under before; the rest are not asked for | |
--retry-failed |
handle only the Anfragen of the window whose last attempt failed (with --ref: those as well) | |
--only-new |
skip every Anfrage the corpus holds with all its documents, without a request; handle the rest | |
--api-key <key> |
credential for sources that need one (overrides the env var) | |
--metadata-only |
download no documents: a new record abstains on qa, a stored one is rebuilt from its archived documents | |
--force |
re-extract even when inputs and extractor version are unchanged | |
--ignore-robots |
fetch documents from a server whose robots.txt disallows it — your decision; the run warns once per host (the records do not record it) | |
--ocr <mode> |
OCR engine for the ocr tier Werte: off, tesseract, tesseract-js | |
--ocr-language <lang> |
traineddata language for OCR | |
--ocr-version <version> |
require exactly this OCR engine version | |
--ocr-traineddata <path> |
traineddata file to hash into the provenance record | |
--json |
print the sync report as JSON (an array of reports for several jobs, --all or --plan) | |
--log-file <path> |
append the run's events and its other log records to this file as JSON Lines, whatever --log-format is | |
--allow-fs <type> |
write to a corpus on this filesystem anyway (repeatable: fat32, exfat); both are refused by default | |
--min-free <size> |
free space to keep on the corpus and blob volumes, e.g. 500M or 20G; 0 checks none (default: 1.0 GB) |