Kommandozeilen-Tool und TypeScript-Bibliothek
ka
Kleine Anfragen aus siebzehn deutschen Parlamenten – Bundestag und alle sechzehn Landtage – in einem einheitlichen, reproduzierbaren Format. ka ist ein Kommandozeilen-Werkzeug, das die Drucksachen dort abholt, wo das Parlament sie selbst veröffentlicht, Frage und Antwort deterministisch aus dem PDF liest und beides als geprüften Datensatz ablegt.
Entscheidend ist, was das Werkzeug nicht tut: Es rät nicht. Wenn ein Extraktor eine Stelle nicht sicher lesen kann, erfindet er nichts, sondern enthält sich – das Feld bleibt leer, wird in abstained_fields benannt und der Datensatz zur Prüfung vorgemerkt. Kein Sprachmodell läuft auf diesem Weg; ka verify rechnet jeden Datensatz aus den archivierten Bytes nach und vergleicht ihn Byte für Byte.
Installation
Benötigt Node.js 22 oder neuer. Installiert den Befehl ka:
npm i -g @maschinenlesbar.org/openka-cliZugang
Kein Konto, kein API-Schlüssel, keine Konfiguration – installieren und loslegen.
Schnellstart
# Ingest a window of Berlin's Schriftliche Anfragen (question + answer + PDF text)
ka sync --source berlin --since 2024-01-01 --limit 50
# Search it
ka search "Brücken Zustand" --parliament berlin --year 2024
ka show berlin-19-18221
# Get the canonical record, or another rendering of it
ka get berlin-19-18221 --format json # canonical JSON: the stored bytes
ka get berlin-19-18221 --format md # readable
ka get berlin-19-18221 --format jsonld # schema.org
# Prove it: re-run the extraction from the archived bytes and compare
ka verify berlin-19-18221
# See what the extractor refused to answer
ka review
# Bulk output
ka export --format csv --out corpus.csv
ka feed --party GRÜNE --out gruene.atomBefehle
Die Beschreibungen der Befehle stammen direkt aus der CLI und sind englisch.
-
syncfetch, extract and store Anfragen from one or more sources
-
searchfull-text search over the corpus
-
getprint one record in a machine-readable format
-
showrender one record for reading
-
openprint the path of a record's archived source document
-
verifyre-run an extraction from the archived bytes and assert identical output
-
reviewwork the abstention queue: records the extractor refused to complete
-
reindexrebuild the search index and catalog from the stored records
-
sourcesthe source map and its health
-
reextractre-extract stored records from their archived bytes with this build's extractor — no network; then rebuild the index
-
rmremove records from the corpus, with their catalog rows and index postings (takes the corpus lock)
-
exportexport the corpus (or a selection of it) in bulk
-
feedan Atom feed of the newest matching Anfragen
-
schemaprint the JSON Schema of the canonical record
-
statswhat is in this corpus: coverage, completeness, extractor versions, disk use, and breakdowns with --by
-
doctorcheck the corpus: filesystem, free space, lock, catalog against records, macOS ._* files (exit 3 on a problem)
-
statuswhat a running sync is doing — progress, rate, time left — or how the last one ended
-
configcredentials kept apart from the corpus, in $XDG_CONFIG_HOME/openka/credentials (bund.api-key)
TypeScript-Bibliothek
@maschinenlesbar.org/openka-cli ist auch ein typisierter API-Client, den Sie in eigenem Code importieren können – ohne HTTP-Abhängigkeiten zur Laufzeit.
Das Werkzeug, nicht die Daten
Die Daten stammen vom Anbieter hinter der API und unterliegen dessen Bedingungen – siehe DATA_LICENSE.md. Der Code steht unter AGPL-3.0-or-later oder einer kommerziellen Lizenz – siehe LICENSING.md.