@maschinenlesbar.org/openka-cli
    Preparing search index...

    Interface Source

    interface Source {
        apiKeyEnv?: string;
        homepage: string;
        key: string;
        label: string;
        minHostIntervalMs?: number;
        minHostIntervalReason?: string;
        notes: string;
        parliament?:
            | "bund"
            | "baden-wuerttemberg"
            | "bayern"
            | "berlin"
            | "brandenburg"
            | "bremen"
            | "hamburg"
            | "hessen"
            | "mecklenburg-vorpommern"
            | "niedersachsen"
            | "nordrhein-westfalen"
            | "rheinland-pfalz"
            | "saarland"
            | "sachsen"
            | "sachsen-anhalt"
            | "schleswig-holstein"
            | "thueringen";
        ruleSets?: readonly SegmentationRules[];
        tier: "structured"
        | "text_layer"
        | "ocr";
        checkRecord?(ref: DocRef, record: KaRecord): string | undefined;
        count?(options: CountOptions): Promise<UpstreamCount>;
        discover(options: DiscoverOptions): Promise<DiscoverResult>;
    }

    Implemented by

    Index
    apiKeyEnv?: string

    Environment variable holding this source's credential, when it needs one.

    homepage: string

    Where a human can see the same data.

    key: string
    label: string
    minHostIntervalMs?: number

    A politeness floor for this source, in milliseconds between requests to one host. Sources that reach a server which has asked not to be crawled set it well above the default: if the operator has decided to fetch anyway, the least the tool can do is go slowly.

    minHostIntervalReason?: string

    Why the source sets minHostIntervalMs, for ka sources show and the sync's note.

    notes: string

    What is genuinely specific about this source — quirks worth knowing.

    parliament?:
        | "bund"
        | "baden-wuerttemberg"
        | "bayern"
        | "berlin"
        | "brandenburg"
        | "bremen"
        | "hamburg"
        | "hessen"
        | "mecklenburg-vorpommern"
        | "niedersachsen"
        | "nordrhein-westfalen"
        | "rheinland-pfalz"
        | "saarland"
        | "sachsen"
        | "sachsen-anhalt"
        | "schleswig-holstein"
        | "thueringen"

    The parliament this adapter is pinned to, when it is pinned to one. An adapter covering several Länder leaves it out rather than naming a placeholder: every ref it yields carries its own parliament, so there is nothing to fall back to.

    ruleSets?: readonly SegmentationRules[]

    Rule sets to use for segmentation; the shared default when omitted.

    tier: "structured" | "text_layer" | "ocr"

    The tier the pipeline runs for documents from this source.

    • Whether an extracted record is the paper its ref names: a reason when it is not, else undefined. The pipeline then stores nothing for the ref and reports the reason. Optional, because only some Länder print their Drucksachennummer in one form a source can rely on; a record is never changed by it, so ka verify and the extractor version are untouched.

      Parameters

      Returns string | undefined

    • How many Anfragen the upstream holds, asked with a request or two and no discovery — what ka sources count sets beside the corpus. Optional: a source that could only count by discovering everything leaves it out. A period the upstream cannot count by is a UsageError, not a number for something else.

      Parameters

      Returns Promise<UpstreamCount>