@maschinenlesbar.org/openka-cli
    Preparing search index...

    Interface SyncOptions

    One source over one window (SyncWindow: what is discovered, and which of it is handled).

    interface SyncOptions {
        apiKey?: string;
        documents?: DocumentMemo;
        engine: FetchEngine;
        force?: boolean;
        ignoreRobots?: boolean;
        limit?: number;
        metadataOnly?: boolean;
        now?: () => Date;
        onDiscovered?: (count: number) => void;
        onlyNew?: boolean;
        onProgress?: (event: ProgressEvent) => void;
        perceiver?: Perceiver;
        period?: number;
        refs?: string[];
        retryFailed?: boolean;
        robots?: RobotsPolicy;
        signal?: AbortSignal;
        since?: string;
        source: Source;
        space?: SpaceGuard;
        store: Store;
        until?: string;
    }

    Hierarchy (View Summary)

    Index
    apiKey?: string
    documents?: DocumentMemo

    The run's documents, shared the same way: a URL one source fetched is not fetched again.

    engine: FetchEngine
    force?: boolean

    Re-extract even when nothing changed.

    ignoreRobots?: boolean

    Fetch from a server whose robots.txt disallows it. Two Länder publish their Drucksachen openly and disallow every client; this is the operator's decision to make. The pipeline checks every document URL against its host's robots.txt before fetching it, whichever source produced the URL, and the override is never silent: the report warns once per host.

    limit?: number
    metadataOnly?: boolean

    Download no documents. A new record gets metadata only and abstains on qa; a stored one is re-extracted from the documents it already archived, so a metadata correction lands and nothing it held is lost.

    now?: () => Date

    Injected clock — the only place the pipeline reads time (retrieved_at).

    onDiscovered?: (count: number) => void

    Called once discovery is done, with the number of Anfragen the run will handle — the total a progress display counts towards. Not called when discovery fails.

    onlyNew?: boolean

    Skip every Anfrage the corpus holds with all its documents, without a request for it; what is missing, or stored without a document, is handled.

    onProgress?: (event: ProgressEvent) => void

    Called after each record, for progress output.

    perceiver?: Perceiver
    period?: number
    refs?: string[]

    Handle only these references of the window, and the ones they were filed under before (DocRef.formerly). With retryFailed too, either one selects.

    retryFailed?: boolean

    Handle only the Anfragen whose last attempt failed (SourceState.failed).

    robots?: RobotsPolicy

    The run's robots.txt policy, shared by every source of one syncSources run so a host's file is read once for all of them. Built for this source when absent; one that is given must have been built with the same ignoreRobots.

    signal?: AbortSignal

    Stop early: checked before each ref, so the ref in hand is finished, the catalog is saved and the report says interrupted. ka sync aborts it on the first Ctrl-C or SIGTERM.

    since?: string
    source: Source
    space?: SpaceGuard

    The disk-space guard (spaceGuard from lib-store). Given, a run whose documents to fetch would not fit — estimated from what the source already archived — is refused with StoreError after discovery and before the first download, and a run stops between two refs, like an aborted one, once a volume drops below the floor (SyncReport.lowSpace).

    store: Store
    until?: string