@maschinenlesbar.org/openka-cli
    Preparing search index...

    Interface SyncSourcesOptions

    The part of the sync options that says what a job covers: the window discovery looks at, and which of the Anfragen it finds are handled (issue #27).

    interface SyncSourcesOptions {
        apiKeyFor?: (source: Source) => string | undefined;
        documents?: DocumentMemo;
        engineFor: (source: Source) => FetchEngine;
        force?: boolean;
        ignoreRobots?: boolean;
        limit?: number;
        metadataOnly?: boolean;
        now?: () => Date;
        onDiscovered?: (job: string, count: number) => void;
        onDone?: (outcome: SourceOutcome) => void;
        onlyNew?: boolean;
        onProgress?: (job: string, event: ProgressEvent) => void;
        onStart?: (job: string) => void;
        perceiver?: Perceiver;
        period?: number;
        refs?: string[];
        retryFailed?: boolean;
        robots?: RobotsPolicy;
        signal?: AbortSignal;
        since?: string;
        sources: readonly Source[];
        space?: SpaceGuard;
        stopOnFailure?: boolean;
        store: Store;
        until?: string;
    }

    Hierarchy (View Summary)

    Index
    apiKeyFor?: (source: Source) => string | undefined

    The credential for a source that takes one (Source.apiKeyEnv), if any.

    documents?: DocumentMemo

    The run's documents, shared the same way: a URL one source fetched is not fetched again.

    engineFor: (source: Source) => FetchEngine

    The engine for one job — a new one per job, best built on one shared HostPacer (EngineOptions.pacer), so a host two jobs reach is paced once.

    force?: boolean

    Re-extract even when nothing changed.

    ignoreRobots?: boolean

    Fetch from a server whose robots.txt disallows it. Two Länder publish their Drucksachen openly and disallow every client; this is the operator's decision to make. The pipeline checks every document URL against its host's robots.txt before fetching it, whichever source produced the URL, and the override is never silent: the report warns once per host.

    limit?: number
    metadataOnly?: boolean

    Download no documents. A new record gets metadata only and abstains on qa; a stored one is re-extracted from the documents it already archived, so a metadata correction lands and nothing it held is lost.

    now?: () => Date

    Injected clock — the only place the pipeline reads time (retrieved_at).

    onDiscovered?: (job: string, count: number) => void
    onDone?: (outcome: SourceOutcome) => void

    A job has finished, failed or been skipped.

    onlyNew?: boolean

    Skip every Anfrage the corpus holds with all its documents, without a request for it; what is missing, or stored without a document, is handled.

    onProgress?: (job: string, event: ProgressEvent) => void
    onStart?: (job: string) => void

    A job is about to start; the callbacks name it by its label.

    perceiver?: Perceiver
    period?: number
    refs?: string[]

    Handle only these references of the window, and the ones they were filed under before (DocRef.formerly). With retryFailed too, either one selects.

    retryFailed?: boolean

    Handle only the Anfragen whose last attempt failed (SourceState.failed).

    robots?: RobotsPolicy

    The run's robots.txt policy, shared by every source of one syncSources run so a host's file is read once for all of them. Built for this source when absent; one that is given must have been built with the same ignoreRobots.

    signal?: AbortSignal

    Stop early: checked before each ref, so the ref in hand is finished, the catalog is saved and the report says interrupted. ka sync aborts it on the first Ctrl-C or SIGTERM.

    since?: string
    sources: readonly Source[]

    The sources to run, each at most once, all over the one window; outcomes come back in this order.

    space?: SpaceGuard

    The disk-space guard (spaceGuard from lib-store). Given, a run whose documents to fetch would not fit — estimated from what the source already archived — is refused with StoreError after discovery and before the first download, and a run stops between two refs, like an aborted one, once a volume drops below the floor (SyncReport.lowSpace).

    stopOnFailure?: boolean

    Start no job once one has failed: the rest are skipped (reason: "after-failure"). Default false.

    store: Store
    until?: string