@maschinenlesbar.org/openka-cli
    Preparing search index...

    Interface PdfTextResult

    interface PdfTextResult {
        imageOnly: boolean;
        lostObjects: number;
        missingPages: number;
        pageCount: number;
        pages: PdfPageText[];
        problems: string[];
        text: string;
        truncated: boolean;
        undecodableStreams: number;
        unmappedRatio: number;
        unreadPages: number[];
        version: string;
    }
    Index
    imageOnly: boolean

    True when the document has no text-showing operators at all but draws images (a scan). A document with neither is not image-only: it has nothing to OCR.

    lostObjects: number

    Objects lost with an object stream that would not decode.

    missingPages: number

    Pages the page tree declares (/Count) beyond those found in the file.

    pageCount: number
    pages: PdfPageText[]
    problems: string[]

    Reasons the reader could not do its job fully — each one drives an abstention.

    text: string

    All pages joined with a form feed, the form the extractors parse.

    truncated: boolean

    True when the file does not end with %%EOF: the download stopped early. What was read may be partial, or an earlier revision of an object that a later, cut-off update replaced.

    undecodableStreams: number

    Content streams refused outright — an encrypted or unsupported filter — plus pages refused because drawing them exceeded the reader's work budget.

    unmappedRatio: number

    Share of character codes no font mapping covered, 0..1.

    unreadPages: number[]

    Pages whose content the reader could not get — named but missing from the file, refused (an undecodable stream, the work budget) — in page order.

    version: string