Pull the embedded images out of a PDF, without rasterising anything.
Scanned documents are almost always one compressed image per page, so handing
those bytes straight to OCR needs no rendering engine — which is why the line can
do the ocr tier with no graphics dependency at all. Images in formats that are
only meaningful once rendered (raw Flate bitmaps) are skipped; the caller then
has fewer images than pages and abstains for the pages it could not cover.
Pull the embedded images out of a PDF, without rasterising anything.
Scanned documents are almost always one compressed image per page, so handing those bytes straight to OCR needs no rendering engine — which is why the line can do the
ocrtier with no graphics dependency at all. Images in formats that are only meaningful once rendered (raw Flate bitmaps) are skipped; the caller then has fewer images than pages and abstains for the pages it could not cover.