@maschinenlesbar.org/openka-cli
    Preparing search index...

    Class RobotsPolicy

    Ask a server's robots.txt whether we may fetch a URL, once per host per run.

    This is where CONCEPT.md §7 is enforced for documents. The two gated connectors check their document server before discovering anything, but a Brandenburg PDF also reaches the pipeline through --source parlamentsspiegel, and until this existed that route fetched it with no flag and no warning. The pipeline now asks here before every document request, so the rule holds for whichever door a URL came in by.

    The file is read at run time rather than baked into a connector, so a Land that lifts its Disallow: / stops blocking us the same day — and one that adds a rule starts being honoured the same day. A server with no robots.txt allows everything, which is what a 404 there means. The rules are matched against the User-Agent the engine actually sends, and a host that is fetched under override is slowed to ROBOTS_OVERRIDE_INTERVAL_MS for the rest of the run.

    A robots.txt that cannot be read is not a missing one. RFC 9309 §2.3.1.3–4: a 4xx means there is none (fetch freely), but a 5xx or a network failure means the file is undefined and the crawler "MUST assume complete disallow". Treating every failure as permission turned a padoka outage into a full-speed crawl of a server whose file says Disallow: / — silently (exploratory test 2026-10-07). A 429 is read like a 5xx: it says "not now", not "no file".

    Index
    • Parameters

      • url: string

      Returns Promise<RobotsVerdict>