Files
openclaw/docs/tools/pdf.md
Peter Steinberger edecdbd05e refactor(config): config-surface reduction tranche 3 — product consolidations (review request) (#111527)
* refactor(config): consolidate media model lists

* refactor(config): unify memory configuration

* refactor(config): consolidate TTS ownership

* refactor(config): move typing policy to agents

* refactor(config): retire product-level config surfaces

* refactor(config): share scoped tool policy type

* chore(config): refresh generated baselines

* fix(config): honor agent typing overrides

* fix(config): migrate sibling config consumers

* refactor(infra): keep base64url decoder private

* fix(config): strip invalid legacy TTS values

* chore(config): refresh rebased baseline hash

* fix(doctor): route legacy messages.tts.realtime voice to talk during tts move

* refactor(config): polish final layout names

* refactor(config): freeze retired tuning defaults

* feat(config): add fast mode default symmetry

* refactor(config): key agent entries by id

* docs(config): update final layout reference

* test(config): cover final layout migrations

* chore(config): refresh final layout baselines

* fix(config): align final layout runtime readers

* fix(config): align remaining readers

* fix(config): stabilize final layout migrations

* fix(config): finalize config projection proof

* fix(config): address final layout review

* docs(release): preserve historical config names

* fix(config): complete keyed agent migration

* fix(config): close final migration gaps

* fix(config): finish full-branch review

* fix(config): complete runtime secret detection

* fix(config): close final review findings

* fix(config): finish canonical docs and heartbeat migration

* fix(config): integrate latest main after rebase

* refactor(env): isolate test-only controls

* refactor(env): isolate build and development controls

* refactor(env): collapse process identity indirection

* refactor(env): remove duplicate config and temp aliases

* docs(env): define the operator-facing allowlist

* ci(env): ratchet production variable count

* fix(env): remove stale provider helper import

* fix(env): make ratchet sorting explicit

* test(env): keep test seam in dead-code audit

* test(env): cover ratchet growth and boundary; document surface budgets

* docs(config): document tier-eval consolidations

* docs(config): clarify speech preference ownership

* test(memory): align retired tuning fixtures

* refactor(memory): freeze engine heuristics

* refactor(config): apply tier-eval tranche

* refactor(tts): move persona shaping to providers

* refactor(compaction): move prompt policy to providers

* test(config): align hookified prompt fixtures

* chore(deadcode): classify test-only exports

* chore(github): remove unused spawn helper

* chore(deadcode): classify queue diagnostics

* chore(deadcode): remove unused lane snapshot export

* chore(plugin-sdk): ratchet consolidated surface

* fix(config): integrate latest main after rebase
2026-07-21 20:28:43 -07:00

7.3 KiB

summary, title, read_when
summary title read_when
Analyze one or more PDF documents with native provider support and extraction fallback PDF tool
You want to analyze PDFs from agents
You need exact pdf tool parameters and limits
You are debugging native PDF mode vs extraction fallback

pdf analyzes one or more PDF documents and returns text. It uses native document input on Anthropic and Google models, and falls back to text/image extraction for every other provider.

Availability

The tool registers only when OpenClaw can resolve a PDF-capable model for the agent. Resolution order:

  1. agents.defaults.pdfModel (explicit primary/fallbacks)
  2. agents.defaults.imageModel (explicit primary/fallbacks)
  3. The agent's resolved session/default model, if its provider supports native PDF input (Anthropic, Google) or already has a configured vision model
  4. Auto-detected image/vision-capable providers with usable auth, preferring native-PDF providers first

Every fallback candidate is auth-checked before use, so a configured provider/model only counts if OpenClaw can authenticate that provider for the agent. If no usable model resolves, the pdf tool is not exposed.

Input reference

One PDF path or URL. Multiple PDF paths or URLs, up to 10 total. Analysis prompt. Page filter like `1-5` or `1,3,7-9`. Not supported in native provider mode. Password for encrypted PDFs. Applies to every PDF in the request; only used by extraction fallback mode. Optional model override in `provider/model` form. Per-PDF size cap in MB. Defaults to `agents.defaults.pdfMaxMb`, or `10` if unset.

Notes:

  • pdf and pdfs are merged and deduplicated before loading; at least one is required.
  • pages is parsed as 1-based page numbers, deduped, sorted, and clamped to agents.defaults.pdfMaxPages (default 20). A range that matches no in-bounds pages errors before the model call.

Supported PDF references

  • Local file path (including ~ expansion)
  • file:// URL
  • http:// and https:// URL
  • OpenClaw-managed inbound refs such as media://inbound/<id>

Other URI schemes (for example ftp://) return details.error = "unsupported_pdf_reference". Remote http(s) URLs are rejected when the tool runs sandboxed. With workspace-only file policy enabled, local paths outside allowed roots are rejected; managed inbound refs and replayed paths under OpenClaw's inbound media store are still allowed.

Execution modes

Native provider mode

Used for provider anthropic and google (the only providers that currently declare native PDF document support). Raw PDF bytes go directly to the provider API as a native document/inline-PDF part per file.

Limits:

  • pages is not supported; if set, the tool throws pages is not supported with native PDF providers.
  • password is not supported; if set, the tool throws password is not supported with native PDF providers. Use a non-native model for encrypted PDFs.

Extraction fallback mode

Used for every other provider.

  1. Extract text from the selected pages (up to agents.defaults.pdfMaxPages, default 20) via the bundled document-extract plugin, which uses the clawpdf package (PDFium WebAssembly) for text and image extraction.
  2. If the extracted text is shorter than 200 characters, render the same pages to PNG images. The render budget is 4,000,000 pixels total, shared across all pages needing images (allocated proportionally per remaining page, not per page), so text pages that already have enough text skip rendering entirely.
  3. Send the extracted text (and any rendered images) plus the prompt to the selected model.

Details:

  • Encrypted PDFs open with the top-level password parameter.
  • If the model has no image input and there is no extractable text, the tool errors.
  • If image rendering fails, OpenClaw drops the images and continues with the extracted text.
  • If the target model is text-only and extraction produced images, OpenClaw drops the images and sends text only.

Config

{
  agents: {
    defaults: {
      pdfModel: {
        primary: "anthropic/claude-opus-4-6",
        fallbacks: ["openai/gpt-5.4-mini"],
      },
      pdfMaxBytesMb: 10,
      pdfMaxPages: 20,
    },
  },
}
Key Default Meaning
agents.defaults.pdfModel unset Explicit primary/fallback PDF models; falls back to imageModel, then the session model.
agents.defaults.pdfMaxMb 10 Per-PDF size cap in MB.
agents.defaults.pdfMaxPages 20 Max pages processed per PDF.

See Configuration Reference for full field details.

Output details

The tool returns text in content[0].text and structured metadata in details.

Common details fields:

  • model: resolved model ref (provider/model)
  • native: true for native provider mode, false for fallback
  • attempts: fallback attempts that failed before success

Path fields:

  • Single PDF input: details.pdf
  • Multiple PDF inputs: details.pdfs[] with pdf entries
  • Sandbox path rewrite metadata (when applicable): rewrittenFrom

Error behavior

Condition Result
No PDF input Throws pdf required: provide a path or URL to a PDF document
More than 10 PDFs details.error = "too_many_pdfs"
Unsupported reference scheme details.error = "unsupported_pdf_reference"
pages with a native provider Throws pages is not supported with native PDF providers
password with a native provider Throws password is not supported with native PDF providers

Examples

Single PDF:

{
  "pdf": "/tmp/report.pdf",
  "prompt": "Summarize this report in 5 bullets"
}

Multiple PDFs:

{
  "pdfs": ["/tmp/q1.pdf", "/tmp/q2.pdf"],
  "prompt": "Compare risks and timeline changes across both documents"
}

Page-filtered fallback model:

{
  "pdf": "https://example.com/report.pdf",
  "pages": "1-3,7",
  "model": "openai/gpt-5.4-mini",
  "prompt": "Extract only customer-impacting incidents"
}

Encrypted PDF with extraction fallback:

{
  "pdf": "/tmp/locked.pdf",
  "password": "example-password",
  "model": "openai/gpt-5.4-mini",
  "prompt": "Summarize this contract"
}