openclaw/extensions/document-extract
Peter Steinberger a1eb59d08b
refactor(plugins): deslop long-tail extensions third pass (#161644)
* refactor(plugins): deslop long-tail extensions third pass

* refactor(beam): keep catalog summary type private
2026-09-29 23:16:58 -07:00
..
assets improve(plugins): give bundled logos consistent white icon tiles (#155259) 2026-09-23 19:09:26 -07:00
document-extractor-worker-entrypoint.ts improve: keep image and PDF processing responsive (#146094) 2026-09-12 10:58:14 -07:00
document-extractor.runtime.test.ts test(discord,crabbox,plugins): remove low-value tests (batch d058) (#159513) 2026-09-27 07:34:12 +00:00
document-extractor.runtime.ts refactor(plugins): deslop long-tail extensions third pass (#161644) 2026-09-29 23:16:58 -07:00
document-extractor.test-support.ts fix(workers): prevent PDF cancellation from aborting Node (#150752) 2026-09-17 03:38:18 -07:00
document-extractor.test.ts fix(pdf): surface partial document extraction (#131922) 2026-09-23 17:52:04 -07:00
document-extractor.ts improve: keep image and PDF processing responsive (#146094) 2026-09-12 10:58:14 -07:00
document-extractor.worker.ts refactor(plugins): deslop tool, search and media plugins (#160552) 2026-09-29 00:31:49 +00:00
index.ts
openclaw.plugin.json feat(plugins): assign one purpose category to every bundled plugin (#142760) 2026-09-10 20:44:20 -07:00
package.json chore(release): close out 2026.9.7 on main (#161587) 2026-09-29 22:39:14 -07:00
README.md feat: show declared plugin capabilities and setup guides (#157956) 2026-09-25 16:43:53 -07:00

Document Extraction

Extract text from PDF attachments locally. When a selected page has too little text, the plugin can render it as an image for a vision-capable model. PDF processing runs in a worker through the bundled PDFium-based extractor.

Get started

The plugin is enabled by default and needs no extraction API key. Configure a model for PDF analysis, then attach a PDF or ask your agent to analyze a local PDF file.

Local extraction and model analysis are separate: the selected model still needs its normal credentials. Models with native PDF support can receive the document directly instead. This extractor handles PDFs, not every document format.

See the PDF guide for model selection, page limits, and encrypted documents.