npx skills add ...
npx skills add allenai/asta-plugins --skill pdf-extraction
Extract text from PDFs using the Asta remote OCR API. Use when the user asks to "extract text from PDF", "OCR a document", "read a PDF", or needs to process scanned documents.
npx skills add allenai/asta-plugins --skill pdf-extraction
Extract text from a local PDF with the Asta remote OCR API. The command returns markdown that preserves document structure, headings, formatting, tables, lists, and emphasis.
<pdf>: local PDF path (required; the file must exist and be readable)-o / --output: markdown output path (default: stdout)--start-page: first page to process, using zero-based page numbering
(default: 0)--max-pages: maximum number of pages to process (default: 50)--images / --no-images: save embedded images alongside the markdown
(default: --no-images)The command requires an authenticated Asta session.
Identify the local PDF and where the user wants the markdown saved. Prefer an
explicit -o path so the result is available for later work. The command
creates missing parent directories for the output file.
For PDFs of 50 pages or fewer, use the defaults:
For longer PDFs, process consecutive ranges in separate commands. Page numbers
are zero-based, so the second 50-page range begins at page 50:
Keep the part numbers zero-padded so the files sort in page order.
Use --images when figures, diagrams, or other embedded images are important:
The images are saved in the output file's directory and referenced by filename
in the markdown. Always provide -o with --images; without it, images are
written to the current directory.
After extraction, check that:
--images was requested.If the document was processed in parts, retain the original PDF and page ranges with the extracted files so their order and provenance remain clear.
To extract pages 21-30 as displayed in a PDF viewer, start at zero-based page
20 and request 10 pages:
Save the markdown under the user's dataset or document directory, then use the
asta-documents or local-paper-index skill to index it. Preserve a link to
the source PDF in the index metadata when possible.
Confirm that the user is logged in to Asta, then retry the same command.
The command reports the HTTP status and response from the service. Verify the
network connection and retry. If a large request repeatedly fails, reduce
--max-pages and process smaller page ranges.
--start-page is within the document.Run the command with both --images and an explicit -o path. Images are saved
next to the output file; if no output file is provided, they are saved in the
current directory.
Use PDF extraction when:
Do not use it when:
asta pdf-extraction remote document.pdf -o document.mdasta pdf-extraction remote document.pdf \
--start-page 0 --max-pages 50 -o document-part-001.md
asta pdf-extraction remote document.pdf \
--start-page 50 --max-pages 50 -o document-part-002.md
asta pdf-extraction remote document.pdf \
--start-page 100 --max-pages 50 -o document-part-003.mdasta pdf-extraction remote document.pdf \
-o extracted/document.md \
--imagesasta pdf-extraction remote scanned-document.pdf \
-o extracted/scanned-document.mdasta pdf-extraction remote report.pdf \
--start-page 20 \
--max-pages 10 \
-o report-pages-21-30.md