npx skills add ...
npx skills add aidenwu0209/paddleocr-skills --skill paddleocr-text-recognition
Use this skill whenever the user wants text extracted from images, photos, scans, screenshots, or scanned PDFs. Returns exact machine-readable strings with line-level text and optional bbox coordinates. Strong accuracy for CJK, small print, and handwritten text. Trigger terms: OCR, 文字识别, 图片转文字, 截图识字, 提取图中文字, 扫描识字, 识字, 纯文字, plain text extraction, 坐标, 检测框, bbox, bounding box, image to text, screenshot, photo scan, recognize text.
npx skills add aidenwu0209/paddleocr-skills --skill paddleocr-text-recognition
Trigger keywords (routing): Bilingual trigger terms (Chinese and English) are listed in the YAML description above—use that field for discovery and routing.
Use this skill for:
Do not use for:
Scripts declare their dependencies inline (PEP 723). No separate install step is needed — uv resolves dependencies automatically:
Working directory: All
uv run scripts/...commands below should be run from this skill's root directory (the directory containing this SKILL.md file).
Identify the input source:
--file-url parameter--file-path parameterExecute OCR:
Or for local files:
Performance note: Parsing time scales with document complexity. Single-page images typically complete in 1-3 seconds; large PDFs (50+ pages) may take several minutes. Allow adequate time before assuming a timeout.
Default behavior: save raw JSON to a temp file:
--output is omitted, the script saves automatically under the system temp directory<system-temp>/paddleocr/text-recognition/results/result_<timestamp>_<id>.json--output is provided, it overrides the default temp-file destination--stdout is provided, JSON is printed to stdout and no file is savedResult saved to: /absolute/path/...--stdout only when you explicitly want to skip file persistenceParse JSON response:
ok field: true means success, false means errortext field contains all recognized text--stdout is used, parse the stdout JSON directlyok is false, display error.messagePresent results to user:
Common next steps once you have the recognized text:
text field to a .txt or .md filetext field is clean plain text, ready for downstream processingAlways display the COMPLETE recognized text to the user. The user typically needs the full content for downstream use — truncation silently loses data they may not notice is missing.
text field, no matter how longExample - Correct:
Example - Incorrect:
The script returns a JSON envelope with ok, text, result, and error fields. Use text for the recognized content; result contains the raw API response for debugging.
For the full schema and field-level details, see references/output_schema.md.
Raw result location (default): the temp-file path printed by the script on stderr
This mirror keeps the bundled scripts/ocr_caller.py as the default path. The upstream PaddleOCR project (since PR #18090, 2026-06-03) also ships an official CLI that calls the same API directly. If the paddleocr package is installed, you can use it as a drop-in alternative — no uv run or local scripts required.
Install (one-time):
Environment: the CLI only needs PADDLEOCR_ACCESS_TOKEN. It resolves the API endpoint internally, so PADDLEOCR_OCR_API_URL is not required when using the CLI (the URL is still required by the script).
Basic OCR:
Common options:
CLI output format — different from the script envelope:
The CLI prints {jobId, pages:[...]} to stdout. It does not wrap the response in the script's {ok, text, result, error} envelope, does not auto-save to a temp file, and does not concatenate text for you. If you switch paths, update your parsing logic accordingly.
Scripts vs CLI — at a glance:
| Scripts (default) | paddleocr CLI (alternative) | |
|---|---|---|
| Install | uv resolves PEP 723 inline deps | pip install "paddleocr>=3.7.0" |
| Required env | PADDLEOCR_OCR_API_URL + PADDLEOCR_ACCESS_TOKEN | PADDLEOCR_ACCESS_TOKEN only |
| Entry | uv run scripts/ocr_caller.py ... | paddleocr api --model_type ocr ... |
| Output | {ok, text, result, error} envelope, auto-saved to temp file | {jobId, pages:[...]} to stdout |
| Result location | Path printed on stderr (or --output/--stdout) | stdout (or --output) |
| Best for | Skills runtimes, offline-friendly, no extra install | Already have paddleocr installed, want the upstream-canonical flow |
Run paddleocr api --help for the full option list.
Example 1: URL OCR
Example 2: Local File OCR
Example 3: OCR With Explicit File Type
--file-type 0: PDF--file-type 1: image.pdf, .png, .jpg, .jpeg, .bmp, .tiff, .tif, .webp) is required; otherwise pass --file-type explicitly. For URLs with unrecognized extensions, the service attempts inference.Example 4: Print JSON Without Saving
When API is not configured, the script outputs:
Configuration workflow:
Show the exact error message to the user.
Guide the user to obtain credentials: Visit the PaddleOCR website, click API, select the PP-OCRv5 model, select the language, then copy the API_URL and Token. They map to these environment variables:
PADDLEOCR_OCR_API_URL — full endpoint URL ending with /ocrPADDLEOCR_ACCESS_TOKEN — 40-character alphanumeric stringOptionally configure PADDLEOCR_OCR_TIMEOUT for request timeout. Recommend using the host application's standard configuration method rather than pasting credentials in chat.
Apply credentials — one of:
All errors return JSON with ok: false. Show the error message and stop — do not fall back to your own vision capabilities. Identify the issue from error.code and error.message:
Authentication failed (403) — error.message contains "Authentication failed"
Quota exceeded (429) — error.message contains "API rate limit exceeded"
Unsupported format — error.message contains "Unsupported file format"
No text detected:
text field is emptyIf recognition quality is poor:
result.result.ocrResults[n].prunedResult.rec_scores) shows per-line confidence scores — low values identify uncertain regions worth reviewingreferences/output_schema.md — Full output schema, field descriptions, and command examplesNote: Model version, capabilities, and supported file formats are determined by your API endpoint (
PADDLEOCR_OCR_API_URL) and its official API documentation.
To verify the skill is working properly:
The first form tests configuration and API connectivity. --skip-api-test checks configuration only. --test-url overrides the default sample image URL.