Files
SnapOtter/apps/docs/tools/pdf-to-text.md
T
SnapOtterandGitHub ecc0a45f0e docs: per-tool reference pages for all 157 tools (+ fix docs build) (#261)
* fix(docs): keep gray-matter on js-yaml 3 so the docs site builds

The js-yaml >=4.2.0 override from #257 forced js-yaml 4 onto gray-matter (used by vitepress and vitepress-plugin-llms), which calls the removed yaml.safeLoad and broke `vitepress build`. Scope a gray-matter>js-yaml ^3.14.1 override so gray-matter keeps the v3 API (build-time, trusted frontmatter only) while app code stays on js-yaml 4.2.0+.

* docs: add per-tool reference pages for all 157 tools, with a modality sidebar

Generate /tools/<id> pages for the 104 tools that lacked one (video 29, audio 17, document 36, data 10, and 12 newer image tools), matching the existing page format (API endpoint, parameters from the OpenAPI spec, curl example, response, notes). Async/AI tools document the 202+SSE flow and feature-bundle requirement.

Sidebar: add Video / Audio / PDF & Documents / Data groups with per-tool links, fold the 12 new image tools into the existing image categories, and replace the placeholder rest.md-anchor group. Docs site builds cleanly (157 pages, no dead links).
2026-06-17 11:11:37 +08:00

1.1 KiB

description
description
Extract plain text from a PDF.

PDF to Text

Extract all readable plain text from a PDF document into a text file.

API Endpoint

POST /api/v1/tools/pdf-to-text

Accepts multipart form data with a PDF file.

Parameters

This tool has no configurable parameters. Upload a PDF and its text content will be extracted.

Example Request

curl -X POST http://localhost:1349/api/v1/tools/pdf-to-text \
  -H "Authorization: Bearer si_your-api-key" \
  -F "file=@report.pdf"

Example Response

{
  "jobId": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
  "downloadUrl": "/api/v1/download/a1b2c3d4-e5f6-7890-abcd-ef1234567890/report.txt",
  "originalSize": 520000,
  "processedSize": 14300,
  "chars": 14300
}

Notes

  • Accepted input format: .pdf.
  • This is a fast (synchronous) tool that returns the result directly.
  • The chars field in the response indicates the number of characters extracted.
  • Only digitally embedded text is extracted. For scanned documents or image-based PDFs, use the PDF OCR tool instead.