Files
SnapOtter/apps/docs/tools/pdf/pdf-to-text.md
T
SnapOtter 8532f3227b docs: audit all 157 tool pages against live schemas
Reconcile every tool page's parameters, defaults, and response shape against the tool's Zod settings schema and executionHint in code. Notable fixes: color-palette (add count + format params, hex output, median-cut algorithm), favicon (add 5 params, was documented as having none), qr-generate (add logoDataUri), convert (add ppm/eps/tga formats), video-loudnorm (-16 LUFS not -14), smart-crop (async 202 not sync 200), images-to-video (1080x1080 square), and several output-filename and behavior-note corrections.

Also normalize API endpoint paths to /api/v1/tools/<id> (no modality segment) and standardize curl examples on the Docker API port 1349. Verified with a clean docs build.
2026-06-18 14:12:28 +08:00

1.1 KiB

description
description
Extract plain text from a PDF.

PDF to Text

Extract all readable plain text from a PDF document into a text file.

API Endpoint

POST /api/v1/tools/pdf-to-text

Accepts multipart form data with a PDF file.

Parameters

This tool has no configurable parameters. Upload a PDF and its text content will be extracted.

Example Request

curl -X POST http://localhost:1349/api/v1/tools/pdf-to-text \
  -H "Authorization: Bearer si_your-api-key" \
  -F "file=@report.pdf"

Example Response

{
  "jobId": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
  "downloadUrl": "/api/v1/download/a1b2c3d4-e5f6-7890-abcd-ef1234567890/report.txt",
  "originalSize": 520000,
  "processedSize": 14300,
  "chars": 14300
}

Notes

  • Accepted input format: .pdf.
  • This is a fast (synchronous) tool that returns the result directly.
  • The chars field in the response indicates the number of characters extracted.
  • Only digitally embedded text is extracted. For scanned documents or image-based PDFs, use the PDF OCR tool instead.