Files
SnapOtter/apps/docs/ja/tools/pdf/ocr-pdf.md
T
SnapOtterandGitHub 4963ab3bbd feat(docs-i18n): translate all documentation into 20 languages
All 181 docs markdown files translated into 20 languages (apps/docs/<locale>/**). Companion to the i18n code PR; admin-merged because the file count exceeds GitHub's per-PR CI trigger limit. Validated by pnpm i18n:check (all surfaces, 0 stale/missing) and a clean all-locale docs build.
2026-07-11 13:52:47 +08:00

2.4 KiB

description, i18n_source_hash, i18n_provenance, i18n_output_hash
description i18n_source_hash i18n_provenance i18n_output_hash
AI 搭載の OCR を使って PDF ドキュメントからテキストを抽出します。 1431fcba180b human d39e66166498

PDF OCR

AI 搭載の光学文字認識を使って PDF ドキュメントからテキストを抽出します。複数の品質ティアと言語に対応しています。OCR 機能バンドルのインストールが必要です。

API Endpoint

POST /api/v1/tools/pdf/ocr-pdf

PDF ファイルと、任意の JSON settings フィールドを含む multipart フォームデータを受け付けます。

Parameters

Parameter Type Required Default Description
quality string No "balanced" OCR 品質ティア: fastbalancedbest
language string No "auto" ドキュメントの言語: autoendefreszhjako
pages string No "all" ページ選択。例: "all""1-3""1,3,5"

Example Request

curl -X POST http://localhost:1349/api/v1/tools/pdf/ocr-pdf \
  -H "Authorization: Bearer si_your-api-key" \
  -F "file=@scanned.pdf" \
  -F 'settings={"quality": "best", "language": "en", "pages": "1-5"}'

Example Response

202 Accepted を返します。進捗は /api/v1/jobs/{jobId}/progress の SSE で追跡できます。

{
  "jobId": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
  "async": true
}

Notes

  • 受け付ける入力形式: .pdf
  • これは OCR 機能バンドル のインストールが必要な AI ツールです。バンドルがインストールされていない場合、API は 501 Not Implemented を返します。
  • fast 品質ティアはより軽量なモデルを使って高速に処理します。best は速度を犠牲にしてより高精度なモデルを使用します。
  • auto の言語設定は、ドキュメントの言語を自動的に検出しようとします。
  • 範囲指定("1-3")、カンマ区切りのリスト("1,3,5")、または全ページを対象とする "all" を使って特定のページを対象にできます。
  • すでに選択可能なテキストを含む PDF については、代わりに高速な PDF to Text ツールの使用を検討してください。