Bumps every version surface to 2.2.0, fixes a latent version-coupling bug in the OCR runtime tests, and stops an absent GPU runner from silently stalling a release. Version surfaces: scripts/sync-version.sh covers the 11 workspaces, APP_VERSION, and the docs release commands across all locales. Root package.json plus the three surfaces the script never reaches are done by hand: the DOCKERHUB.md banner and tag table, the docker-tags.md pinning table in 21 locales, and the example runtimeVersion in tools/image/ocr.md in 21 locales. The release-notes archive step is deliberately not pre-run, so the notes text stays editable until the release. Latent bug: runtime-state rejects any runtime whose compatibility.snapotterVersion is not exactly APP_VERSION, and five fixtures pinned the literal 2.1.0. Since semantic-release rewrites APP_VERSION on every release, the first PR after any bump would have gone red for a reason nobody would trace to the release. The fixtures now derive from APP_VERSION. GPU runner: sign-ocr-index needs verify-ocr-nvidia on self-hosted hardware, and the gated manifest job needs ai-bundles, so a missing runner queued instead of failing and produced no image tags. preflight-gpu-runner claims the same labels with no dependencies, so it is scheduled first and validates the GPU before the 90-minute build. An API preflight is impossible because listing self-hosted runners needs Administration:read, which GITHUB_TOKEN cannot hold, so RELEASE.md carries the maintainer-side check.
7.7 KiB
description, i18n_output_hash, i18n_source_hash, i18n_provenance
| description | i18n_output_hash | i18n_source_hash | i18n_provenance |
|---|---|---|---|
| 組み込みの Tesseract またはオプションの高精度 RapidOCR ランタイムを使用して、画像からテキストをローカルに抽出します。 | 360daff93230 | 0d453b49db02 | human |
OCR / テキスト抽出
画像を外部サービスに送信せずに、画像からテキストを抽出します。組み込みの fast 層は、Tesseract を使用します。オプションの balanced 層と best 層は、固定された PP-OCR ONNX モデルで RapidOCR を使用します。
::: info 韓国語 OCR の互換性
高速 OCR は auto、en、de、es、fr、zh、ja に対応しますが、韓国語 (ko) には対応しません。韓国語には高精度 OCR パックと balanced または best が必要です。パックは公式 Linux amd64/arm64 コンテナで動作し、NVIDIA ホストでも OCR は CPU 上で実行されます。非対応システムでは明示的な互換性エラーを返し、暗黙に fast へ切り替えません。韓国語で fast または旧 tesseract エイリアスを指定すると、キュー投入前に FEATURE_INCOMPATIBLE と fast-korean-unsupported で拒否されます。
:::
API エンドポイント
POST /api/v1/tools/image/ocr
処理: OCR は常に非同期で実行されます。検証してキューに追加した後、エンドポイントは jobId とともに直ちに 202 Accepted を返します。ジョブの SSE 進行ストリームを終端の complete または failed イベントまで追跡してください。成功イベントの result に OCR フィールドが含まれます。
正確な OCR パック: オプションの ocr ランタイム (ターゲットに応じて、約 208 ~ 234 の MiB をダウンロードし、409 ~ 488 の MiB をインストールします)。 fast にはこのパックは必要ありません。インストーラーは、署名付きインデックスによって制限される正確なサイズを検証します。
パラメーター
| パラメーター | 型 | 必須 | デフォルト | 説明 |
|---|---|---|---|---|
| file | file | はい | - | 画像ファイル (マルチパート)、最大 512 MiB エンコードおよび 40 メガピクセルのデコード。オペレータのアップロード制限の下限は引き続き適用されます |
| quality | string | いいえ | 動的 | 品質レベル: fast (Tesseract)、balanced (小型 PP-OCRv6 モデルを備えた RapidOCR)、または best (校正されたバリアント スコアリングを備えた高精度の中型 PP-OCRv6 モデル) |
| language | string | No | "auto" |
言語ヒント: auto, en, de, fr, es, zh, ja, ko |
| enhance | boolean | いいえ | ティアに依存 | 認識前に局所的なコントラストを改善します。高速ではそれを直接適用します。 Balanced および Best は、調整されたスコアによって結果が改善された場合にのみバリアントを保持します。デフォルトは、best の場合は true、fast/balanced の場合は false です。 |
| engine | string | いいえ | - | 非推奨の互換性エイリアス。代わりに quality を使用してください。 tesseract は fast にマップされます。従来の paddleocr 値は balanced にマップされますが、PaddlePaddle はロードされません |
quality と engine を省略すると、SnapOtter は best、balanced、fast の順で利用可能な最上位の層を選びます。韓国語では fast を選択せず、best、次に balanced を使用し、どちらもなければ高精度ランタイムのインストールまたは互換性エラーを返します。
リクエスト例
curl -X POST http://localhost:1349/api/v1/tools/image/ocr \
-F "file=@document.png" \
-F 'settings={"quality":"best","language":"en","enhance":true}'
受付レスポンス(202)
{
"jobId": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
"async": true
}
進捗と結果(SSE)
202 レスポンスで返された jobId(または指定した clientJobId)を使って GET /api/v1/jobs/{jobId}/progress に接続します。終端の complete または failed イベントまでストリームを開いたままにしてください。成功した終端フレームでは、result に OCR 出力が含まれます。
{
"jobId": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
"type": "single",
"phase": "complete",
"stage": "complete",
"percent": 100,
"result": {
"jobId": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
"downloadUrl": "/api/v1/download/a1b2c3d4-e5f6-7890-abcd-ef1234567890/document_ocr.txt",
"originalSize": 12345,
"processedSize": 47,
"text": "Extracted text content from the image...",
"engine": "rapidocr-onnx",
"requestedQuality": "best",
"actualQuality": "best",
"device": "cpu",
"provider": "CPUExecutionProvider",
"degraded": false,
"warnings": [],
"runtimeVersion": "2.2.0",
"modelVersion": "PP-OCRv6-best-v1-medium"
}
}
処理エラーは終端の failed イベントの error フィールドで通知されます。キューへの追加後に HTTP 422 として返されることはありません。
注意事項
fastは、サポートされている SnapOtter イメージで常に使用できます。balancedおよびbestには、オプションの正確な OCR パックが必要です。- 内蔵の Tesseract は、公式イメージに約 25 個の MiB を追加します。正確なパックはイメージに焼き付けられるのではなく、
/data/aiに保存されます。 - 公式 Linux amd64 および arm64 コンテナー用に正確なパックが公開されています。 NVIDIA ホストを含む ONNX Runtime の CPU プロバイダーを意図的に使用するため、CUDA ライブラリや GPU の互換性に依存しません。 ソースおよびビルド済みの bare-metal インストールは、独自の互換性のあるランタイムを提供しない限り、高速 OCR を使用します。
- 成功した終端の
resultには、textの抽出テキストとdownloadUrlのダウンロード可能な.txtアーティファクトの両方が含まれます。 - SnapOtter は、明示的に要求された層を尊重します。
balancedまたはbestが使用できない場合、API は、FEATURE_NOT_INSTALLEDまたはFEATURE_INCOMPATIBLEを含む501を返します。リクエストを黙って別の層にダウングレードすることはありません。 - 成功した空の結果は空の結果のままになります。実行時にエラーが発生した場合は、低品質のエンジンで再試行するのではなく、エラーが返されます。
- 成功した終端の
resultでは、requestedQualityとactualQualityの両方に加えて、エンジン、デバイス、プロバイダー、ランタイムとモデルのバージョン、および警告が報告されます。 - HEIC/HEIF、RAW、TGA、PSD、EXR、HDR の入力フォーマットを自動デコードでサポートします。
- オーバーサイズのエンコードされた入力は、
413を返します。 40 メガピクセルを超える画像と、制限された出力制限を超える OCR 応答は、部分的に処理される代わりに拒否されます。