Read text from scans, photos and PDFs with Google Vision, AWS Textract, Azure Document Intelligence, Mistral, Tesseract or your own service.
Text recognition works the same way as storage: one shape for results, and a provider
you can swap. Every engine returns an OcrResult:
interface OcrResult {
text: string // all text, pages separated by blank lines
pages: { number: number; text: string; lines?: { text: string; confidence?: number }[] }[]
confidence?: number // 0 to 1
fields?: Record<string, { value: string; confidence?: number }> // merchant, date, total…
markdown?: string // Mistral keeps headings and tables
provider: string
}"use client"
import * as React from "react"
import { tesseractOcr } from "@uploadcn/core"
import { Button } from "@/components/ui/button"
import { OcrUpload } from "@/components/ocr-upload"
/**
* Real OCR in the browser with Tesseract: no server and no key, the file
* never leaves this page. In your app, `ocrEndpoint("/api/ocr")` sends it
* to Google Vision, Textract, Azure, Mistral or your own service instead.
*/
const recognize = tesseractOcr({ load: () => import("tesseract.js") })
/** Draws a receipt so there is something to read without your own file. */
async function sampleReceipt() {
const canvas = document.createElement("canvas")
canvas.width = 560
canvas.height = 640
const context = canvas.getContext("2d")!
context.fillStyle = "#ffffff"
context.fillRect(0, 0, canvas.width, canvas.height)
context.fillStyle = "#111111"
const lines: [string, number, string?][] = [
["BLUE BOTTLE COFFEE", 34, "bold"],
["300 Webster St, Oakland", 22],
["2026-03-14 09:42", 22],
["", 22],
["Cappuccino 4.75", 24],
["Almond croissant 5.25", 24],
["Cold brew 16oz 4.50", 24],
["", 22],
["Subtotal 14.50", 24],
["Tax 1.31", 24],
["TOTAL 15.81", 30, "bold"],
["", 22],
["Thank you!", 24],
]
let y = 70
for (const [text, size, weight] of lines) {
context.font = `${weight ?? "normal"} ${size}px monospace`
context.fillText(text, 40, y)
y += size + 18
}
const blob = await new Promise<Blob>((resolve) =>
canvas.toBlob((value) => resolve(value!), "image/png")
)
return new File([blob], "receipt.png", { type: "image/png" })
}
export default function OcrUploadExample() {
const root = React.useRef<HTMLDivElement>(null)
const trySample = async () => {
const input = [
...(root.current?.querySelectorAll<HTMLInputElement>(
'input[type="file"]'
) ?? []),
][0]
if (!input) return
const transfer = new DataTransfer()
transfer.items.add(await sampleReceipt())
input.files = transfer.files
input.dispatchEvent(new Event("change", { bubbles: true }))
}
return (
<div ref={root} className="flex w-full max-w-2xl flex-col gap-3">
<OcrUpload
recognize={recognize}
accept="image/png,image/jpeg,image/webp"
/>
<div className="flex items-center justify-between gap-3 text-xs text-muted-foreground">
<span>
Runs Tesseract in your browser. The first read downloads the engine.
</span>
<Button variant="outline" size="sm" onClick={() => void trySample()}>
Try a sample receipt
</Button>
</div>
</div>
)
}
Pick an engine
| Engine | Runs | Good at | Fields |
|---|---|---|---|
| Tesseract | Browser | Private documents, no backend, no cost | - |
| Google Cloud Vision | Server | Photos, handwriting, 50+ languages | - |
| AWS Textract | Server | Forms, receipts and invoices on AWS | Receipts, invoices |
| Azure Document Intelligence | Server | Receipts, invoices, IDs, custom models | Yes |
| Mistral OCR | Server | Long PDFs, tables, Markdown for LLMs | - |
| Your own service | Server | PaddleOCR, docTR, EasyOCR, internal models | Yours |
How it fits
- Uploading
- Reading text
- Success
OCR runs as the uploader's process step, after the file is stored. It doesn't care
where the file went, so it works with every adapter, including localAdapter.
import { ocrEndpoint, ocrProcess } from "@uploadcn/core"
<Upload adapter={adapter} process={ocrProcess(ocrEndpoint("/api/ocr"))} />
// later: item.meta.ocr is the OcrResultThe OCR Upload block does this for you and shows the text,
confidence and fields. Receipt Upload takes the same
recognize prop and fills in merchant, date and amount.
On the server
API keys must never reach the browser, so server engines sit behind a route. Add it:
npx shadcn@latest add @uploadcn/ocr-routeThat creates app/api/ocr/route.ts, which picks a provider from your environment
variables. Or write it yourself:
import { createOcrRoute, googleVisionOcr, OcrRouteError } from "@uploadcn/server/ocr"
export const { POST } = createOcrRoute({
provider: googleVisionOcr({ apiKey: process.env.GOOGLE_VISION_API_KEY! }),
maxFileSize: 20 * 1024 * 1024,
async authorize({ request }) {
const session = await auth(request)
if (!session) throw new OcrRouteError("Unauthorized", 401)
return session
},
})Then point the component at it:
<OcrUpload recognize={ocrEndpoint("/api/ocr")} />The route:
- authorizes before it reads the body,
- accepts PNG, JPEG, WebP, TIFF, GIF, BMP and PDF; SVG is refused (it can carry script),
- enforces
maxFileSizefromcontent-lengthand again on the actual file, - hides provider errors (they can contain account details) and logs them instead.
Read files that are already stored
Pass storage and clients can send { "key": "…" } instead of the file. Check that the
caller owns the key in authorize:
createOcrRoute({
provider,
storage: s3Storage({ /* … */ }),
async authorize({ request, key }) {
const user = await requireUser(request)
if (key && !key.startsWith(`${user.id}/`)) throw new OcrRouteError("Forbidden", 403)
},
})Providers
Google Cloud Vision
import { googleVisionOcr } from "@uploadcn/server/ocr"
googleVisionOcr({
apiKey: process.env.GOOGLE_VISION_API_KEY!, // restrict the key to the Vision API
languageHints: ["en"],
})Uses DOCUMENT_TEXT_DETECTION. Images go to images:annotate; PDFs, TIFFs and GIFs to
files:annotate (the first 5 pages). Prefer a service account? Pass
accessToken: () => getAccessToken() instead of apiKey. The key is sent in a header,
never in the URL.
In the browser
Tesseract runs in a Web Worker. Nothing is uploaded for OCR and there's no key, which suits private documents and offline apps. Install it, then:
npm install tesseract.jsimport { tesseractOcr } from "@uploadcn/core"
const recognize = tesseractOcr({
load: () => import("tesseract.js"), // loaded on first use
lang: ["eng", "deu"],
})
<OcrUpload recognize={recognize} accept="image/png,image/jpeg,image/webp" />Tesseract reads images, not PDFs, and is slower and less accurate than the cloud engines on
photos. By default it downloads its engine and language data from a CDN; to self-host them,
pass workerOptions: { workerPath, corePath, langPath }.
Without the block
ocrProcess stores the result on item.meta.ocr; render it however you like:
const uploader = useUploader({
adapter,
process: ocrProcess(ocrEndpoint("/api/ocr"), { required: true }),
})With required: true, a file whose text can't be read is rejected; by default it is kept and
item.meta.ocrError explains why.