OCR

Read text from scans, photos and PDFs with Google Vision, AWS Textract, Azure Document Intelligence, Mistral, Tesseract or your own service.

Text recognition works the same way as storage: one shape for results, and a provider you can swap. Every engine returns an OcrResult:

interface OcrResult {
  text: string // all text, pages separated by blank lines
  pages: { number: number; text: string; lines?: { text: string; confidence?: number }[] }[]
  confidence?: number // 0 to 1
  fields?: Record<string, { value: string; confidence?: number }> // merchant, date, total…
  markdown?: string // Mistral keeps headings and tables
  provider: string
}
"use client"

import * as React from "react"
import { tesseractOcr } from "@uploadcn/core"

import { Button } from "@/components/ui/button"
import { OcrUpload } from "@/components/ocr-upload"

/**
 * Real OCR in the browser with Tesseract: no server and no key, the file
 * never leaves this page. In your app, `ocrEndpoint("/api/ocr")` sends it
 * to Google Vision, Textract, Azure, Mistral or your own service instead.
 */
const recognize = tesseractOcr({ load: () => import("tesseract.js") })

/** Draws a receipt so there is something to read without your own file. */
async function sampleReceipt() {
  const canvas = document.createElement("canvas")
  canvas.width = 560
  canvas.height = 640
  const context = canvas.getContext("2d")!
  context.fillStyle = "#ffffff"
  context.fillRect(0, 0, canvas.width, canvas.height)
  context.fillStyle = "#111111"
  const lines: [string, number, string?][] = [
    ["BLUE BOTTLE COFFEE", 34, "bold"],
    ["300 Webster St, Oakland", 22],
    ["2026-03-14  09:42", 22],
    ["", 22],
    ["Cappuccino            4.75", 24],
    ["Almond croissant      5.25", 24],
    ["Cold brew 16oz        4.50", 24],
    ["", 22],
    ["Subtotal             14.50", 24],
    ["Tax                   1.31", 24],
    ["TOTAL                15.81", 30, "bold"],
    ["", 22],
    ["Thank you!", 24],
  ]
  let y = 70
  for (const [text, size, weight] of lines) {
    context.font = `${weight ?? "normal"} ${size}px monospace`
    context.fillText(text, 40, y)
    y += size + 18
  }
  const blob = await new Promise<Blob>((resolve) =>
    canvas.toBlob((value) => resolve(value!), "image/png")
  )
  return new File([blob], "receipt.png", { type: "image/png" })
}

export default function OcrUploadExample() {
  const root = React.useRef<HTMLDivElement>(null)

  const trySample = async () => {
    const input = [
      ...(root.current?.querySelectorAll<HTMLInputElement>(
        'input[type="file"]'
      ) ?? []),
    ][0]
    if (!input) return
    const transfer = new DataTransfer()
    transfer.items.add(await sampleReceipt())
    input.files = transfer.files
    input.dispatchEvent(new Event("change", { bubbles: true }))
  }

  return (
    <div ref={root} className="flex w-full max-w-2xl flex-col gap-3">
      <OcrUpload
        recognize={recognize}
        accept="image/png,image/jpeg,image/webp"
      />
      <div className="flex items-center justify-between gap-3 text-xs text-muted-foreground">
        <span>
          Runs Tesseract in your browser. The first read downloads the engine.
        </span>
        <Button variant="outline" size="sm" onClick={() => void trySample()}>
          Try a sample receipt
        </Button>
      </div>
    </div>
  )
}

Pick an engine

EngineRunsGood atFields
TesseractBrowserPrivate documents, no backend, no cost-
Google Cloud VisionServerPhotos, handwriting, 50+ languages-
AWS TextractServerForms, receipts and invoices on AWSReceipts, invoices
Azure Document IntelligenceServerReceipts, invoices, IDs, custom modelsYes
Mistral OCRServerLong PDFs, tables, Markdown for LLMs-
Your own serviceServerPaddleOCR, docTR, EasyOCR, internal modelsYours

How it fits

  1. Uploading
  2. Reading text
  3. Success

OCR runs as the uploader's process step, after the file is stored. It doesn't care where the file went, so it works with every adapter, including localAdapter.

import { ocrEndpoint, ocrProcess } from "@uploadcn/core"

<Upload adapter={adapter} process={ocrProcess(ocrEndpoint("/api/ocr"))} />
// later: item.meta.ocr is the OcrResult

The OCR Upload block does this for you and shows the text, confidence and fields. Receipt Upload takes the same recognize prop and fills in merchant, date and amount.

On the server

API keys must never reach the browser, so server engines sit behind a route. Add it:

npx shadcn@latest add @uploadcn/ocr-route

That creates app/api/ocr/route.ts, which picks a provider from your environment variables. Or write it yourself:

app/api/ocr/route.ts
import { createOcrRoute, googleVisionOcr, OcrRouteError } from "@uploadcn/server/ocr"

export const { POST } = createOcrRoute({
  provider: googleVisionOcr({ apiKey: process.env.GOOGLE_VISION_API_KEY! }),
  maxFileSize: 20 * 1024 * 1024,
  async authorize({ request }) {
    const session = await auth(request)
    if (!session) throw new OcrRouteError("Unauthorized", 401)
    return session
  },
})

Then point the component at it:

<OcrUpload recognize={ocrEndpoint("/api/ocr")} />

The route:

  • authorizes before it reads the body,
  • accepts PNG, JPEG, WebP, TIFF, GIF, BMP and PDF; SVG is refused (it can carry script),
  • enforces maxFileSize from content-length and again on the actual file,
  • hides provider errors (they can contain account details) and logs them instead.

Read files that are already stored

Pass storage and clients can send { "key": "…" } instead of the file. Check that the caller owns the key in authorize:

createOcrRoute({
  provider,
  storage: s3Storage({ /* … */ }),
  async authorize({ request, key }) {
    const user = await requireUser(request)
    if (key && !key.startsWith(`${user.id}/`)) throw new OcrRouteError("Forbidden", 403)
  },
})

Providers

Google Cloud Vision

import { googleVisionOcr } from "@uploadcn/server/ocr"

googleVisionOcr({
  apiKey: process.env.GOOGLE_VISION_API_KEY!, // restrict the key to the Vision API
  languageHints: ["en"],
})

Uses DOCUMENT_TEXT_DETECTION. Images go to images:annotate; PDFs, TIFFs and GIFs to files:annotate (the first 5 pages). Prefer a service account? Pass accessToken: () => getAccessToken() instead of apiKey. The key is sent in a header, never in the URL.

In the browser

Tesseract runs in a Web Worker. Nothing is uploaded for OCR and there's no key, which suits private documents and offline apps. Install it, then:

npm install tesseract.js
import { tesseractOcr } from "@uploadcn/core"

const recognize = tesseractOcr({
  load: () => import("tesseract.js"), // loaded on first use
  lang: ["eng", "deu"],
})

<OcrUpload recognize={recognize} accept="image/png,image/jpeg,image/webp" />

Tesseract reads images, not PDFs, and is slower and less accurate than the cloud engines on photos. By default it downloads its engine and language data from a CDN; to self-host them, pass workerOptions: { workerPath, corePath, langPath }.

Without the block

ocrProcess stores the result on item.meta.ocr; render it however you like:

const uploader = useUploader({
  adapter,
  process: ocrProcess(ocrEndpoint("/api/ocr"), { required: true }),
})

With required: true, a file whose text can't be read is rejected; by default it is kept and item.meta.ocrError explains why.

On this page