Self-Hosted OCR in 2026: DeepSeek-OCR on vLLM vs. Template Pipelines

We built two document-intelligence systems this year — DocuVerse and SemaScan — and they took opposite approaches to the same problem. One runs a modern OCR model behind vLLM for beautiful Markdown. The other uses classical OCR plus a template rule engine for structured JSON. Both shipped. Both taught us something the other approach hides.
Approach A: DeepSeek-OCR on vLLM (DocuVerse)
DeepSeek-OCR is a visual-compression OCR model: it reads a page and emits Markdown with tables and formulas. The catch is operational — it cannot be a simple REST call. It needs a GPU with roughly 24 GB of VRAM and runs as a persistent model server behind vLLM.
# docker-compose.yml (excerpt)
ocr:
image: vllm/vllm-openai:latest
command: --model deepseek-ai/DeepSeek-OCR
deploy:
resources:
reservations:
devices: [{ driver: nvidia, count: 1 }]
Around that server, the product work is mostly queue management: upload → Celery job → OCR → store Markdown, with WebSocket progress, retries, and workspaces with RBAC on top. The result is documents you can read like a web page, formulas included. The cost is a GPU server and a 24 GB floor you cannot negotiate down.
Approach B: OCR + templates (SemaScan)
SemaScan assumes OCR output is never deterministic — skewed scans, mixed JP/EN text, layout drift. Instead of chasing a better model, it constrains the problem:
- Preprocess: deskew, denoise, normalize DPI.
- OCR with whichever engine fits (PaddleOCR for JP/EN, Tesseract, EasyOCR, cloud fallbacks).
- Match against a JSON template: region anchors + regex + keyword fallbacks.
- Emit structured JSON with a confidence score per field.
A vendor changes their invoice layout? Ship a new template — no retraining, no GPU. The trade: someone has to write templates, and wildly novel documents fall outside the rules.
Where each one wins
| DeepSeek-OCR + vLLM | OCR + templates | |
|---|---|---|
| Output | Markdown, tables, formulas | Structured JSON, typed fields |
| Hardware | GPU (≈24 GB VRAM) | CPU is fine |
| New document types | Prompt/model behaviour | Write a template |
| Per-field trust | Model confidence (opaque) | Explicit, testable |
| Best for | Reading, search, archives | Receipts, tickets, invoices, ERP feeds |
If your users need to read documents, Approach A is unmatched. If your systems need to act on fields, Approach B gives you something the model can't: a score you can gate on, and a boundary you can test.
The lesson that transfers
Neither project ships raw OCR text to users. DocuVerse renders Markdown in a workspace; SemaScan refuses to emit a field below a confidence threshold. The extraction pipeline is the product — OCR is an ingredient. And both pipelines are queues: background processing, retries, progress, idempotent steps. The document AI part is the easy part; the job orchestration is where projects die.
One more thing we would tell our past selves: benchmark on your own worst documents, not on samples. A 0.93-confidence average means nothing when the 7% is every handwritten receipt in the folder.
