📄

Text Recognition (OCR) - Sarashina2.2-OCR

🛠

Japanese Document OCR — Vertical Text, Tables and Formulas

Reads Japanese and English documents and returns markdown, with tables as HTML and formulas as LaTeX. Handles vertical Japanese text, which the other OCR models in the catalog do not. 4B VLM from SB Intuitions, MIT licensed. Japanese and English only.

What it does

Sarashina2.2-OCR is a 4-billion-parameter vision-language model built for Japanese document reading. Give it a page image and it returns two things: the extracted text as markdown, and an image with the detected figures and charts marked.

It is the only model here that reads vertical Japanese text (縦書き), and the structure it recovers goes beyond plain text — tables come back as HTML, formulas as LaTeX. That makes it the right pick for Japanese contracts, technical papers, textbooks, and financial statements, where the layout carries meaning.

Japanese and English only. It will not read Korean, Chinese, or anything else.

About the image output: it marks figures and charts only, not every text block. Visual components are located with bounding boxes on a normalized 0–1000 coordinate range. If you want the full page layout boxed, use GLM-OCR instead.

Problem it solves

Japanese documents
Contracts, filings, manuals, textbooks — including vertical text
Tables
Returned as HTML, ready to parse downstream instead of retyping
Formulas
Returned as LaTeX from technical papers and textbooks
Figure & chart detection
Marked on the output image so you can crop or count them
Markdown output
Headings and structure preserved, not a flat text dump

Input/Output

  • Input: A document page image
    • Format: JPG, PNG, WEBP, AVIF
    • Content: Any document page — contracts, papers, forms, textbook pages, statements
    • Language: Japanese or English only
  • Output: Two ports
    • Image — the page with detected figures and charts marked, using normalized 0–1000 bounding boxes. Not a full layout box map; GLM-OCR does that
    • Text — extracted text as markdown, with tables as HTML and formulas as LaTeX

Choosing an OCR model

Your input
Use
Japanese documents — vertical text, tables, formulas
Sarashina2.2-OCR
Korean, Chinese or English printed text, high volume
PaddleOCR
Full-page layout understanding, general multi-language documents
GLM-OCR
Handwriting
DeepSeek OCR

Performance

  • Latency: 7–20 seconds per page, depending on how dense the page is
  • Strongest on: printed Japanese documents with real structure — tables, columns, vertical text, formulas
  • Weaker on: headers and footers, and math on old scanned pages — these are the model's two lowest olmOCR-bench categories. If your pages are aged scans with formulas, check a sample before running a batch
  • Not for: Korean or Chinese (use PaddleOCR or GLM-OCR), handwriting (use DeepSeek OCR), or photos of casual text like signs and menus, where PaddleOCR is faster and sufficient

Technical Details

Architecture
4B vision-language model (3B language component + vision encoder)
Developer
SB Intuitions
Languages
Japanese, English
Structure recovery
Markdown text · tables as HTML · formulas as LaTeX · vertical text · figure boxes on a normalized 0–1000 range
Runtime
GPU, cloud — no local setup

Compliance & Provenance

Provider
SB Intuitions
Provider type
Specialized
License
MIT
EU AI Act risk class
Minimal Risk
Art. 50 transparency
OCR output is labeled "AI-extracted", not "AI-generated"
Region availability
Available globally
Training data summary
Pending — provider has not yet published per Art. 53(d)

For more on how we classify models and mark outputs, see our AI Policy.

Model Source

  • Model card