Japanese Document OCR — Vertical Text, Tables and Formulas
Reads Japanese and English documents and returns markdown, with tables as HTML and formulas as LaTeX. Handles vertical Japanese text, which the other OCR models in the catalog do not. 4B VLM from SB Intuitions, MIT licensed. Japanese and English only.
What it does
Sarashina2.2-OCR is a 4-billion-parameter vision-language model built for Japanese document reading. Give it a page image and it returns two things: the extracted text as markdown, and an image with the detected figures and charts marked.
It is the only model here that reads vertical Japanese text (縦書き), and the structure it recovers goes beyond plain text — tables come back as HTML, formulas as LaTeX. That makes it the right pick for Japanese contracts, technical papers, textbooks, and financial statements, where the layout carries meaning.
Japanese and English only. It will not read Korean, Chinese, or anything else.
About the image output: it marks figures and charts only, not every text block. Visual components are located with bounding boxes on a normalized 0–1000 coordinate range. If you want the full page layout boxed, use GLM-OCR instead.
Problem it solves
Japanese documents | Contracts, filings, manuals, textbooks — including vertical text |
Tables | Returned as HTML, ready to parse downstream instead of retyping |
Formulas | Returned as LaTeX from technical papers and textbooks |
Figure & chart detection | Marked on the output image so you can crop or count them |
Markdown output | Headings and structure preserved, not a flat text dump |
Input/Output
- Input: A document page image
- Format: JPG, PNG, WEBP, AVIF
- Content: Any document page — contracts, papers, forms, textbook pages, statements
- Language: Japanese or English only
- Output: Two ports
- Image — the page with detected figures and charts marked, using normalized 0–1000 bounding boxes. Not a full layout box map; GLM-OCR does that
- Text — extracted text as markdown, with tables as HTML and formulas as LaTeX
Choosing an OCR model
Your input | Use |
Japanese documents — vertical text, tables, formulas | Sarashina2.2-OCR |
Korean, Chinese or English printed text, high volume | PaddleOCR |
Full-page layout understanding, general multi-language documents | GLM-OCR |
Handwriting | DeepSeek OCR |
Performance
- Latency: 7–20 seconds per page, depending on how dense the page is
- Strongest on: printed Japanese documents with real structure — tables, columns, vertical text, formulas
- Weaker on: headers and footers, and math on old scanned pages — these are the model's two lowest olmOCR-bench categories. If your pages are aged scans with formulas, check a sample before running a batch
- Not for: Korean or Chinese (use PaddleOCR or GLM-OCR), handwriting (use DeepSeek OCR), or photos of casual text like signs and menus, where PaddleOCR is faster and sufficient
Technical Details
Architecture | 4B vision-language model (3B language component + vision encoder) |
Developer | SB Intuitions |
Languages | Japanese, English |
Structure recovery | Markdown text · tables as HTML · formulas as LaTeX · vertical text · figure boxes on a normalized 0–1000 range |
Runtime | GPU, cloud — no local setup |
Compliance & Provenance
Provider | SB Intuitions |
Provider type | Specialized |
License | MIT |
EU AI Act risk class | Minimal Risk |
Art. 50 transparency | OCR output is labeled "AI-extracted", not "AI-generated" |
Region availability | Available globally |
Training data summary | Pending — provider has not yet published per Art. 53(d) |
For more on how we classify models and mark outputs, see our AI Policy.
Model Source
- Model card
- License: MIT