Tool and Model Comparison

Interactive comparison for OCR, classification, layout, bounding boxes, structured output and deployment.

10/10 = top within this comparison category · 8–9/10 = very strong · 6–7/10 = situational · 0/10 / – = not applicable
Tool / ModelOCRDocument ClassificationLayout / TablesBounding BoxesJSON / Structured OutputLocal / Self-hostedSpeed / EfficiencyAdaptability to defined document classesProduction ReadinessOverall for
Gemini 2.5 Flash-Lite8/1010/108/108/1010/100/1010/1010/1010/108.5/10
Gemini 3.5 Flash-Lite9/1010/109/106/1010/100/1010/1010/1010/108.5/10
Qwen3-VL 4B8/108/108/107/109/1010/109/1010/108/108.5/10
Qwen3-VL 8B9/109/109/107/109/1010/108/1010/108/109/10
Qwen3-VL 30B-A3B9/1010/10*9/108/109/1010/108/1010/108/109.5/10
Qwen3-VL 32B9/1010/10*10/108/109/1010/106/1010/108/109.5/10
Qwen3-VL 235B-A22B10/1010/10*10/109/109/109/103/1010/108/108.5/10
Gemma 4 E2B7/108/107/106/109/1010/1010/1010/108/108/10
Gemma 4 E4B8/109/108/107/109/1010/109/1010/108/108.5/10
Azure Document Intelligence Read/Layout10/105/1010/1010/1010/101/109/106/1010/108.5/10
Azure Custom Classifier5/1010/108/106/1010/101/109/109/1010/108.5/10
PP-OCRv6 Medium10/102/105/1010/108/1010/1010/104/109/109/10 as OCR
PaddleOCR-VL10/107/1010/109/109/1010/108/108/108/109/10
PP-StructureV39/105/1010/109/1010/1010/108/107/109/108.5/10
Docling7/105/1010/108/1010/1010/107/107/109/108/10
MinerU7/104/1010/108/109/1010/109/106/108/107.5/10
Xberg7/107/109/109/1010/1010/105/10 in test9/107/107.5/10
Tesseract6/101/103/108/105/1010/1010/104/1010/106.5/10
OpenCV10/1010/1010/1010/1010/10 as Preprocessing
EmbeddingGemma0/108/10 for text0/100/109/1010/1010/1010/108/108/10 as Pre-Classifier
PaddleX Custom Classifier0/1010/10 potential5/100/109/1010/1010/1010/109/109/10 with enough training data
Note: Qwen3-VL 30B/32B/235B classification scores are capability-based candidates and should still be validated against the golden dataset. 10/10 means top within this comparison, not universally perfect.