Tool and Model Comparison — benchmark-informed update
OCR, classification, layout, bounding boxes, structured output, deployment and operational readiness · Updated 24 August 2026
Independent local test results
98.3%OCR completion · 176/179
74.3%Historical end-to-end completion · 133/179
100%Latest stabilized runs · 53/53
Five full runs are included. Documents overlap between runs, so 179 is a processing-event count, not 179 unique files. Weighted model confidence for successful classifications is about 92.0%; confidence is not ground-truth accuracy.
Benchmark-informed overall score
Value axis: overall score (/10) · Category axis: tool/model · utilities excluded · sources current to 24 August 2026
Detailed comparison
10/10 = strongest fit within this comparison, not universal perfect accuracy · 8–9 = very strong · 6–7 = situational · – = not applicable
Tool / Model
OCR
Document Classification
Layout / Tables
Bounding Boxes
JSON / Structured Output
Local / Self-hosted
Speed / Efficiency
Adaptability
Production Readiness
Overall
Interpretation and corrections
The previous overall values were not consistently reproducible from the category values. Here, every overall value is calculated automatically. OpenCV, EmbeddingGemma and PaddleX Custom Classifier are rated only for the role they can actually perform.
The current local test stack scores 8.7/10 as a practical local pipeline. Its main remaining deduction is the absence of expert-confirmed labels and production-scale monitoring, not OCR completion.