Tool and Model Comparison — benchmark-informed update

OCR, classification, layout, bounding boxes, structured output, deployment and operational readiness · Updated 24 August 2026

Independent local test results

98.3%OCR completion · 176/179
74.3%Historical end-to-end completion · 133/179
100%Latest stabilized runs · 53/53

Five full runs are included. Documents overlap between runs, so 179 is a processing-event count, not 179 unique files. Weighted model confidence for successful classifications is about 92.0%; confidence is not ground-truth accuracy.

Benchmark-informed overall score

Value axis: overall score (/10) · Category axis: tool/model · utilities excluded · sources current to 24 August 2026

Detailed comparison

10/10 = strongest fit within this comparison, not universal perfect accuracy · 8–9 = very strong · 6–7 = situational · – = not applicable

Tool / ModelOCRDocument
Classification
Layout /
Tables
Bounding
Boxes
JSON / Structured
Output
Local /
Self-hosted
Speed /
Efficiency
AdaptabilityProduction
Readiness
Overall

Interpretation and corrections

The previous overall values were not consistently reproducible from the category values. Here, every overall value is calculated automatically. OpenCV, EmbeddingGemma and PaddleX Custom Classifier are rated only for the role they can actually perform.

The current local test stack scores 8.7/10 as a practical local pipeline. Its main remaining deduction is the absence of expert-confirmed labels and production-scale monitoring, not OCR completion.

Sources and evidence scope