papersTODAY 04:00 UTC
arXiv paper examines robustness, cost and governance trade-offs in VLM document extraction
A new arXiv preprint argues that evaluations of vision-language models for extracting structured fields from business documents focus too heavily on accuracy against clean benchmarks. The authors propose assessing approaches along additional dimensions such as robustness, cost, and governance considerations, aiming to help practitioners pick a method suited to a given task complexity. No specific model or tool is released with the work.