
Text, Pixels, or Both? Evaluating Input Representations for Multimodal Document QA
arXiv:2609.22628v1 Announce Type: new Abstract: Every document QA system begins with a choice that is rarely studied on its own: whether to feed the model page images, extracted text, or both. We isolate this choice, holding the prompt, judge, and scoring pipeline fixed, across four commercial model endpoints, two corpora, and two context regimes (gold evidence pages and the full document). On…
Read original article on cs.AI updates on arXiv.org →