Text, Pixels, or Both? Evaluating Input Representations for Multimodal Document QA

cs.AI updates on arXiv.org · 1d ago
Research Papers

arXiv:2609.22628v1 Announce Type: new Abstract: Every document QA system begins with a choice that is rarely studied on its own: whether to feed the model page images, extracted text, or both. We isolate this choice, holding the prompt, judge, and scoring pipeline fixed, across four commercial model endpoints, two corpora, and two context regimes (gold evidence pages and the full document). On…

Read original article on cs.AI updates on arXiv.org →