Improving OCR Faithfulness via Gated and Attenuated On-Policy Distillation

cs.AI updates on arXiv.org · 3h ago

arXiv:2609.38282v1 Announce Type: new Abstract: Vision-language models may rewrite anomalous text in images into linguistically plausible expressions, compromising OCR transcription faithfulness. Sequence-level task rewards and local teacher guidance are complementary, but guidance from the same teacher may not remain equally effective as the student improves. Offline analysis shows that…

Read original article on cs.AI updates on arXiv.org →