Didactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models

cs.AI updates on arXiv.org · 1d ago
Research Papers

arXiv:2609.22161v1 Announce Type: new Abstract: Medical large language models are commonly trained on mixtures of didactic data (e.g., textbooks) and clinical data (e.g., patient records), yet how these data types differentially shape model capabilities remains unclear. We address this issue with token-matched experiments that vary the didactic-to-clinical ratio and analyze how data composition…

Read original article on cs.AI updates on arXiv.org →