Abstract Medical vision-language models (MVLMs) offer promise in clinical practice but face limitations in generalizability, data quality, and clinically meaningful evaluation. We propose RadiSim-CL, an MVLM trained via curriculum learning by simulating the three-phase pathway of a radiologist: foundational knowledge understanding, anatomical knowledge, and advanced diagnostic reasoning. To support this, we curate RadiSim, a 12-million image-text pair dataset aligned to these phases. We evaluate the model using a five-stage coarse-to-fine validation framework: (1) modality recognition, (2) anatomical recognition, (3) anatomical localization, (4) abnormality and disease diagnosis, and (5) disease differentiation and grading. This framework spans 24 zero-shot subtasks across MR, CT, and DR imaging. RadiSim-CL achieves comparable performance to state-of-the-art baselines in both foundational and anatomical tasks, and demonstrates superior capabilities in complex reasoning (e.g., an AUC of 0.953 for brain tumor diagnosis and an accuracy of 0.764 for meningioma grading). Ablation studies further confirm the curriculum’s effectiveness. RadiSim-CL thus offers a scalable, clinically aligned solution to enhance diagnostic precision. Similar content being viewed by others Acknowledgements This work was supported in part by National Natural Science Foundation of China (grant numbers 82441023, U23A20295, 62131015), National Key Research and Development Program of China (No. 2022YFE0205700), Beijing Natural Science Foundation (IS24053), and HPC Platform of ShanghaiTech University and Shanghai United Imaging Intelligence Co., Ltd. Author information Authors and Affiliations Corresponding authors Ethics declarations Competing interests M.T. is an intern at Shanghai United Imaging Intelligence Co., Ltd. B.Z., G.R., J.N., Z.X., Y.Z., S.Z., X.C., and D.S. are employees of Shanghai United Imaging Intelligence Co., Ltd. The companies have no role in designing and performing the surveillance and analyzing and interpreting the data. Additional information Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Supplementary information Rights and permissions Open Access This article is licensed
a medical vision-language model for radiological <b>image</b> analysis via curriculum learning
Read the original article
nature.com →