Abstract
Alzheimer's disease poses a critical challenge in medical imaging, where accurate diagnosis requires integrating structural brain MRI with clinical descriptions. Visionlanguage models like CLIP have revolutionized natural image understanding, but their direct application to neuroimaging remains limited due to substantial domain shift and data scarcity. To address this gap, we present NeuroVLM, the first vision-language model for image-text retrieval in Alzheimer's disease neuroimaging. By adapting CLIP on thousands of T1-weighted MRI scans from the ADNI dataset paired with structured clinical captions, NeuroVLM establishes a novel multimodal retrieval benchmark for Alzheimer's neuroimaging. Our model substantially outperforms general-purpose vision-language models and biomedical adaptations, achieving more than double their retrieval accuracy through domain-specific fine-tuning combined with universal retrieval enhancements including soft-LME pooling, CSLS de-hubbing, and reciprocal rank fusion. With median rank of one and comprehensive evaluation across multiple protocols, this work provides the specialized vision-language framework for Alzheimer's disease with rigorous benchmarking, enabling natural language queries over brain MRI repositories for clinical decision support.