Citation
Mehda, Rafid and Oishi, Ramisa Anjum and Alam, Tamzid Tanvi and Morol, Md Kishor and Hui, Liew Tze (2026) Cross-modal bias in medical vision-language models: a pipeline-aware framework for mechanisms, evaluation, and mitigation. Frontiers in Digital Health, 8. ISSN 2673-253X|
Text
fdgth-8-1904053.pdf - Published Version Restricted to Repository staff only Download (827kB) |
Abstract
Medical vision-language models encode images and clinical text in a shared representation. Across radiology and ophthalmology, their diagnostic performance now approaches that of specialist clinicians. The mechanism behind that performance is also the source of a problem that has gone largely unexamined. These models are trained by contrastive alignment, so bias from the image encoder and bias from the text encoder meet at a single point: the alignment interface. There they can interact and compound in ways that single-modality systems never experience. Fairness research has so far studied the two modalities separately, and medical vision-language models have fallen into the gap between those literatures. We organize the evidence into a three-tier taxonomy keyed to the pretraining pipeline. Tier 1 is data-level bias in the pretraining corpus. Tier 2 is alignment bias produced at the contrastive interface. Tier 3 is inference-time bias that surfaces during deployment. Within this structure, we compare the major evaluation benchmarks, show where they disagree, and assign each mitigation strategy to the tier it actually addresses. Three points emerge. First, no published method spans all three tiers; mitigation is fragmented by pipeline stage. Second, fine-tuning does not remove alignment-stage bias, even in parameter-efficient form, which shifts the burden of debiasing onto pretraining rather than adaptation. Third, the inference-time failures are more dangerous than the literature suggests. Medical-specialist models will abandon a correct reading and defer to a confident user on most trials, and specialization appears to make this worse, not better. We close with a research agenda.
| Item Type: | Article |
|---|---|
| Uncontrolled Keywords: | algorithmic bias, contrastive learning, cross-modal alignment, fairness, healthcare disparities, medical vision-language models, sycophancy, trustworthy AI |
| Subjects: | R Medicine > R Medicine (General) > R855-855.5 Medical technology |
| Divisions: | Faculty of Information Science and Technology (FIST) |
| Depositing User: | Ms Suzilawati Abu Samah |
| Date Deposited: | 03 Sep 2026 06:50 |
| Last Modified: | 03 Sep 2026 06:50 |
| URII: | http://shdl.mmu.edu.my/id/eprint/16627 |
Downloads
Downloads per month over past year
Edit (login required) |
