Bringing Clinical Reasoning into AI: Concept-Based Explanations for Medical Imaging
Research Visit Summary — René Vogels Foundation Valentina Corbetta, PhD Student, Netherlands Cancer Institute UiT The Arctic University of Norway, August 2025 – March 2026
The challenge: when AI follows the wrong signal
Breast cancer is one of the most common cancers worldwide, and mammography screening plays a crucial role in early detection. AI models are increasingly being developed to support radiologists in interpreting mammograms — for example, by classifying findings according to the BI-RADS scoring system, which radiologists use to assess the likelihood of malignancy.
However, AI models trained on medical images can learn to rely on the wrong features. Instead of learning to recognize clinically meaningful patterns (e.g., the shape of a mass, its margins, its density), a model may latch onto irrelevant visual artifacts introduced during image acquisition, such as grid misplacement or collimator misalignment. These artifacts can vary systematically across hospitals and scanners, creating spurious associations between visual noise and diagnostic labels. When the model encounters a different hospital’s data, or when artifact patterns change, performance can collapse, not because the model never learned the right features, but because it learned both the right and the wrong ones simultaneously, and followed the wrong one when the two conflicted.
This is the problem my research addresses: how do we build AI models that are robust to these spurious correlations and can be trusted to follow clinical reasoning rather than visual shortcuts?
What we did in Tromsø
During my research visit to the Machine Learning Group at UiT, I worked with Kristoffer Wickstrøm and Elisabeth Wetzer on a study investigating this problem specifically in the context of medical vision-language models (VLMs), a new generation of AI models pre-trained on large collections of paired medical images and clinical reports. These models are increasingly used as the backbone for medical image analysis, and their clinical pre-training was expected to confer robustness to spurious correlations. Our study tested whether this expectation holds.
We designed a controlled evaluation framework that injects synthetic acquisition artifacts into training images at varying levels, and evaluates models under two complementary scenarios: one where artifacts are absent at test time, and one where artifact-class associations are deliberately inverted. Applying this framework to mammography diagnosis and diabetic retinopathy grading, we compared five model architectures ranging from standard vision models to fully concept-supervised architectures.
Our results showed that VLMs encode clinical and artifactual signal in parallel. When artifacts are absent at test time, models can still classify based on clinical features. But when artifacts conflict with pathology, models follow the spurious association, and VLM backbones degraded faster than standard visual backbones despite their higher absolute performance. Among the architectures tested, only models that both reshape their internal representations toward clinical concepts and restrict predictions to concept-based information showed meaningful robustness.
As part of this work, we also produced a new set of expert annotations for 400 mammography images, with pixel-wise delineation of findings following the complete BI-RADS lexicon, a resource that will be publicly released to support future research in this area. The manuscript is now under review at the International Conference on Medical Image Computing and Computer Assisted Intervention 2026, the biggest conference in our field.
Figure: Feature representations of diabetic retinopathy images under increasing spurious correlations, illustrating model behaviour observed across both tasks in our study. A standard vision-language model loses clinically organised structure as artifact prevalence increases (top row), while a concept-supervised model maintains it (bottom row). Each panel shows test-set representations at 100%, 50%, and 0% spurious fraction.

Tromsø
My time in Tromsø was one of the most enriching experiences of my PhD. The Machine Learning Group is exceptionally welcoming and collaborative, and being part of it gave me the opportunity to engage with a much broader range of research than my own project, including through the Visual Intelligence Days and the NLDL 2026 conference. I also had the privilege of visiting the Norwegian Cancer Registry, which gave me a new perspective on the challenges of working with real-world cancer data at a national scale.
I loved Tromsø so much that I extended my stay. Living through both the midnight sun and the polar night, learning to kayak among otters and sea eagles, and cross-country skiing across the island are experiences I will not forget. I am deeply grateful to the René Vogels Foundation for making this visit possible, and I look forward to continuing the collaborations that grew out of it.