Notableevaluation benchmarks

Do Center Biases Propagate? Robustness of Pathology Foundation Models in Whole-Slide Image Classification

Ilán Carretero, Pablo Meseguer, Rocío del Amor, Valery Naranjo

Published
Sep 23, 2026 14:55 UTC

Problem

This work addresses the gap in understanding center-associated robustness in whole-slide image classification, particularly how biases related to class centers affect model predictions. The study is motivated by the need to evaluate the robustness of pathology foundation models (PFMs) in the context of increasing class-center correlations, quantified by Cramér's V. The paper is a preprint and has not undergone peer review.

Method

The authors conducted experiments in a controlled training setting to evaluate center-associated robustness. They benchmarked six pathology foundation models across four datasets, employing two multiple instance learning (MIL) aggregators. The robustness of the models was assessed using a robustification strategy called ComBat. The performance metric introduced is the Area Under the Cramér's V Curve (AUCC), which captures both classification performance and degradation due to center biases. The study systematically varied class-center correlations to analyze their impact on model predictions.

Results

The findings indicate that PFMs propagate center-related information to whole-slide image (WSI) level predictions, suggesting that the models are sensitive to class-center biases. The robustness of the models was found to depend on both the representation of the PFM and the chosen MIL aggregation strategy. However, the application of ComBat as a robustification strategy did not yield consistent robustness gains across the different datasets evaluated, indicating variability in effectiveness depending on the context.

Limitations

The authors did not report any limitations in their study. However, the lack of reported limitations may suggest a need for further exploration of the robustness across more diverse datasets or additional robustness strategies.

Why it matters

This research has significant implications for the deployment of pathology foundation models in clinical settings, as it highlights the importance of understanding and mitigating center biases in model predictions. The insights gained from this study can inform future work on improving the robustness of PFMs, potentially leading to more reliable diagnostic tools in pathology.

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: arXiv cs.AI