Test-fairness deep learning with influence score
Source: PLOS (Public Library of Science) · 2026-07-16
Author summary Artificial intelligence is increasingly used to support medical image interpretation, but models that look accurate overall can still perform differently across patient populations or healthcare settings. In practice, medical image datasets often come from specific hospitals or regions, and images may differ in camera type, lighting, workflow, and population characteristics. As a result, a model may learn patterns that are tied to the data source rather than to clinically meaningful signs of disease, which can lead to unequal performance when the model is applied elsewhere. In this study, we present a practical framework to evaluate and reduce such uneven behavior without collecting new data or changing the original datasets. Using skin lesion classification as an example, we identify internal model signals that are strongly linked to the dataset source and reduce the model’s reliance on them. We test the approach on two widely used datasets from different populations (ASAN and ISIC 2019) and assess external generalization on a third dataset (PAD-UFES-20). The results show that the revised model maintains strong diagnostic performance while reducing cross-dataset performance disparity under our test-fairness evaluation setting, supporting more reliable deployment across clinical contexts.
1 Introduction The development of artificial intelligence is becoming more and more mature, and people can solve various problems to a great extent through artificial intelligence, such as image recognition, language translation, and classification p... [60399 chars]