Section 6 of 6
Conclusion and Limitations
Yifei Zhang, James Song, Siyi Gu, Tianxu Jiang, Bo Pan, Guangji Bai, and Liang Zhao · about 2 minutes
We introduce Saliency-Bench, a comprehensive benchmark suite for evaluating visual explanations generated by saliency methods in image classification. Our benchmark includes eight diverse datasets spanning gender classification, environment classification, action classification, object classification, cancer diagnosis, disease estimation, pet type classification, and security check classification, each with ground-truth explanations. We conduct extensive benchmarking experiments using six widely adopted saliency methods, evaluating them with multiple performance metrics, including mIoU, Pointing Game, and iAUC. These methods are tested across different image classifier architectures, including ResNet-18 and VGG-19, providing a comprehensive analysis of their performance. Additionally, we explore ViT-B/16 as a saliency method and perform an inter-method reliability analysis. To facilitate future research, we provide an user-friendly toolkit for dataset loading, saliency map generation, and evaluation, streamlining the benchmarking process. By standardizing the evaluation of visual explanations, Saliency-Bench aims to drive progress in XAI.
Despite its contributions, our benchmark has limitations. Human-annotated explanations in datasets like Gender-XAI and Scene-XAI may still introduce biases despite independent assessments. The pancreatic tumor detection dataset combines samples from different sources, which could cause unintended dataset biases. Additionally, the inclusion of gender classification may raise ethical concerns related to reinforcing gender stereotypes. These challenges underscore the need for continued refinement in dataset construction and evaluation methodologies.