ReportGem ReportGem

Academic paper

Measuring Explainer Stability via Attribution Separability

Authors: Eddie Conti, \'Alvaro Parafita, Axel BrandoPublished: 2026-08-03Paper ID: 2608.02697Category: cs.LGLicense: CC BY 4.0

Abstract

Attribution methods (AMs) assign an importance score to each feature and are widely adopted to explain black-box models. However, most methods can produce variable attribution scores due to stochastic components in their definition. In this paper, we propose a distribution-based framework to capture the stability of attribution scores. In particular, our approach allows to understand the degree of separability in the ranked attribution vector and obtain the largest index for which a feature ranking remains reliable. We further extend this framework to compare AMs based on the robustness of their rankings across a dataset. Through experiments, we demonstrate how to apply our method to evaluate explainer stability. Overall, our approach provides a complementary criterion for evaluating the stability of AMs.

This public page contains bibliographic metadata and the author abstract. Use the reader for licensed document access.

Open licensed paper reader