TY - GEN
T1 - Bagel
T2 - European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases, ECML PKDD 2024
AU - Rathee, Mandeep
AU - Funke, Thorben
AU - Anand, Avishek
AU - Khosla, Megha
PY - 2026
Y1 - 2026
N2 - Evaluating interpretability approaches for graph neural networks (GNN) specifically is known to be challenging due to the lack of a commonly accepted benchmark. Given a GNN model, several interpretability approaches exist to explain GNN models with diverse (sometimes conflicting) evaluation methodologies. In this paper, we propose a benchmark for evaluating the explainability approaches for GNNs called Bagel. In Bagel, we first propose four diverse GNN explanation evaluation regimes – 1) faithfulness, 2) sparsity, 3) correctness, and 4) plausibility. We reconcile multiple evaluation metrics in the existing literature and cover diverse notions for a holistic evaluation. Our graph datasets range from citation networks and document graphs to graphs from molecules and proteins. We conduct an extensive empirical study on four GNN models and nine post-hoc explanation approaches for node and graph classification tasks. We release both the benchmarks and reference implementations and make them available at https://github.com/Mandeep-Rathee/Bagel-benchmark.
AB - Evaluating interpretability approaches for graph neural networks (GNN) specifically is known to be challenging due to the lack of a commonly accepted benchmark. Given a GNN model, several interpretability approaches exist to explain GNN models with diverse (sometimes conflicting) evaluation methodologies. In this paper, we propose a benchmark for evaluating the explainability approaches for GNNs called Bagel. In Bagel, we first propose four diverse GNN explanation evaluation regimes – 1) faithfulness, 2) sparsity, 3) correctness, and 4) plausibility. We reconcile multiple evaluation metrics in the existing literature and cover diverse notions for a holistic evaluation. Our graph datasets range from citation networks and document graphs to graphs from molecules and proteins. We conduct an extensive empirical study on four GNN models and nine post-hoc explanation approaches for node and graph classification tasks. We release both the benchmarks and reference implementations and make them available at https://github.com/Mandeep-Rathee/Bagel-benchmark.
KW - Explainability
KW - Graph Neural Networks
KW - Interpretability
UR - https://www.scopus.com/pages/publications/105040154058
U2 - 10.1007/978-3-032-25305-7_3
DO - 10.1007/978-3-032-25305-7_3
M3 - Conference contribution
AN - SCOPUS:105040154058
SN - 9783032253040
T3 - Communications in Computer and Information Science
SP - 27
EP - 42
BT - Machine Learning and Principles and Practice of Knowledge Discovery in Databases - International Workshops of ECML PKDD 2024, Revised Selected Papers
A2 - Cerrato, Mattia
A2 - Kalinauskaite, Danguole
A2 - Lukoševicius, Mantas
A2 - Šutiene, Kristina
A2 - Pechenizkiy, Mykola
PB - Springer
CY - Cham
Y2 - 9 September 2024 through 13 September 2024
ER -