A detailed study of interpretability of deep neural network based top taggers

Abstract
Recent developments in the methods of explainable AI (XAI) allow researchers to explore the inner workings of deep neural networks (DNNs), revealing crucial information about input-output relationships and realizing how data connects with machine learning models. In this paper we explore interpretability of DNN models designed to identify jets coming from top quark decay in high energy proton-proton collisions at the Large Hadron Collider (LHC). We review a subset of existing top tagger models and explore different quantitative methods to identify which features play the most important roles in identifying the top jets. We also investigate how and why feature importance varies across different XAI metrics, how correlations among features impact their explainability, and how latent space representations encode information as well as correlate with physically meaningful quantities. Our studies uncover some major pitfalls of existing XAI methods and illustrate how they can be overcome to obtain consistent and meaningful interpretation of these models. We additionally illustrate the activity of hidden layers as Neural Activation Pattern (NAP) diagrams and demonstrate how they can be used to understand how DNNs relay information across the layers and how this understanding can help to make such models significantly simpler by allowing effective model reoptimization and hyperparameter tuning. These studies not only facilitate a methodological approach to interpreting models but also unveil new insights about what these models learn. Incorporating these observations into augmented model design, we propose the Particle Flow Interaction Network (PFIN) model and demonstrate how interpretability-inspired model augmentation can improve top tagging performance.
Type
Publication
Mach. Learn. Sci. Tech.
This open-access paper presents a systematic study of the interpretability of deep neural network (DNN) models used for top-quark jet tagging in high-energy proton–proton collisions at the LHC.
Key points
- The authors review a subset of existing top-tagger models (primarily MLP-based) and apply multiple quantitative explainable AI (XAI) methods to identify which input features contribute most to classifying jets originating from top-quark decays versus QCD background.
- They examine how feature importance varies across different XAI metrics, the impact of feature correlations on explainability, and how information is encoded in the models’ latent-space representations (including correlations with physically meaningful quantities).
- Major pitfalls of existing XAI techniques are identified, along with ways to obtain consistent and meaningful interpretations.
- Neural activation pattern diagrams are used to visualize information flow through hidden layers, enabling model simplification, re-optimization, and hyperparameter tuning.
- Building on these interpretability insights, the authors propose the Particle Flow Interaction Network (PFIN) model and demonstrate that interpretability-inspired architectural improvements can enhance top-tagging performance.
Published 11 July 2023 in Machine Learning: Science and Technology. This work laid the groundwork for the later evidential deep learning studies by the same group.

Authors
University of Illinois at Urbana-Champaign
I am a professor at the University of Illinois. My research is highly interdisciplinary at the intersection of particle physics, AI/ML, and quantum, aiming to understand the universe at its fundamental level and to accelerate scientific discovery through innovation.
Authors
Authors