INTERPRETABLE MACHINE LEARNING FOR HIGGS BOSON EVENT CLASSIFICATION: A REPRODUCIBLE GRADIENT-BOOSTING BASELINE WITH SHAP ANALYSIS

Authors

  • Suhail A. Chandio
  • Shahnawaz Talpur
  • Sanam Narejo
  • Athar Ali
  • Uzaif Talpur

Keywords:

Higgs boson; collider physics; gradient boosting; Random Forest; SHAP; interpretability; reproducibility; ROC/AUC.

Abstract

Separating Higgs-boson signal events from large Standard-Model backgrounds is a standard supervised-learning task on high-level collider observables, and the classifiers used for it are most useful when their decisions can be inspected and reproduced rather than merely scored. We present a compact, fully reproducible pipeline for Higgs event classification on the ATLAS-format dataset popularised by the Higgs Machine Learning Challenge (HiggsML;  labelled training events and  unlabelled test events). The pipeline applies a transparent preprocessing protocol (sentinel-value imputation by training-set medians, standardisation, and leakage-checked stratified splitting), augments twelve high-level observables with six kinematically motivated composite features, and trains two tree-ensemble learners: a Random Forest (RF) baseline and a gradient-boosted decision-tree model (LightGBM). On a held-out validation fold of  events, the RF baseline attains AUC  and , while the LightGBM model attains AUC  and ; both operate in the  AUC range reported for high-level tabular solutions to this dataset. We interpret the LightGBM model with gain-based feature importance and Shapley Additive Explanations (SHAP). Both attributions place the approximate di-tau mass DER_mass_MMC as the dominant discriminant, followed by the transverse-mass surrogate DER_mass_transverse_met_lep and missing-transverse-energy–related quantities, and the SHAP dependence of DER_mass_MMC is smooth and single-peaked, consistent with resonance-like behaviour. The contribution of this work is not a new state of the art but a readable, auditable, and reproducible baseline: we release the preprocessing, feature-construction, training, and evaluation scripts together with all predictions, metrics, and figures so that every reported number can be regenerated. We also report, without embellishment, that the untuned LightGBM model does not improve on the RF baseline here, and we discuss the calibration, robustness, and systematic-uncertainty studies that a downstream physics analysis would additionally require.

Downloads

Published

2026-03-30

How to Cite

Suhail A. Chandio, Shahnawaz Talpur, Sanam Narejo, Athar Ali, & Uzaif Talpur. (2026). INTERPRETABLE MACHINE LEARNING FOR HIGGS BOSON EVENT CLASSIFICATION: A REPRODUCIBLE GRADIENT-BOOSTING BASELINE WITH SHAP ANALYSIS. Spectrum of Engineering Sciences, 4(3), 4562–4571. Retrieved from https://www.thesesjournal.com/index.php/1/article/view/3642