London International Conference on Artificial Intelligence and Data Science Volume 1 Issue 1 (2026) · pp. 3-10

Robust Feature Selection for Imbalanced Tabular Datasets Using Ensemble Gradient Boosting

Whitfield, Daniel, Tan, Mei Ling

DOI: 10.5281/zenodo.22544795 Read at source PDF

Abstract

Imbalanced tabular datasets remain a persistent obstacle for supervised learning pipelines deployed in domains such as fraud detection, industrial defect screening, and rare disease diagnosis, where minority class instances carry disproportionate operational significance. This study proposes a feature selection framework that couples recursive elimination with an ensemble of gradient boosting learners trained under class-aware resampling, aiming to identify predictive variables that remain stable across resampled folds rather than variables that merely correlate with majority class noise. We evaluate the approach across several benchmark tabular collections spanning finance, manufacturing, and healthcare, comparing selected feature subsets against filter-based and wrapper-based baselines using precision-recall area, minority class recall, and selection stability indices. Results indicate that ensemble-guided selection consistently yields smaller, more interpretable feature subsets while preserving minority class detection performance relative to using the full feature space. We further examine how resampling ratio and boosting depth interact with selection stability, offering practical guidance for practitioners tuning these hyperparameters jointly rather than independently. The findings suggest that treating feature selection and imbalance correction as a joint optimization problem, rather than sequential preprocessing steps, produces more robust and auditable machine learning pipelines for high-stakes tabular prediction tasks.

feature selectionclass imbalancegradient boostingensemble learningtabular data

Metadata source: the journal's OAI-PMH archive · Sindex does not store the full text; it links to the source.

Cite

APA 7
Whitfield, Daniel & Tan, Mei Ling (2026). Robust Feature Selection for Imbalanced Tabular Datasets Using Ensemble Gradient Boosting. London International Conference on Artificial Intelligence and Data Science, 1(1), 3-10. https://doi.org/10.5281/zenodo.22544795
GOST R 7.0.5
Whitfield, Daniel, Tan, Mei Ling Robust Feature Selection for Imbalanced Tabular Datasets Using Ensemble Gradient Boosting // London International Conference on Artificial Intelligence and Data Science. 2026. Т. 1. № 1. С. 3-10. URL: https://doi.org/10.5281/zenodo.22544795
BibTeX
@article{daniel2026,
  author  = {Whitfield, Daniel and Tan, Mei Ling},
  title   = {Robust Feature Selection for Imbalanced Tabular Datasets Using Ensemble Gradient Boosting},
  journal = {London International Conference on Artificial Intelligence and Data Science},
  year    = {2026},
  volume  = {1},
  number  = {1},
  pages   = {3-10},
  doi     = {10.5281/zenodo.22544795}
}
RIS
TY  - JOUR
AU  - Whitfield, Daniel
AU  - Tan, Mei Ling
TI  - Robust Feature Selection for Imbalanced Tabular Datasets Using Ensemble Gradient Boosting
JO  - London International Conference on Artificial Intelligence and Data Science
PY  - 2026
VL  - 1
IS  - 1
SP  - 3
EP  - 10
DO  - 10.5281/zenodo.22544795
ER  -