PaperScope
LIVE · 2026-09-29 05:40 UTC

Adaptive Ensemble Selection for Noisy Labels on Tabular Data

Faizaan Ali, Inwon Kang, Oshani Seneviratne

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32976 v1
Category
Submitted
2026-09-26

Abstract

Incorrect or corrupted labels in tabular datasets can significantly degrade supervised learning performance, particularly when mislabeling is subtle and not easily detectable from feature space alone. In the context of automated or AI-augmented data science workflows, robust detection of such label noise is critical for building reliable models. We propose a data-centric reasoning module for AI data science systems that automatically diagnoses dataset quality and selects appropriate cleaning strategies. Given a dataset, a meta-model predicts weights over a diverse set of detectors, including confidence-based, neighborhood-based, and distributional methods. Across benchmark datasets with controlled noise, our approach achieves performance comparable to a Confident Learning baseline on average, with dataset-dependent gains and losses, particularly in heterogeneous regimes. We further show that detector effectiveness is systematically linked to dataset properties. These results demonstrate the value of descriptor-driven, data-centric ensembling as a component of AI-assisted data-science pipelines for robust dataset assessment and model reliability.

arXiv abs page · PDF · same-day batch