Resumen
The class imbalance problem occurs when one class far outnumbers the other classes, causing most traditional classifiers perform poorly on the minority classes. To tackle this problem, a plethora of techniques have been proposed, especially centered around resampling methods. This paper introduces a two-stage method that combines the DBSCAN clustering algorithm to filter noisy majority class instances with a graph-based procedure to overcome the class imbalance. We then experimentally evaluate the behavior of the proposed method on a collection of two-class imbalanced data sets. The experimental results show an improvement in the classification performance measured by the geometric mean of the accuracy on each class and also a higher reduction in the imbalance ratio when compared to several state-of-the-art under-sampling techniques.