OdeSIS: in-depth analysis of outlier detection positioning using median absolute deviation
Abstract
Outliers are easy to overlook, but they quietly skew data analysis and the decisions built on it. This study examined how the median absolute deviation (MAD) algorithm performs when applied before or after K-means clustering. To test this, the study used a synthetic dataset with five deliberately inserted outliers and an augmented version of the Iris setosa dataset containing three anomalies. MAD was applied first, then again after clustering with two and three clusters. The results were clear. When MAD ran first, it detected every outlier with 100% accuracy rate and zero false positives. After clustering, its performance slipped, false positives appeared, and it failed to flag genuinely present outliers. K-means is likely the culprit. Its mean-based centroids get dragged toward anomalous points, so outliers get absorbed into clusters where they no longer look unusual, and MAD stops flagging them. These findings suggest MAD works best as a pre-processing step. Detecting anomalies before clustering improves outlier identification, preserves the underlying data, and prevents stray points from steering cluster formation.
Keywords
Algorithm; Clustering; Data mining; K-means; Outlier
Full Text:
PDFDOI: https://doi.org/10.11591/eei.v15i5.11397
Refbacks
- There are currently no refbacks.

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Bulletin of Electrical Engineering and Informatics (BEEI)
ISSN: 2089-3191
,
e-ISSN: 2302-9285
This journal is published by the
Institute of Advanced Engineering and Science (IAES)
.