Chapter Four · failure evidence
What K-Means Clustering got wrong, from 54 dissertations
The records document numerous experimental failures, benchmark losses, and methodological rejections of k-means clustering across diverse domains such as single-cell biology, computer vision, and text processing. Researchers frequently found that the algorithm struggles with non-spherical geometries, initialization instability, severe class imbalance, high-dimensional spaces, and non-Euclidean data types. These records come from PhD theses at 28 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.
Sensitivity to initial centroid seeding and convergence instability
Standard k-means clustering is highly vulnerable to seed point selection, which often leads to inconsistent or noisy groupings across successive replicate runs. In several studies, naive iterations failed to improve upon initializations, suffered from degenerate single-cluster collapse, or failed to converge on unbalanced datasets.
Tried and failed
k-means clustering initialized with hierarchical clustering applied to spatial tracking coordinate data. Outcome: worse than baseline. Reason: Lloyd iterations failed to improve upon the initial hierarchical assignments across all k values.
New methods in home-range overlap and clustering · Iowa State
Tried and failed
k-means clustering applied to gene expression time-series data. Outcome: unstable. Reason: Returned inconsistent results over successive runs compared to Gaussian mixture models
Extraction of genetic network from microarray data using Bayesian framework · Cranfield
Considered and rejected
Considered and rejected: Rejected simultaneous clustering and representation learning via naive k-means (DeepCluster style) due to susceptibility to degenerate single-cluster collapse; adopted PU-anchored k-means++ seeding instead
Robust and efficient learning in high dimensions from noisy data · UT Austin
Considered and rejected
Considered and rejected: Rejected K-means relocation clustering method in DAPC in favor of Ward's agglomerative hierarchical clustering method because K-means failed to converge across replicate runs for taxa with small/unbalanced sample sizes.
Considered and rejected
Considered and rejected: Rejected K-Means clustering because of sensitivity to initial seed points, opting for hierarchical clustering.
Glass on the Silk Roads: an SEM-EDS study of Islamic period artifacts from Rayy, Iran: their manufacture and trade connections · University of Nottingham Repository
Considered and rejected
Considered and rejected: Rejected using global clustering (k-means) for pseudo-labeling in hidden unit discovery because of seed selection sensitivity and noisy/arbitrary clustering results.
Self-supervised learning for automatic speech recognition In low-resource environments · University of Nottingham Repository
Considered and rejected
Considered and rejected: Rejected simple K-means clustering (with purely unsupervised centroids) for algospeak topic categorization due to inability to generate clear, stable, and balanced clusters.
Checking, Moderating, Adapting: Collective Engagement for Digital Integrity · Cornell
Tried and failed
K-means clustering without manual centroid presets applied to short text term clustering. Outcome: unstable. Reason: failed to generate clear, stable, and balanced topic clusters without manual centroid presets
Checking, Moderating, Adapting: Collective Engagement for Digital Integrity · Cornell
Considered and rejected
Considered and rejected: Rejected k-means clustering-based evaluation metrics (Normalized Mutual Information and F1 score) in metric learning due to seed sensitivity, failure to separate class cluster quality, and artificial inflation on datasets with few samples per class.
Inability to handle non-spherical geometries and unequal cluster densities
The rigid assumption of convex, isotropic, and equal-sized partitions prevents k-means from accurately resolving complex data distributions. Authors rejected or abandoned the method when dealing with non-spherical spectral variances, varying spatial densities, and non-linearly separable structures.
Considered and rejected
Considered and rejected: Decided against standard K-means and density-based clustering for scRNA-seq because they falsely assume equal cluster sizes and uniform within-population variance.
Considered and rejected
Considered and rejected: Rejected k-means clustering for spatial cell clustering because it requires a predefined cluster count and forces geometrically equal partitions regardless of biological cluster density/noise.
Considered and rejected
Considered and rejected: Avoided standard k-means 1D splits in hierarchical bisection when cluster sizes/densities differ significantly, preferring KDE-based thresholding.
Algorithms for Data Fusion, Representation Learning, and Scalable Clustering based on Constrained Low-Rank Approximation · Georgia Tech
Tried and failed
k-means clustering instead of Gaussian mixture models applied to multispectral image binarization. Outcome: worse than baseline. Reason: Spherical cluster assumption in k-means fails to capture non-spherical spectral distribution variances.
Restoration of multispectral images of ancient documents · DSpace-CRIS at TU Wien
Considered and rejected
Considered and rejected: Rejected K-means and its variants for identifying lower-level subnetwork islands because they struggle with clusters of highly varying sizes, shapes, and densities.
Synthetic models of distribution gas networks in low-carbon energy systems · IRIS - POLITO - prod
Considered and rejected
Considered and rejected: Rejected K-Means and DBSCAN clustering because K-Means assumes convex isotropic clusters and DBSCAN cannot handle large spatial density variations across CFD meshes
A machine learning approach to develop turbulence closures using clustering and neural networks · Imperial
Considered and rejected
Considered and rejected: Rejected K-Means and DBSCAN clustering due to sensitivity to outliers and poor performance on non-linearly separable OCR character shapes.
Automated identification of press variants in old documents · De Montfort Open Research Archive (DORA)
Severe cluster size imbalance and partition collapse
Unconstrained clustering frequently yields degenerate solutions dominated by one or two massive clusters while collapsing minority groups into trivial partitions. This behavior resulted in severe distribution bias, empty clusters, and isolated single-document groups that obscured critical data variability.
Tried and failed
k-means and hierarchical clustering applied to pairwise co-occurrence and embedding distances. Reason: Produced highly unbalanced partitions dominated by one or two massive clusters
Tried and failed
clustering-based multi-instance feature aggregation applied to unimodal histopathology image features. Reason: clustering single-latent-group data concentrates samples into one cluster, causing subject loss from downstream balance requirements
STATISTICAL METHODS FOR VARIABLE SELECTION AND PREDICTION WITH PATHOMIC FEATURES · Penn
Considered and rejected
Considered and rejected: Rejected using k-means/k-modes cluster analysis as primary dimensionality reduction because clusters failed to capture variability in limitations subcodes and minority answer categories.
Tried and failed
k-means clustering for color histogram binning applied to multispectral image calibration. Outcome: unstable. Reason: Fewer clusters caused severe distribution bias, while more clusters created empty partitions.
Considered and rejected
Considered and rejected: Rejected standard unconstrained K-means (SKM) because it creates severely imbalanced, non-contiguous spatial partitions.
Spatial Optimization Techniques for School Redistricting · Virginia Tech
Considered and rejected
Considered and rejected: Rejected k-means clustering with 3 to 5 clusters for semi-supervised labeling because it severely exacerbated class imbalance.
A COMPARISON OF ON-THE-EDGE MACHINE LEARNING CLASSIFICATION METHODS FOR FLIGHT PLANNING · Calhoun
Considered and rejected
Considered and rejected: Rejected unconstrained k-means clustering for speech-to-manifesto authorship attribution, implementing minimum-size constrained k-means clustering instead to prevent the manifesto from forming an isolated single-document cluster.
Speech as data for the politics and partisanship of legislators · Oxford
Failure on categorical, ordinal, and non-Euclidean data
Applying Euclidean distance and arithmetic mean centroids to categorical, ordinal, or textual attributes produces incoherent and uninterpretable partitions. The algorithm cannot naturally capture negatively correlated variables or discrete survey codes, resulting in lower silhouette scores than categorical matching alternatives.
Lost to a baseline
k-means and k-medoids achieved lower silhouette scores (e.g. 0.0887 to 0.1703) on categorical taxonomy benchmarks compared to LSHFk-Centers and k-SCC (up to 0.3536)
Information systems and decision support for sustainable urban transport, energy systems, and emerging methods in research and practice · Leibniz Universität Hannover Repository
Considered and rejected
Considered and rejected: Rejected treating ordinal survey data as interval and applying k-means clustering in favor of LCCA, citing poor performance of naive clustering on categorical data.
Method versatility in analysing human attitudes towards technology · Leibniz Universität Hannover Repository
Considered and rejected
Considered and rejected: Rejected k-means clustering using Euclidean distance across all data, because it treats textual similarities poorly compared to categorical simple matching.
Optimizing Data Compression via Data Reordering Strategies · YorkSpace
Tried and failed
k-modes clustering for categorical dimensionality reduction applied to high-dimensional binary categorical subject tags. Reason: produced incoherent clusters that could not be used as features in downstream regression models
Considered and rejected
Considered and rejected: Rejected K-means clustering because it uses cluster averages instead of actual data points as medoids and cannot naturally cluster negatively correlated variables.
Analysis of High-dimensional Data with Variable Clustering and Selection · Georgia Tech
Tried and failed
k-means and k-modes clustering applied to qualitative categorical survey response codes. Outcome: no signal. Reason: failed to yield distinct representative clusters without obscuring essential data variance
Vulnerability to noise and inability to reject outliers
Because k-means forces every observation into a partition, it cannot discard noise points or isolate localized anomalies. This sensitivity causes outliers and noisy features to distort cluster boundaries and degrade separation quality across spatial and time series data.
Tried and failed
k-means clustering on higher-order principal components applied to functional connectivity feature representations. Outcome: worse than baseline. Reason: higher-order principal components introduced noise that progressively degraded cluster separation quality
Tried and failed
k-means clustering for anomaly detection applied to subsurface drilling sensor time series. Reason: produced overly broad clusters with low silhouette scores rather than isolating localized anomalies
Integrating Machine Learning Techniques with Measurement-While-Drilling Data for Subsurface Characterization in Open-Pit Mines · Virginia Tech
Considered and rejected
Considered and rejected: Rejected k-means clustering for track identification because it requires pre-specifying k for unknown scenes and fails to discard noisy data points.
MARITIME DOMAIN AWARENESS THROUGH THE CHARACTERIZATION OF SHIP BEHAVIOR WITH AIS DATA · Calhoun
Considered and rejected
Considered and rejected: Rejected K-Means clustering and Density-based Clustering for decompositional clustering of seasonal signals due to high sensitivity to outliers (K-Means) and suboptimal cluster bordering/distinction (Density-based).
Characterizing non-E. coli coliforms as indicators of groundwater susceptibility via “big data”, geostatistical analysis, and machine learning · Queens University Institutional Repository
Considered and rejected
Considered and rejected: Rejected K-means and hierarchical clustering for final spatio-temporal hotspot analysis because they cannot remove noise points and require pre-specifying cluster counts.
IDENTIFYING PROBABLE MARITIME PIRACY EVENTS USING MARITIME INCIDENT DATA · Calhoun
Curse of dimensionality and metric degradation in high dimensions
Clustering directly in high-dimensional vector spaces impairs density estimation and leads to the formation of fragmented micro-clusters. Network degree heterogeneity and unreduced sentence embeddings severely degraded cluster quality without prior dimensionality reduction.
Tried and failed
K-means clustering on latent vectors applied to degree-heterogeneous network nodes. Outcome: worse than baseline. Reason: Severe degree heterogeneity in the network biased clustering on weighted latent vectors.
Statistical viewpoints on network model, PDE Identification, low-rank matrix estimation and deep learning · Georgia Tech
Considered and rejected
Considered and rejected: Rejected K-means and Gaussian Mixture Models / kernel-density clustering because cluster counts are unknown and high dimensionality (64D) impairs density estimation.
Quantifying the Effects of Knee Joint Biomechanics on Acoustical Emissions · Georgia Tech
Tried and failed
direct clustering of high-dimensional sentence embeddings applied to text embedding vectors. Reason: Curse of dimensionality prevented effective clustering without intermediate dimensionality reduction.
Innovating the Study of Self-Regulated Learning: An Exploration through NLP, Generative AI, and LLMs · Virginia Tech
Considered and rejected
Considered and rejected: Rejected direct K-means clustering on 768-dimensional BERT embeddings for topic modeling due to the formation of micro-clusters and the curse of dimensionality, adopting TopClus projection instead.
Inability to preserve spatial and temporal continuity
Treating data points as independent observations causes k-means to corrupt continuous sequential structures in image and time-series domains. The algorithm failed to preserve longitudinal spatial continuity and exhibited extreme sensitivity to noise when temporal consistency was ignored.
Tried and failed
K-means clustering on color channels applied to image segmentation. Outcome: no signal. Reason: strong surface reflections and discarding spatial pixel coordinates corrupted cluster separation
Information Extraction from Messy Data, Noisy Spectra, Incomplete Data, and Unlabeled Images · Georgia Tech
Tried and failed
k-means clustering on raw time series features applied to multivariate sensor time series. Outcome: worse than baseline. Reason: sensitivity to noise and lack of temporal consistency in feature space
Fault diagnosis in time series data with application to railway assets · Cranfield
Tried and failed
k-means clustering on 1D spatial projections applied to continuous longitudinal anatomical functional patterns. Reason: Failed to preserve spatial continuity along the longitudinal axis
Spinal Cord fMRI: Functional Connectivity, Network Modeling, and Brain-Spine Interactions · EPFL
Considered and rejected
Considered and rejected: Rejected generic K-means clustering for weekly sub-dataset evaluation because it is not tailored to time-series data.
Clustering and dimensionality reduction for time-series service monitoring data · oURspace
Incompatibility with hierarchical relationships and overlapping memberships
Partitional clustering forces discrete discriminant grouping onto phenomena that are inherently hierarchically nested or continuous. Unsupervised cluster count selection frequently collapsed data into trivial counts, failing to accommodate multidimensional memberships and lineage relationships.
Lost to a baseline
Unsupervised BIC-driven cluster selection frequently selected K=1 or K=2, performing worse than fixing cluster count a priori to K=4.
STATISTICAL METHODS FOR IDENTIFYING AGING-RELATED VULNERABILITY STATES · JScholarship
Considered and rejected
Considered and rejected: Partitional k-means clustering was rejected because biological relatedness is hierarchically nested
The Impact of Medieval and Early Modern Migrations on Dental Nonmetric Variation in Hungary · unevada
Considered and rejected
Considered and rejected: K-means clustering for comment vectors was rejected in favor of agglomerative hierarchical clustering after dendrogram analysis.
Modelling Human Behaviour Based on Similarity Measurements Between Event Sequences · Queens University Institutional Repository
Considered and rejected
Considered and rejected: Rejected hard clustering algorithms (e.g., standard k-means) for textual analysis because they force discrete discriminant grouping rather than accommodating multidimensional cluster memberships.
El comportamiento del donante de sangre en España desde la perspectiva del marketing social · accedaCRIS
Left open by the authors
Problems the authors named and did not get to.
Left open
Compare the satellite clustering performance of k-means against hierarchical clustering algorithms like CHAMELEON or HDBSCAN on the longitudinal tracking dataset. Blocker: None
Left open
Analyze why Ward hierarchical clustering outperforms Neighbor-Joining on specific metric-learned cell embedding spaces. Blocker: None
Reconstructing Cell Lineage Trees from Phenotypic Features with Metric Learning · Penn
Left open
Subcluster regulatory T cells from the post-HSCT scRNA-seq dataset using iterative clustering algorithms like ARBOL to identify rare cell states. Blocker: Requires access to the patient single-cell sequencing dataset generated in the thesis.
Single-cell immune analysis of chronic GVHD after hematopoietic stem cell transplantation · Harvard
Left open
Develop hierarchical metrics to quantify clusterness and trajectoriness on subsets and individual clusters in single-cell datasets. Blocker: None
Developing Graph-based Computational Algorithms for Single-cell Data Science · Georgia Tech
Left open
Perform k-means clustering with k > 14 on mxbai-embed-large-v1 sentence embeddings to separate quantum from generic uncertainty themes in student physics responses. Blocker: Access to the private Cornell student physics text dataset.
Evaluating language models applied to student thinking about experiments · Cornell
Left open
Implement consensus clustering methods, such as voting or mixture models, using LSSM sparse similarity matrices on single-cell sequencing data. Blocker: None
Identifying cell types with single cell sequencing data · HARVEST
Left open
Evaluate deep clustering algorithms on a dedicated, expert-annotated ground truth single-cell dataset with identical marker panels. Blocker: Requires a dedicated, expert-annotated ground-truth single-cell multiplexed image dataset with identical marker panels.
Single-cell Methods and Spatial Analysis for Highly Multiplexed Tissue Images · Harvard
Left open
Apply k-means clustering with PCA or confidence filtering to separate experimental XSW or dF/dz spectra of C60 and H2O@C60. Blocker: Requires experimental XSW or dF/dz spectral measurement data of C60 and H2O@C60 from STM apparatus
Machine learning at the nanoscale · University of Nottingham Repository
Left open
Benchmark alternative clustering algorithms beyond k-means within the iterative filtering and relabeling framework for neuroimaging data. Blocker: Vague task direction with no specific algorithms or validation targets specified
Data driven approaches to address inaccurate nosology in mental health from neuroimaging data · Georgia Tech
Left open
Integrate alternative clustering objectives beyond K-means and hierarchical clustering into the multi-task learning cancer subtyping framework. Blocker: The specific clustering algorithms and their mathematical integration into the multi-task objective are unspecified.
Supervised-Unsupervised Cancer Subtyping Based on Multi-Task Learning · DSpace at SUNY Buffalo
Checking a claim in this area?
We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.