Physical Sciences › Computer Science › Artificial Intelligence
Advanced Clustering Algorithms Research
127 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos33 % · 23 artículos
- China17 % · 12 artículos
- Canadá14 % · 10 artículos
- Reino Unido10 % · 7 artículos
- Italia7,2 % · 5 artículos
- Francia4,3 % · 3 artículos
- RAE de Hong Kong (China)4,3 % · 3 artículos
- España4,3 % · 3 artículos
Sobre 69 artículos de este tema con al menos un laboratorio localizado. 26 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- Scaffold-Constrained Subset Dynamic Programming for Exact SSE Clustering
Yordan P. Raykov, Max A. Little · 28 de septiembre de 2026
Exact Euclidean \(K\)-means partitions \(n\) observations into \(K\) unlabelled clusters, but the unrestricted search is generally exponential. We use data-derived geometric graphs to precondition an exact subset dynamic program: as a result only connected vertex subsets are admitted as clusters, wh…
- Selective Inference for Deep Clustering in Latent Spaces
Eina Mizui, Tomohiro Shiraishi, Shunichi Nishino, Ichiro Takeuchi · 25 de septiembre de 2026
Deep clustering is a powerful approach for discovering meaningful structures in high-dimensional data by learning a low-dimensional latent representation prior to clustering. Despite its empirical success, assessing the statistical reliability of the resulting clusters remains challenging. Testing d…
- Automatic depth-based local center clustering via $\beta$-integrated local depth and adaptive grouping
Siyi Wang, Alexandre Leblanc, Paul D. McNicholas · 23 de septiembre de 2026
Clustering is an unsupervised learning technique that partitions unlabeled data into groups. Most existing methods require user-specified parameters, such as the number of clusters or neighborhood size. Conversely, we propose automatic depth-based local center clustering (A-DLCC), a fully data-drive…
- Multi-View Fair Clustering Guided by Cross-View Sensitive Information Discrepancy
Mudi Jiang, Jiahui Zhou, Xinying Liu, Zengyou He, Zhikui Chen · 23 de septiembre de 2026
Multi-view clustering (MVC) aims to uncover latent cluster structures by exploiting complementary information from multiple views. Despite substantial progress in clustering performance, fairness remains an important concern when MVC is applied to socially sensitive scenarios. Recent fair multi-view…
- Multi-Domain Clustering via Measure Quantization
Rafael Pereira Eufrazio, Eduardo Fernandes Montesuma, Charles Casimiro Cavalcante · 21 de septiembre de 2026
Clustering is a fundamental task in data analysis, typically addressed through centroid-based methods such as K-means. In this work, we present a general framework for multi-domain clustering via measure quantization: given samples from multiple domains, we learn a shared set of cluster prototypes b…
- Randomized SVD Approximations for Spectral Co-Clustering of Word-Document Matrices
Fateme Mazdarani, Carlos Toxtli · 18 de septiembre de 2026
Spectral co-clustering is a useful tool for discovering latent structure in word-document matrices, but its reliance on singular value decomposition (SVD) can make standard formulations expensive on high-dimensional data. This paper presents two randomized approximations for normalized spectral co-c…
- Hierarchical Clustering Can Jointly Satisfy Richness, Consistency, and Scale Invariance
Daichi Kuroda, Maximilien Dreveton, Matthias Grossglauser, Patrick Thiran · 11 de septiembre de 2026
Despite its ubiquity, clustering lacks a universally accepted definition of what is a cluster. Kleinberg's Impossibility Theorem formalizes this difficulty by showing that no flat clustering method can simultaneously satisfy three natural axioms: scale invariance, richness, and consistency. In this …
- An Empirical Study into Clustering of Unseen Datasets with Self-Supervised Encoders
Scott C. Lowe, Joakim Bruslund Haurum, Sageev Oore, Thomas B. Moeslund, Graham W. Taylor · 7 de septiembre de 2026
Can pretrained models generalize to new datasets without any retraining? We deploy pretrained image models on datasets they were not trained for, and investigate whether their embeddings form meaningful clusters. Our suite of benchmarking experiments uses encoders pretrained solely on ImageNet-1k wi…
- Quality-diversity in dissimilarity spaces
Steve Huntsman · 7 de septiembre de 2026
The theory of magnitude provides a mathematical framework for quantifying and maximizing diversity. We apply this framework to formulate quality-diversity algorithms in generic dissimilarity spaces. In particular, we instantiate and demonstrate a very general version of Go-Explore with promising per…
- Anisotropic View Distance Metric for High-Dimensional Data: Theory, Geometry, and Fast Computation
Yiqun Zhang, Hou-biao Li · 4 de septiembre de 2026
K-Means clustering algorithm is one of the most commonly used clustering algorithms because of its simplicity and efficiency. K-Means clustering algorithm based on Euclidean distance only pays attention to the linear distance between Euclidean distance is an efficient and interpretable similarity me…
- DK-GBMKKM: Dynamic Kernel-Space Granular-Ball Multiple Kernel $k$-Means Clustering
Xiaoyu Lian, Yuchao Zhang, Shuyin Xia, Siqi Zhong, Xuzhao Xiang · 2 de septiembre de 2026
Multiple kernel $k$-means integrates complementary nonlinear similarities by learning a combination of base kernels. Its pointwise optimization, however, is sensitive to noisy and boundary samples and repeatedly operates on sample-scale kernel matrices. Granular-ball representations organize local s…
- Stochastic complexity of vectors containing cluster structure
Daniel Nicorici, Olli Yli-Harja, Jaakko Astola · 2 de septiembre de 2026
This paper studies the problem of computing the stochastic probability (shortest code length) of the encoded vectors containing cluster structure using Normalized Maximum Likelihood (NML) model. This is of great theoretical and practical importance in data clustering based on Minimum Description Len…
- Individual Fairness in Hierarchical Clustering
Binita Maity, Shrutimoy Das · 27 de agosto de 2026
Hierarchical clustering produces ultrametric representations that impose strong global geometric constraints and may distort local similarities in ways that disproportionately affect individual data points. We study hierarchical clustering under an individual fairness requirement that bounds relativ…
- Hyperbolic Hierarchical Clustering for Visual Representation Learning
Jianan Wei, Guikun Chen, Zhiyuan Weng, Chunchao Guo, Yujia Wang, Wenguan Wang · 25 de agosto de 2026
We investigate the token mixer in vision backbones by revisiting clustering, one of the most classic approaches in machine learning. An effective token mixer is a fundamental component of modern vision backbones like vision Transformers, facilitating information exchange between image patches. Mains…
- DICS: Data-Informed Centroid Splitting for Decision Tree Classifiers
MD Saifur Rahman Mazumder, Feng Yu · 21 de agosto de 2026
Decision tree-based models are widely used in machine learning due to their interpretability and strong empirical performance. However, training decision trees can be computationally expensive, particularly for large and high-dimensional datasets, largely due to the exhaustive search over candidate …
- Stop Indexing at Full Precision: Revisiting Clustering for Vector Embeddings
Leonardo Kuffo, Peter Boncz · 18 de agosto de 2026
In this study, we revisit three widely used techniques in vector search and utilize them to optimize vector embedding indexing through clustering: dimensionality reduction, quantization, and dimension pruning. We propose an indexing pipeline in which these techniques are applied before clustering, a…
- Connected Subspace Clustering: Hardness, a Scalable Heuristic, and an Application to Sea Level Geodesy
Johanna Hillebrand, Jan H\"ockendorff, J\"urgen Kusche, Kelin Luo, Heiko R\"oglin, Melanie Schmidt, Christian Sohler, Bernd Uebbing · 17 de agosto de 2026
Constrained optimization extends classical optimization by integrating side information, making it widely applicable across scientific and engineering domains. Consider a setting where we measure variables at different physical locations. When grouping these measurements, we often want clusters that…
- Local Cluster Cardinality Estimation for Adaptive Mean Shift
\'Etienne Pepin · 13 de agosto de 2026
This article presents an adaptive mean shift algorithm in which every parameter used at a point is derived from that point's own distance distribution. The distance distribution from a point to all others is used to estimate the cardinality of the local cluster by identifying a local minimum in the …
- Uncertainty-Aware Probabilistic Constrained Clustering from Entangled Pairwise Supervision
Shaojie Zhang, Ke Chen · 13 de agosto de 2026
Pairwise constrained clustering typically relies on hard must-link/cannot-link labels, whereas realistic pairwise supervision may be real-valued and entangle intrinsic ambiguity, expert judgment, and stochastic corruption. Existing deep constrained clustering (DCC) methods mainly target hard, expert…
- Gromov-Wasserstein Quantization and Clustering: Structure, Rates, and Algorithms
Florian Beier, Stephan Eckstein · 12 de agosto de 2026
Clustering is a fundamental class of data analysis techniques with the most important representatives being centroid-based methods like $k$-means. Such methods are strongly connected to quantization problems, which aim to approximate general probability measures with discrete ones. For example, $k$-…
- Multi-kernel spectral clustering: Entrywise eigenvector perturbation bounds and exact recovery
Zeqin Lin, Guangming Pan, Zhixiang Zhang, Yinbing Zhou · 11 de agosto de 2026
Kernel spectral clustering with a single bandwidth can be inadequate for data exhibiting multiple characteristic pairwise-distance scales, a problem particularly prevalent in the high-dimensional regime. We address this issue through a multi-kernel formulation that aggregates kernels with different …
- GRACE: LLM-Grounded Semantic Metric Spaces for Scalable Mixed-Data Clustering
Zihua Yang, Zhencheng Xie, Junyang Chen, Liang Xie, Yiqun Zhang, Mengke Li, Yang Lu · 11 de agosto de 2026
Clustering mixed tabular data requires a unified metric space to bridge the inherent heterogeneity between continuous numerical measurements and discrete categorical symbols. Traditionally, algorithms rely entirely on dataset-internal statistics to estimate categorical relationships, which confines …
- Density-aware Hierarchical Clustering Based on Element-Categorized Connection Subgraphs
Yuning Yu, Jos\'e Rodr\'iguez-Pi\~neiro, Xuefeng Yin, Bin Feng · 10 de agosto de 2026
Clustering is a fundamental data mining technique for pattern recognition through unsupervised learning. Among various clustering methods, hierarchical clustering, density-based clustering, and graph clustering stand out as representative approaches. For hierarchical clustering, it can be categorize…
- Beyond Feature Importance: A Comparative Analysis of Pattern Detection Methods in Cluster Interpretation
Benjamin Connor, Anna Jurek-Loughrey, Lu Bai, Muhammad Fahim · 7 de agosto de 2026
Interpreting clustering outcomes remains a fundamental challenge in data analysis, particularly in domains such as healthcare where meaningful patterns must be extracted from high-dimensional data. While numerous explainability techniques exist, they are primarily designed to assess feature importan…
- Strong bounds for large-scale Minimum Sum-of-Squares Clustering
Anna Livia Croella, Veronica Piccialli, Antonio M. Sudoso · 5 de agosto de 2026
Clustering is a fundamental technique in data analysis and machine learning, used to group similar data points together. Among various clustering methods, the Minimum Sum-of-Squares Clustering (MSSC) is one of the most widely used. MSSC aims to minimize the total squared Euclidean distance between d…
Otros asuntos del tema Inteligencia artificial
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Large Language Models7407 artículos / 12 meses+247 %
- Adversarial Robustness in Machine Learning3552 artículos / 12 meses+118 %
- Reinforcement Learning in Robotics2519 artículos / 12 meses+117 %
- Explainable Artificial Intelligence (XAI)2319 artículos / 12 meses+200 %
- Domain Adaptation and Few-Shot Learning2059 artículos / 12 meses+67 %
- Advanced Graph Neural Networks1926 artículos / 12 meses+38 %
