Physical Sciences › Computer Science › Information Systems
Data Mining Algorithms and Applications
22 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Latest papers
- ReLaG: A Scalable Framework Generalizing Random Splits to Data with Latent Relations
Anthony Lavertu, Jacob Cote, Sophie Gobeil, Jacques Corbeil, Isabeau Premont-Schwarz, Pascal Germain · 1 October 2026
Random splitting can yield non-independent train--test subsets when a dataset contains related samples, as is common in certain applications such as biochemical studies. This leads to overly optimistic generalization estimates. Here, we introduce ReLaG, a modality-agnostic framework that models samp…
- A Moving-Horizon Approximate Branch-and-Reduce Method for Deep Classification Trees
Chenxuanyin Zou, Jiayang Ren, Qiangqiang Mao, Jing Liu, Marcus Lai, Yankai Cao · 1 October 2026
Despite the importance for interpretability, decision trees face severe scalability challenges. Existing global optimal methods are often limited by binary feature selection and shallow tree depths, whereas traditional heuristic approaches frequently sacrifice predictive accuracy. To overcome these …
- A New Gap Sequence for Shellsort: RL-Driven Algorithm Discovery Beyond $N^{4/3}$
Bo Liu · 25 September 2026
Choosing Shellsort gaps is a well-known open problem. For over sixty years, successful sequences have relied on human-designed formulas, numerical searches, or number-theoretic constructions. Although stronger general bounds exist for dense or mainly theoretical families, the worst-case upper bound …
- Falling Trees: A Model Class for Interpretable Risk Prioritization
Varun Babbar, Zachery Boner, Margo Seltzer, Cynthia Rudin · 22 September 2026
Many real-world decisions require prioritizing high-risk cases, such as clinicians prioritizing high-risk patients before lower-risk ones. Falling rule lists (FRLs), which are ordered if--then rules with monotonically decreasing risks, provide an interpretable framework for such tasks; however, thei…
- Support Thresholds, Not Algorithms, Limit Rare-Association Recovery in Co-Purchase Networks
Xiao Han, Zhen Zhang, Xin Zhao, Jiechun Lei, Moxuan Zheng, Youting Wang · 18 September 2026
The support threshold of the Apriori algorithm involves a trade-off in conducting market basket analysis: the associations that occur frequently are noted with high threshold; however, the low ones lead to generating the large amount of rules. The paper compares five methods for co-purchase edge fil…
- Learned Look-Ahead Splitting Rule for CART
Andrew Gao, Tianlin Liu, Ruichen Han, Lu Tian · 16 September 2026
Classification and regression trees are typically constructed using a greedy splitting rule that maximizes the immediate reduction in prediction error at each node. Although this strategy is computationally efficient, it can miss splits that yield small short-term gains but create substantial downst…
- Literati: Towards Anytime Optimal Shape Generalized Trees via AO*
Nakul Upadhya, Eldan Cohen · 10 September 2026
Decision trees are prized for their interpretability and strong performance on tabular data, but popular greedy top-down induction algorithms can yield suboptimal and unnecessarily complex structures. Optimal decision tree methods address this through global optimization, yet remain restricted to ax…
- Hierarchical and Permutation-Invariant Feature Transformation Learning via Policy-Guided Embedding Search
Rui Liu, Tao Zhe, Yanyong Huang, Sankha Narayan Guria, Xiao Luo, Wei Fan, Yanjie Fu, Dongjie Wang · 10 September 2026
Feature transformation improves predictive performance on tabular data by constructing informative abstractions from raw features. Recent generative approaches encode transformation knowledge into continuous embedding spaces for efficient exploration of candidate strategies, but face three key limit…
- Learning Sparse Decision Trees via Transformer Variational Auto-Encoders
Giacomo Fidone, Alessio Cascione, Riccardo Guidotti · 2 September 2026
Decision trees are among the most widely used models in machine learning, largely due to their transparent decision logic, making them well-suited for high-stakes decision-making contexts. However, most existing learning algorithms focus on predictive performance, overlooking the joint optimization …
- Harnessing the Synergy between LLM Agents and Knowledge Graphs for Urban Socioeconomic Prediction
Zhilun Zhou, Jingyang Fan, Yu Liu, Fengli Xu, Depeng Jin, Yong Li · 10 August 2026
Socioeconomic prediction aims to leverage various urban data to predict the socioeconomic indicators of regions such as population and commercial activity level, which plays an important role in understanding urban regions and supporting decision-making. Existing studies leverage knowledge graphs (K…
- Optimal or Greedy Decision Trees? Revisiting their Objectives, Tuning, and Performance
Jacobus G. M. van der Linden, Dani\"el Vos, Mathijs M. de Weerdt, Sicco Verwer, Emir Demirovi\'c · 7 August 2026
Recently there has been a surge of interest in optimal decision tree (ODT) methods that globally optimize accuracy directly, in contrast to traditional approaches that locally optimize an impurity or information metric. However, the literature shows conflicting evidence on the value of ODTs, with so…
- Fast Discovery of Inclusion Dependencies with Desbordante
Alexander Smirnov, Anton Chizhov, Ilya Shchuckin, Nikita Bobrov, George Chernishev · 4 August 2026
Inclusion dependency is a relation between attributes of tables that indicates possible Primary Key-Foreign Key references. Automatic discovery of inclusion dependencies is a relevant problem for both academic and industrial communities. The core concern for this problem is the efficiency of discove…
- Understanding Context Sampling in TabPFN on Small Tabular Datasets
Mohammed Abdullah · 30 July 2026
TabPFN performs classification through in-context learning: it conditions on a set of labeled training rows (the context, or prototypes) and predicts test labels without gradient updates. On small tabular datasets, practitioners must still choose the context size and which rows constitute the contex…
- Discovering Frequent Closed Embedded Sub-DAGs in Spatio-Temporal Event Data
Piotr S. Maci\k{a}g · 8 July 2026
We propose a novel approach to mine patterns in spatio-temporal event data based on discovering frequent closed embedded sub-Directed Acyclic Graphs (DAGs). In our method, event instances are represented as nodes labelled by event types, while edges capture spatio-temporal following relationships. W…
- Scalable Maximal Frequent Episode Mining with Desbordante
Maxim Ivanov, Matvei Smirnov, Alisa Strazdina, George Chernishev · 7 July 2026
Episode mining aims to extract subsequences of events that possess certain distinctive properties and constitute facts valuable to the user. Maximal frequent episode mining concentrates on discovery of frequently-appearing subsequences, which are not included into any other larger frequent subsequen…
- Clustering with Non-adaptive Subset Queries
Hadley Black, Euiwoong Lee, Arya Mazumdar, Barna Saha · 30 June 2026
Recovering the underlying $k$-clustering of a set $U$ of $n$ points by asking pair-wise same-cluster queries has garnered significant interest in the past few years. Given a query $S \subset U$, $|S|=2$, the oracle returns "yes" if the points are in the same cluster and "no" otherwise. For adaptive …
- AlphaMemo: Structured Search-Process Memory for Self-Evolving Alpha Mining Agents
Hang Yu, Zifan Zheng, Jeff Z. Pan, Tongliang Liu, Zhiyong Wang, Fengxiang He · 23 June 2026
LLM agents are promising for alpha mining via combining financial priors, symbolic reasoning, executable factor generation, and feedback-driven refinement. Yet, they face a combinatorial search space, noisy non-stationary feedback, redundant discoveries, and overfitting risks from naively reusing pa…
- Few-Shot Resampling for Scalable Statistically-Sound Data Mining
Leonardo Pellegrina, Fabio Vandin · 11 June 2026
A key step in knowledge discovery is the evaluation of data mining results. In several applications, including pattern mining, graph analysis, and others, this step includes the evaluation of the statistical significance of the results, to avoid spurious discoveries due only to noise or random fluct…
- Discovering Data Structures: Nearest Neighbor Search and Beyond
Omar Salemohamed, Laurent Charlin, Shivam Garg, Vatsal Sharan, Gregory Valiant · 9 June 2026
We propose a general framework for end-to-end learning of data structures. Our framework adapts to the underlying data distribution and provides fine-grained control over query and space complexity. Crucially, the data structure is learned from scratch, and does not require careful initialization or…
- Frequency-based Constrained Sampling for Interval Patterns
Djawad Bekkoucha, Abdelkader Ouali, Bruno Cr\'emilleux · 9 June 2026
Output space pattern sampling is a powerful alternative to exhaustive pattern mining for exploring large pattern spaces, as it enables users to focus on representative patterns drawn according to a chosen interestingness measure. In this paper, we address the problem of sampling interval patterns un…
- Small Language Model Agents Enable Efficient and High-Quality Knowledge Mining
Sipeng Zhang, Shuhuai Lin, Xinpeng Wei, Yihang Chen, Pin Qian, Su Wang, Huan Xu · 8 June 2026
At the core of Deep Research is knowledge mining, the task of extracting structured information from massive unstructured text in response to user instructions. Large language models (LLMs) excel at interpreting such instructions but are prohibitively expensive to deploy at scale, while traditional …
- Optimal Pattern Detection Tree for Symbolic Rule-Based Classification
Young-Chae Hong, Yangho Chen · 15 May 2026
Pattern discovery in data plays a crucial role across diverse domains, including healthcare, risk assessment, and machinery maintenance. In contrast to black-box deep learning models, symbolic rule discovery emerges as a key data mining task, generating human-interpretable rules that offer both tran…
- Decision Tree Learning on Product Spaces
Arshia Soltani Moakahr, Faraz Ghahremani, Kiarash Banihashem, MohammadTaghi Hajiaghayi · 14 May 2026
Decision tree learning has long been a central topic in theoretical computer science, driven by its practical importance. A fundamental and widely used method for decision tree construction is the top-down greedy heuristic, which recursively splits on the most influential variable. Despite its empir…
- Out of Spuriousity: Improving Robustness to Spurious Correlations without Group Annotations
Phuong Quynh Le, J\"org Schl\"otterer, Christin Seifert · 28 April 2026
Machine learning models are known to learn spurious correlations, i.e., features having strong relations with class labels but no causal relation. Relying on those correlations leads to poor performance in the data groups without these correlations and poor generalization ability. To improve the rob…
- Variable Selection Using Relative Importance Rankings
Tien-En Chang, Argon Chen · 14 April 2026
Although conceptually related, variable selection and relative importance (RI) analysis have been treated quite differently in the literature. While RI is typically used for post-hoc model explanation, this paper explores its potential for variable or feature ranking and filter-based selection befor…
Other topics in Information systems
The topics the OpenAlex classification attaches to the same theme, most active first.
- Software Engineering Research853 papers / 12 months+218%
- Information Retrieval and Search Behavior727 papers / 12 months+506%
- Recommender Systems and Techniques600 papers / 12 months+154%
- Expert finding and Q&A systems205 papers / 12 months+1175%
- Information and Cyber Security169 papers / 12 months+1500%
- Blockchain Technology Applications and Security146 papers / 12 months+650%
