Physical Sciences › Computer Science › Artificial Intelligence
Data Stream Mining Techniques
171 artículos indexados
El estudio de los flujos de datos en inteligencia artificial se centra en los métodos que permiten analizar información que llega de manera continua, sin almacenamiento previo. Estas técnicas buscan, en particular, detectar y adaptarse a los cambios en los datos, como el concept drift, donde las relaciones entre las variables evolucionan con el tiempo. También abordan desafíos como la clasificación automática, la detección de anomalías, la preservación de la privacidad o la mejora de modelos en tiempo real, apoyándose en enfoques como los Gaussian Mixture Models, el reinforcement learning o las arquitecturas cognitivas.
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos33 % · 36 artículos
- China19 % · 21 artículos
- Reino Unido7,4 % · 8 artículos
- Francia6,5 % · 7 artículos
- Alemania6,5 % · 7 artículos
- Japón6,5 % · 7 artículos
- Corea del Sur5,6 % · 6 artículos
- Australia5,6 % · 6 artículos
Sobre 108 artículos de este tema con al menos un laboratorio localizado. 37 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- Foundations of Reinforcement Learning and Interactive Decision Making
Dylan J. Foster, Alexander Rakhlin · 28 de septiembre de 2026
Interactive decision making is the problem of learning to act well in an unknown environment, using the data that one's own actions generate to continuously improve, and arises in situations ranging from online platforms and robotics to medical treatments. This monograph gives a statistical perspect…
- Rolling Conformal Prediction in Sequential Model Training
Chen Cheng, Ruiting Liang, Rina Foygel Barber · 24 de septiembre de 2026
We introduce Rolling Conformal Prediction (rolling-CP), a distribution-free predictive inference method for the setting of sequential model training. Specifically, given a data stream $(X_1,Y_1),(X_2,Y_2),\dots$, at each time $n$ the trained model may depend on the observed history $\{(X_i,Y_i)\}_{i…
- HYDRA: Proactive Android Malware Drift Adaptation via Hierarchical Graph Contrastive Learning
Han Chen, Hanchen Wang, Hongmei Chen, Lu Qin, Wenjie Zhang, Ying Zhang · 23 de septiembre de 2026
Concept drift, driven by the rapid evolution of Android malware, severely degrades the performance of machine learning detectors. Current adaptation strategies are often reactive, responding only after performance has dropped and imposing a significant manual annotation burden, or they are proactive…
- Concept Drift from a Causal Perspective
Eduardo V. L. Barboza, Jean Paul Barddal, Robert Sabourin, Rafael M. O. Cruz · 23 de septiembre de 2026
Concept drift is a common phenomenon in real-world data streams, in which changes in the data-generating distribution can degrade predictive model performance. Most existing definitions characterize drift as changes in the joint distribution $P(\mathbf{x}, y)$, without distinguishing which component…
- The Anatomy and Boundary of Adaptation under Temporal Tabular Shift
Tianyu Wang, Xi Vincent Wang, Lihui Wang, Mian Li, Zhihao Liu · 14 de septiembre de 2026
Prequential adaptation of frozen tabular foundation models under temporal drift, with each label revealed only after prediction, helps some deployments and harms others, yet current practice does not predict which. We study the sources and limits of these gains. A diagnostic anatomy attributes gains…
- General Quantification of Covariate and Concept Shifts
Hongbo Chen, Li Charlie Xia · 11 de septiembre de 2026
Generalization under distribution shift remains a core challenge in modern machine learning, yet existing learning bound theory is limited to narrow, idealized settings and is non-estimable from samples. In this paper, we bridge the gap between theory and practical applications. We first show that e…
- Measuring the Value of World-Model Updates: A Counterfactual Utility Protocol for Continual Adaptation
Anqi Peter Li, Kaden Kim · 11 de septiembre de 2026
Continual world models must decide whether new data justify changing the model. Fixed replay schedules and prediction-error triggers specify when to update, but neither reveals the value of an individual update: one deployment run cannot show how the same model would have performed at that moment ha…
- Evaluating Model Retraining under Drift: Paired Comparisons of Cumulative Subgroup Disparity
Aaron Ceross · 10 de septiembre de 2026
Choosing when to retrain a deployed classifier requires assessing subgroup error rates across the sequence of models used, including periods between updates. We compare complete scheduled, loss-triggered, and subgroup-gap-triggered policies with retaining the initial model on the same observations a…
- SCCM : Stream Cruise Control Method for Automated Drift Detection and Adaptation
Mohammad Abu-Shaira, Weishi Shi · 10 de septiembre de 2026
Real-world datasets often exhibit evolving distributions, known as concept drift. Ignoring drift degrades predictive performance, while reliance on fixed hyperparameters further limits model adaptability under changing conditions. Adaptive learning addresses this challenge by continuously updating m…
- When Does Online Adaptation Pay on the Edge? A Leakage-Free Evaluation of Warmup, Learning-Rate Selection, and Resource Trade-offs for Time-Series Forecasting
Takumi Fujimoto, Hiroaki Nishi · 2 de septiembre de 2026
Online adaptation can help edge time-series forecasting under distribution drift, but its measured benefit is sensitive to evaluation choices. We study six public multivariate streams, including building-sensor and smart-meter data, under a leakage-free streaming protocol. We identify two additional…
- Efficient Online Continual Foundation Model Fine-Tuning for Predictive Process Monitoring
Sjoerd van Straten, Marwan Hassani · 31 de agosto de 2026
Predictive Process Monitoring (PPM) models are increasingly deployed in dynamic environments where concept drift causes the underlying process distribution to shift over time. While recent work has moved toward online continual learning, existing methods train compact, task-specific networks entirel…
- RECAST: Recent & Context-Aware Sampling for Test-Time Adaptation in Streaming Biosignals
Yong-Yeon Jo, Junho Song, Joon-myoung Kwon · 31 de agosto de 2026
Streaming biosignals vary across subjects and drift over time, so population-trained models lose accuracy during long-term monitoring. Test-time adaptation (TTA) enables online personalization by updating the model on incoming samples. But in a stream, a basic question is left open: \emph{which samp…
- Provenance Guided Incremental Learning Under Evolving Concept Definitions
Ismail Lamaakal · 26 de agosto de 2026
Learning systems deployed over long periods must adapt not only to statistical changes in incoming data, but also to revisions of the definitions that generate their prediction targets. Conventional concept-drift methods typically infer such changes from observations or prediction errors, even when …
- FreKoo++: Learning Continuous Spectral Dynamics for Temporal Domain Generalization
En Yu, Xiaoyu Yang, Wei Duan, Guangquan Zhang, Jie Lu · 25 de agosto de 2026
Temporal Domain Generalization (TDG) aims to learn from historical domains and generalize to unseen future distributions under concept drift. Nevertheless, prevailing TDG methods struggle with complex real-world streaming scenarios involving both multi-scale drift patterns (e.g., long-term periodici…
- skchange: Fast and Flexible Algorithms for Changepoint Detection
Martin Tveten, Johannes Voll Kolst{\o}, Per August Jarval Moen · 21 de agosto de 2026
Skchange is an open-source Python library for detecting structural changes in time series. It implements modern change detection algorithms within a unified and extensible framework. The algorithms are modular and composable, and they include changepoint search methods based on both cost minimisatio…
- End-to-end Early Classification of Time Series in Non-Stationary Environments
Aur\'elien Renault, Alexis Bondu, Antoine Cornu\'ejols, Vincent Lemaire · 21 de agosto de 2026
Early Classification of Time Series (ECTS) requires making accurate decisions as early as possible in inherently online and evolving environments. Yet, most existing methods assume stationarity and rely on separable designs, where classification and triggering are optimized independently, an assumpt…
- Evidence Before Expansion: Reuse, Spawn, or Defer in Lifelong Expert Pools
Kentaro Oda · 21 de agosto de 2026
Streaming systems that maintain a pool of expert models must repeatedly decide whether to reuse an existing expert for arriving data, spawn a new one, or defer. We present a decision layer that makes all three outcomes statistically meaningful. Reuse and spawn are posed as one-sided sequential hypot…
- Separating Covariate Shift from Mechanism Change with Two Discriminators: CJSD, a Conditional Discrepancy with an Exact Covariate-Concept Decomposition
Kentaro Oda · 21 de agosto de 2026
Streaming systems that maintain a pool of expert models must repeatedly decide whether to reuse an existing expert for arriving data, spawn a new one, or defer. We present a decision layer that makes all three outcomes statistically meaningful. Reuse and spawn are posed as one-sided sequential hypot…
- When to Retrain: An Empirical Study of Retraining Policies for Streaming ML Under Concept Drift, Budget, and Latency Constraints
Sawan Dasari · 21 de agosto de 2026
Production machine learning systems degrade under concept drift, yet practitioners have little principled guidance on when to retrain. Retraining is costly, retraining budgets are finite, and a retrained model does not take effect instantly: training and deployment latency leave a stale model servin…
- When Does Dynamic Ensembling Pay Off? Diagnosing Regionwise Gains in Regression under Distribution Shift
Tianxin Zhou, Ruixi Lin · 20 de agosto de 2026
Whether input-dependent ("dynamic") combination of a regression model pool beats the best static blend depends on the shift and is rarely known before deployment. Can a small labeled target-domain probe tell us when reallocating trust across regions of the input space will pay off? We answer this wi…
- Performance Drift Detection in Machine Learning as a Service (MLaaS) for IoT Environments
Deepak Kanneganti, Sajib Mistry, Sheik Mohammad Mostakim Fattah, Erik Elmroth, Aneesh Krishna, Monowar Bhuyan · 20 de agosto de 2026
Machine Learning as a Service (MLaaS) is a powerful cloud paradigm enabling data-driven intelligent applications in Internet of Things (IoT) environments, widely adopted across healthcare, smart homes, and industry due to its cost-effectiveness. However, the dynamic nature of IoT frequently alters d…
- Hoeffding adaptive splitting trees for data stream classification with concept drift and ensemble learning
Daniel Nowak Assis, Jean Paul Barddal, Fabrício Enembreck · 18 de agosto de 2026
Ensembles of decision trees are well-established methods for data stream classification. In ensemble learning, Hoeffding Trees are widely adopted as base learners, performing periodic split attempts according to the Hoeffding bound. Recent studies, however, indicate that this standard splitting mech…
- Maintaining IoT Device Identification under Concept Drift via Budget-Aware Traffic Labeling
Shayan Azizi, Norihiro Okui, Masataka Nakahara, Ayumu Kubota, Gustavo Batista, Hassan Habibi Gharakaheili · 18 de agosto de 2026
Identification of IoT device types from passive traffic is increasingly used for security management in enterprise and ISP networks. However, the performance of machine learning-based classifiers gradually degrades under concept drift as device behavior evolves. Therefore, maintaining classification…
- Concept Drift Detection and Adaptive Retraining of Malware Classification Models
Christofer Washington Berruz Chungata, Martin Jurecek, Katerina Potika, William B. Andreopoulos, Mark Stamp · 14 de agosto de 2026
Concept drift refers to changes over time in the statistical properties of data, as compared to the data that was used to train a learning model. Machine learning models for malware detection or classification are particularly susceptible to performance degradation caused by concept drift, as attack…
- Privacy-Preserving Data Drift Detection and Recovery for Large-Scale LLM Applications via Proxy Representations
Michael Levit, Josh Ledgard, Haoyu Dong, Vishwas Suryanarayanan, Eyal Kolman, Sharon Tan, Qiang Gan, Vishal Chowdhary · 11 de agosto de 2026
LLM applications deployed at scale face a fundamental challenge: privacy constraints prevent direct inspection of user interactions, making it difficult to obtain any representative evaluation dataset or to track the ongoing evolution of production traffic. We present ProxyDrift, a framework that (i…
Otros asuntos del tema Inteligencia artificial
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Large Language Models7407 artículos / 12 meses+247 %
- Adversarial Robustness in Machine Learning3552 artículos / 12 meses+118 %
- Reinforcement Learning in Robotics2519 artículos / 12 meses+117 %
- Explainable Artificial Intelligence (XAI)2319 artículos / 12 meses+200 %
- Domain Adaptation and Few-Shot Learning2059 artículos / 12 meses+67 %
- Advanced Graph Neural Networks1926 artículos / 12 meses+38 %
