Physical Sciences › Computer Science › Artificial Intelligence
Machine Learning and Data Classification
524 artículos indexados
El estudio de los métodos de Machine Learning y clasificación de datos explora cómo mejorar la precisión y la robustez de los modelos de inteligencia artificial. Las investigaciones abordan técnicas como el semi-supervised learning, donde etiquetas parciales o generadas guían el aprendizaje, o la adaptación de modelos preentrenados a nuevas tareas sin perder rendimiento. Enfoques como la calibración de predicciones, la selección optimizada de prototipos o la gestión de datos ruidosos buscan reforzar la fiabilidad de los sistemas frente a distribuciones variadas o condiciones imperfectas.
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos34 % · 120 artículos
- China32 % · 112 artículos
- Alemania12 % · 42 artículos
- Canadá6,8 % · 24 artículos
- Reino Unido5,6 % · 20 artículos
- Francia5,6 % · 20 artículos
- Australia3,7 % · 13 artículos
- Corea del Sur3,4 % · 12 artículos
Sobre 355 artículos de este tema con al menos un laboratorio localizado. 52 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- Integrating Fairness and Explainability in a Multiple Instance Reinforcement Learning System
Bente Hinkenhuis, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag · 2 de octubre de 2026
Predicting student performance from educational interaction data requires models that are both accurate and sufficiently transparent to support meaningful intervention, while demographic information introduces an additional risk of unfair predictions. This study investigates a multi-objective framew…
- Effective Synthetic Data Curation Requires Group-Level Signals
Cathy Jiao, Chenyan Xiong · 2 de octubre de 2026
Synthetic data now is essential to LLM training, used to strengthen advanced capabilities such as autonomous and long-horizon task execution. Yet recent work shows that training on it at scale can degrade model generation, making it important to decide what synthetic data is worth training on. While…
- Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic)
Zilin Du, Bowen Yang, Boyang Albert Li · 2 de octubre de 2026
Data selection is critical for training large language models on massive and heterogeneous corpora. Meta-learning for Training-data Selection offers a principled alternative to heuristic scoring by learning data weights from a target validation objective, but existing methods face a trade-off betwee…
- Parameter-Efficient Distributionally Robust Adaptation of Tabular Foundation Models under Subpopulation Shift
Seonghwi Kim, Sung Ho Jo, Minwoo Chae · 2 de octubre de 2026
Despite strong mean accuracy, tabular foundation models (TFMs) can perform poorly on underrepresented groups under subpopulation shift, where group proportions change between training and deployment. We propose DR-TFM, a parameter-efficient distributionally robust adaptation framework that requires …
- Optimal Transport Reweighting for Robust Learning under Spurious Correlations and Label Noise
Sung Ho Jo, Seonghwi Kim, Wonsang Yun, Minwoo Chae · 2 de octubre de 2026
Machine learning models often suffer performance degradation under subpopulation shift, particularly when spurious correlations cause models to rely on shortcut features that fail to generalize across subgroups. A recent line of work mitigates this issue by using loss-based signals to identify infor…
- Training-Aware Target Coverage for Synthetic Data Selection
Yang Ba, Michelle V. Mancenido, Rong Pan · 2 de octubre de 2026
Synthetic data are increasingly used to scale LLM training, yet more synthetic data do not necessarily produce better models. Useful synthetic data must add information relevant to the target task without introducing errors that offset their benefit, and the value of an example can change as the tra…
- ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization
Sungho Park, Wonjoong Kim, Jue Zhang, Wook-Shin Han, Pengfei Gao, Chanyoung Park, Yongqiang Yao, Rao Fu, Elsie Nallipogu, Qingwei Lin, Victor R\"uhle · 2 de octubre de 2026
Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scenarios generate the fe…
- A Generalisation Signal Need Not Be a Model-Selection Signal
Aditya Nagarsekar, M P Ashish Bhat, Aadi Nesarkar, Vrishti Godhwani, Rahul Yedida, Aditya Challa, Danda Sravan, Snehanshu Saha · 1 de octubre de 2026
Model selection in computational biology often relies on validation data drawn from the training regime, even when deployment lies outside it. When validation no longer preserves which model is best, a natural alternative is to rank candidates using properties of the trained network itself. We test …
- Signal-Routed Temperature Scaling: Low-Capacity Risk-Conditioned Calibration for Small Validation Budgets
Wenhao Liang, Liangwei Nathan Zheng, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen · 1 de octubre de 2026
When a classifier is recalibrated from only a few thousand held-out examples, the capacity of the calibration map becomes a statistical design choice rather than a purely architectural one: a scalar map can underfit structured residual miscalibration, while a highly adaptive map can be hard to estim…
- LLMs as Adaptive Meta-Solvers: Strategy-Diverse RL for Industrial-Scale Optimization
Shihao Zhang, Weiting Liu, Siyu Shao, Yitian Chen, Jianfeng Feng, Dongdong Ge, Yinyu Ye · 30 de septiembre de 2026
Scaling LLM-based optimization from textbook-scale instances to real-world, industrial tasks remains a critical open challenge. Existing approaches are predominantly evaluated on small, self-contained textual problems and often commit to a solver-integrated paradigm, limiting their ability to handle…
- CPATTA: Conformal Supervision Allocation For Active Test-Time Adaptation
Tingyu Shi, Fan Lyu, Haihua Zhu, Dadi Wang, Shaoliang Peng · 30 de septiembre de 2026
Active Test-Time Adaptation (ATTA) improves model robustness under domain shift by selectively querying human annotations at deployment, but existing methods use heuristic uncertainty measures and suffer from low data selection efficiency, wasting human annotation budget. We propose Conformal Predic…
- From Checkpoint Variation to Selection Gains in Supervised Fine-Tuning
Yupeng Chang, Wenxuan Zhang, Yuan Wu · 30 de septiembre de 2026
Checkpoint selection is a routine decision in supervised fine-tuning (SFT): training produces multiple checkpoints, but only one is retained. Yet fixed-budget comparisons do not by themselves distinguish three empirical claims: whether more validation data improve checkpoint selection, whether a sel…
- When Noise Meets Long-Tail: Feature-Threshold Dual Calibration for Robust Pseudo-Labeling
Ping Guo, Zhiqi Huang, Xinran Li · 29 de septiembre de 2026
Pseudo-labeling has become a cornerstone of learning from unlabeled data in semantic segmentation. Yet its effectiveness drops sharply in real-world scenarios where strong imaging noise and long-tailed class distributions occur together. We trace this failure to a vicious cycle of pseudo-label degra…
- Refreshing Less, Selecting Better: Reusing Stale Gradient Features for Efficient Influence-Based Data Selection
Jianchang Su, Yifan Zhang, Wei Zhang · 29 de septiembre de 2026
Gradient-based data selection methods such as LESS score each candidate by the alignment between its gradient and a target validation gradient, and recomputing per-example gradient features at every new checkpoint dominates their cost. Across three selection seeds, two model families, two candidate …
- A Solvable Theory of Pre-training Data Poisoning: Regime-Dependent Scaling Exponents
Indranil Halder, Rastri Dey, Cengiz Pehlevan · 29 de septiembre de 2026
Pre-training data poisoning of large language models is usually studied using targeted backdoors and their survival through safety post-training, which leaves open a more basic question: how does a model's clean data performance degrade as the poison rate $\varepsilon$ grows? Motivated by our contro…
- GERIS: A Game-Theoretic Framework for Filtering Instance-Dependent Label Noise in License Plate Data Augmentation
Seyedeh Sara Jalili Shani (Department of Computer Science, University of Alberta, Alberta, Canada), Rouhollah Ahmadian (Department of Mathematics and Computer Science, Amirkabir University of Technology, Iran), Amin Rahmani (Department of Mathematics and Computer Science, Amirkabir University of Technology, Iran), Mahdi Bideh (Department of Mathematics and Computer Science, Amirkabir University of Technology, Iran), Mehdi Ghatee (Department of Mathematics and Computer Science, Amirkabir University of Technology, Iran) · 29 de septiembre de 2026
In this paper, we propose GERIS, a game-theoretic framework for instance selection in the data augmentation phase of license plate recognition systems. During augmentation, synthetic license plate images are generated and transformed using stochastic noise to simulate real-world conditions. However,…
- SpecRegMatch: Robust Semi-Supervised Regression for Vehicle Interior Noise Prediction
Sejin Sim, Jinsoo Bae, Seoung Bum Kim · 29 de septiembre de 2026
The rapid advancement of artificial intelligence has observed increased application in predicting vehicle interior noise levels within the automotive industry. However, the collection of labeled data for training models in this context involves significant costs. Previous studies in semi-supervised …
- Efficient Constrained Graph Search for Post-hoc Error Correction in Binary Classifiers
Qinwu Xu · 28 de septiembre de 2026
We introduce a model-agnostic framework for constrained post-hoc error correction in binary classifiers. Given a frozen base classifier, the method searches for an interpretable conjunction of feature--threshold rules that corrects residual false-positive or false-negative errors while explicitly co…
- Teacher-Anchored Selection of Post-Training Quantized Models under Domain Shift
Alejandro Rodriguez Dominguez, Muhammad Shahzad, Xia Hong · 28 de septiembre de 2026
Compressing a trained model yields a family of deployment candidates, and under domain shift the most compressed one need not be the one to deploy. We study selection over such a family, with candidates and teacher fixed and target labels absent or scarce. Two findings organize the label-free case. …
- Let Training Guide Selection: Online Synthetic Data Filtering via Real-Anchored Utility
Yanran Wu, Sana Lakdawala, Renzo Tassara Miller, Chongyang Bai, Sharath Ciddu, Shivendra Pratap Singh, Kungang Li, Sandeep Pandey, Chunwei Liu · 25 de septiembre de 2026
Synthetic data can scale training supervision when real-world data are limited, but noise and distribution mismatch can reduce its value. Existing synthetic data selection methods often emphasize fidelity or diversity rather than the learner's evolving needs. We propose FROST, an online framework th…
- Speculative Evaluation of Stochastic LLMs
Qianli Shen, Xiang Li, Ruomeng Ding, Yanxi Chen, Daoyuan Chen, Yaliang Li · 25 de septiembre de 2026
Evaluating a stochastic large language model is costly: benchmark scores estimate expected performance from randomized rollouts, yet uniform repetition ignores sharp differences in task-level rollout variance. We ask how to minimize the variance of a fixed-benchmark mean under an exact rollout budge…
- Right Choice of Classification Algorithms Based on Reinforcement Learning for Prediction of Non-Alcoholic Fatty Liver
Hasan Samadbin, Arman Daliri · 25 de septiembre de 2026
There are many complex issues in the world of artificial intelligence. Some of these problems are solved using other artificial intelligence methods, which are called artificial intelligence for artificial intelligence. Finding an appropriate classifier algorithm is a time-consuming task. For this r…
- NPBoost: Neural Processes with Gradient-Boosted Fixed Effects
Andrea Nava, Ken R\"olli, Armin Begic, Fabio Sigrist · 24 de septiembre de 2026
Neural Processes (NPs) are model-based meta-learners that implicitly learn a stochastic process and adapt to a new task from a small context set. Most extensions of NPs focus on improving the neural network architecture. We instead develop an extension motivated by the shared hierarchical interpreta…
- LWCal: Loss-Weighted Calibration for Tabular Classifiers with Noisy Calibration Labels
Zeming Liu, Hang Lyu, Jingtao Zhang, Yuan Xie · 24 de septiembre de 2026
Post-hoc probability calibration is usually evaluated under an optimistic assumption: the held-out calibration labels are clean. In many AI deployment settings, however, labels come from weak annotators, historical decisions, heuristics, or distant supervision, so the same label noise that corrupts …
- CORE-STACK+: Meta-Learning for Deep Stacked Generalization
Noor Islam S. Mohammad · 24 de septiembre de 2026
Stacking heterogeneous vision backbones (CNNs, ViTs, and hybrids) is the de facto recipe for accuracy, calibration, and robustness, yet two coupled pathologies limit its returns. Prediction-space multicollinearity ill-conditions the meta-learner's Gram matrix, inflating weight variance and producing…
Otros asuntos del tema Inteligencia artificial
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Large Language Models7407 artículos / 12 meses+247 %
- Adversarial Robustness in Machine Learning3552 artículos / 12 meses+118 %
- Reinforcement Learning in Robotics2519 artículos / 12 meses+117 %
- Explainable Artificial Intelligence (XAI)2319 artículos / 12 meses+200 %
- Domain Adaptation and Few-Shot Learning2059 artículos / 12 meses+67 %
- Advanced Graph Neural Networks1926 artículos / 12 meses+38 %
