Physical Sciences › Computer Science › Artificial Intelligence
Imbalanced Data Classification Techniques
194 artículos indexados
La clasificación de datos desbalanceados examina los métodos para mejorar el rendimiento de los modelos cuando ciertas categorías están subrepresentadas. Estas técnicas abordan desafíos como la calibración de las predicciones, la generación de datos sintéticos o la adaptación de algoritmos para manejar distribuciones desiguales, especialmente en contextos donde los errores tienen consecuencias variables. Los enfoques explorados incluyen variantes de Softmax, métodos ensemble, garantías de seguridad por instancia o estrategias de reequilibrio basadas en la geometría de los errores, aplicadas a tareas que van desde el reconocimiento de imágenes hasta la detección de fraudes.
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos38 % · 48 artículos
- China21 % · 27 artículos
- India9,4 % · 12 artículos
- Francia7,8 % · 10 artículos
- Alemania7 % · 9 artículos
- Canadá5,5 % · 7 artículos
- Irán3,9 % · 5 artículos
- Corea del Sur3,9 % · 5 artículos
Sobre 128 artículos de este tema con al menos un laboratorio localizado. 41 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- Enhancing Foundation Models for Imbalanced SAR Ship Classification via Targeted Oversampling
Ch Muhammad Awais, Marco Reggiannini, Davide Moroni · 29 de septiembre de 2026
Remote-sensing foundation models offer strong representations for SAR imagery, but their behavior under severe long-tail class imbalance is still not well characterized. We benchmark DOFA and SAR-JEPA on the imbalanced OpenSARShip dataset and compare them with ImageNet-pretrained baselines under a f…
- Predicting Symptoms of Amotivation and Anhedonia among University Students with a Novel Oversampling Method
Dang Nguyen, Bao Duong, Arun Kumar, Dat Phan-Trong, Julian Berk, Taylor Braund, Kien Do, Debopriyo Bal, Wu Yi Zheng, Leonard Hoon, Jill Newby, Helen Christensen, Svetha Venkatesh, Alexis Whitton, Sunil Gupta · 25 de septiembre de 2026
University students experience disproportionately high rates of common mental health conditions, such as depression, which can impair learning, social functioning, and overall well-being. Within this context, symptoms of amotivation (i.e. loss of motivational drive) and anhedonia (i.e. diminished in…
- When Identical Rows Disagree: From Benchmark Identifiability to Replication-Robust Anomaly Detection
Jie Deng · 25 de septiembre de 2026
A released table is often treated as an i.i.d. sample, although its repeated rows may encode business frequency, repeated entities, joins, resampling, or extraction errors. We show that this ambiguity creates a hidden measurement layer with three consequences: feature-identical rows impose an attain…
- CRISP: Scalable Importance-Stratified Coresets for Imbalanced Tabular Learning
Hardhik Mohanty, Indrayana Rustandi, Mohamadreza Sheibani · 24 de septiembre de 2026
Large imbalanced tabular datasets make repeated gradient-boosted tree training expensive. Existing coreset methods often lose accuracy when most majority examples are removed. We present CRISP (Coreset Reduction via Importance-Stratified Pruning), a linear-time method that allocates a negative-class…
- Density-Ratio Rescoring for Imbalanced Classification Using Raking Duals and Classifier Scores
Dongha Kim, Seunghwan Park · 23 de septiembre de 2026
Density-Ratio Rescoring (DRR) augments a classifier trained at the original class prior with a survey-raking dual score. Raking reweights the majority sample to match minority feature moments within a tolerance. DRR marginally standardizes the dual and base scores and combines them with a fixed weig…
- Exposing Blind Spots in Deep Imbalanced Regression Evaluation
Noah C. Puetz, Jens U. Brandt, Marc Hilbert, Elena Raponi, Thomas B\"ack, Thomas Bartz-Beielstein · 23 de septiembre de 2026
Deep Imbalanced Regression (DIR) addresses a common failure mode of regression models: target distributions are highly non-uniform, causing models to perform best in densely populated target regions even when reliable performance is required across the full target range. Despite rapid methodological…
- Rethinking Class Imbalance for Single-Cell Foundation Models: A Systematic Benchmark Across Architectures and Long-Tail Loss Functions
Zeyu Dong, Jiahui Zhong · 22 de septiembre de 2026
Single-cell foundation models (scGPT, scBERT, Geneformer) achieve cell-type classification accuracy up to 97.5% in our experiments, yet this aggregate accuracy can mask systematic failure on rare, often disease-relevant cell populations that long-tail loss functions are widely assumed to address. We…
- Improving the Predictive Performance of Bootstrap Aggregating by Dirichlet Resampling
Quoc Viet Le, Joonha Park · 21 de septiembre de 2026
We revisit Breiman's observation that reducing inter-tree correlation without weakening individual trees can improve random forests. Building on this principle, we introduce two variants: Dirichlet-Multinomial Bagging Random Forest (DM) and Dirichlet-Weighted Random Forest (DW). Both modulate sample…
- Auditing a KB Elicitation of Frontier LLM Knowledge: A Multi-dimensional Analysis of GPTKB v1.5
Shrestha Ghosh, Luca Giordano, Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski · 21 de septiembre de 2026
LLMs are remarkable artifacts that have revolutionized a range of knowledge-intensive tasks. A significant contributor is their factual knowledge, which, to date, remains poorly understood, and is usually analyzed from biased samples. In this paper, we provide a framework and the results of a multi-…
- TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection
Huiyuan Liu, Zhiming Ma, Yanxing Liu, Shun Zhang, Qifan Wang, Di Liu, Yifan Wang, Yuyang Deng, Haoyang Meng, Yijin Zhou, Yuxi Zhao, Chengxian Hu, Peidong Wang, Peng Chen · 18 de septiembre de 2026
Telecom fraud scripts evolve rapidly and are often designed to resemble routine service conversations, creating two key requirements for audio-based telecom-fraud evaluation. First, benchmarks must incorporate newly observed scam patterns without overwriting previously established test sets. Second,…
- When majority rules, minority loses: bias amplification of gradient descent
Fran\c{c}ois Bachoc (LPP), J\'er\^ome Bolte (TSE-R), Ryan Boustany (TSE-R), Jean-Michel Loubes (IMT, REGALIA) · 16 de septiembre de 2026
Despite growing empirical evidence of bias amplification in machine learning, its theoretical foundations remain poorly understood. We develop a formal framework for majority-minority learning tasks, showing how standard training can favor majority groups and produce stereotypical predictors that ne…
- Bounded Adjustment with Reliability-Guided Embedding for Imbalanced Learning with Noisy Labels
Mushir Akhtar, Akarsh J., M. Tanveer, Mohd. Arshad · 16 de septiembre de 2026
Class-balanced learning and label noise create a coupled failure mode: frequency correction prevents majority classes from dominating the decision rule, but can amplify incorrectly labeled minority examples. We introduce BARGE (Bounded Adjustment with Reliability-Guided Embeddings), a single-stage o…
- Mini-batch Sampling Strategies for Long-Tailed Image Classification: An Empirical Study on CIFAR-100-LT
Siyu Yuan · 16 de septiembre de 2026
Real-world datasets often exhibit long-tailed class distributions, where a few head classes contain a large number of training samples while a large number of tail classes have only a few. The composition of each mini-batch, determined by the sampling strategy, governs which classes contribute to th…
- PU classification under Non-SCAR: clustering-assisted logistic model with oversampling enhancement
Konrad Furma\'nczyk, Kacper Paczutkowski · 15 de septiembre de 2026
This study addresses the PU classification problem under violations of the SCAR assumption. We investigate logistic regression-based approaches, namely the cluster method and its extensions with strict and non-strict Lasso regularization. The primary contribution of this work is the integration of t…
- Optimizing for the decision not the prediction: an exploration of Smooth Net Benefit as a training objective
Koen M. F. Gorgels, Lasai Barre\~nada, Maarten van Smeden, Ben Van Calster, Ewout W. Steyerberg, Wouter A. C. van Amsterdam · 14 de septiembre de 2026
Objective Prediction models are commonly trained using objectives such as Bernoulli negative log-likelihood (NLL), although downstream clinical decisions may depend on specific risk thresholds. We introduce Smooth Net Benefit ($\sigma$NB), a differentiable approximation of Net Benefit designed to al…
- AUC Maximization from Biased Positive-unlabeled Data with Confidence
Atsutoshi Kumagai, Tomoharu Iwata, Hiroshi Takahashi, Taishi Nishiyama, Kazuki Adachi, Yasuhiro Fujiwara · 11 de septiembre de 2026
Maximizing the area under the receiver operating characteristic curve (AUC) is a standard approach to imbalanced binary classification. Although positive and negative data are required for maximizing the AUC, negative data are often difficult to collect in some real-world applications due to privacy…
- Temporal Heterogeneous Graph Transformer for Credit Card Fraud Detection
Qinwen Yan · 9 de septiembre de 2026
Credit card fraud detection typically relies on tabular features, while repeated attributes can also provide useful relational signals. This paper proposes THGT-FD, a Temporal Heterogeneous Graph Transformer for Fraud Detection. Each transaction is represented using one transaction token and six typ…
- Beyond Homoscedasticity: Decoupled Uncertainty Optimization for Deep Imbalanced Regression
Juncheng Zhou, Jiaxi Lu, Weijing Zeng, Zhong Li, Hao Qi, Jingsong Cui · 7 de septiembre de 2026
Deep Imbalanced Regression (DIR) is pervasive in continuous prediction tasks across diverse modalities, such as age estimation, depth prediction, and protein mutation activity prediction, where label-scarce tail samples often carry higher practical value. However, most existing methods still learn d…
- On the Reliability of Generative Augmentation: A Wasserstein-Based Theoretical and Empirical Study
Chathurika S Abeykoon, Mathias Nthiani Muia, Mallory Goldstein · 2 de septiembre de 2026
Generative data augmentation is widely used to mitigate class imbalance, yet its theoretical effect on downstream generalization remains poorly understood. In this work, we develop a statistical framework for conditional generative augmentation and analyze its impact on classification risk. We forma…
- Local Reference Geometry Residual Augmentation for Imbalanced Time Series Classification
Chuanhang Qiu, Yanran Xu, Yue Wang, Anthony Bagnall · 2 de septiembre de 2026
Imbalanced time series classification is often addressed by changing the training distribution, objective, logits, or final threshold. These interventions address important biases, yet leave a representation-level question unmeasured: after minority support is reduced, does a learned feature space r…
- AIA$^{2}$: Attribute-Agnostic Imbalance Augmentation for Subgroup Robustness
Hanshu Rao, Guangzeng Han, Xiaolei Huang · 1 de septiembre de 2026
Attributes describing data content and context can induce diverse imbalance patterns that go beyond label imbalance alone. However, existing studies primarily address label imbalance while overlooking data attributes, such as topics and demographics, which can induce meaningful subgroup structure wh…
- FraudBench: Protocol-Sensitive Benchmarking of Adversarial Robustness for Financial Risk Assessment
Xitong Zeng, Zhaoge Bi, Yitian Yang, Huaming Chen, Quan Z. Sheng · 26 de agosto de 2026
Machine learning models are widely used in financial fraud and credit-risk detection, yet their adversarial robustness remains difficult to evaluate because financial tabular data involve domain-specific constraints, severe class imbalance, and asymmetric attacker capability. We argue that, in this …
- Temporal Graph Prototype-conditioned Conformal Prediction for Fraud Detection
Xudong Chen, Shengbo Gong, Lu Cheng, Wei Jin · 18 de agosto de 2026
Conformal prediction (CP) provides distribution-free coverage guarantees and has emerged as a principled tool for uncertainty quantification. In edge-level fraud detection on temporal interaction graphs, where false positives and false negatives both carry substantial cost, such coverage guarantees …
- No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage
Shiven Khurdi · 18 de agosto de 2026
We introduce AgentRelBench, an environment-agnostic reliability instrument that computes ground-truth, severity-priced damage from database state diffs across repeated runs, with no LLM in the measurement path, demonstrated on EnterpriseOps-Gym. Across 2,128 evaluation runs spanning nine models in s…
- FinFraudBench: A Heterogeneous Graph Benchmark for Financial Fraud Detection
Yixuan Chen, Hongyu Zhan, Jie Sheng, Weiyu Han, Shuai Chen, Tianyi Zhang, Xiao Tan, Jun Xia · 18 de agosto de 2026
The increasing complexity of digital financial systems has reshaped financial fraud detection from isolated transaction classification into relational risk reasoning over interconnected financial entities. This shift has motivated graph-based fraud detection, where models identify fraudulent nodes b…
Otros asuntos del tema Inteligencia artificial
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Large Language Models7407 artículos / 12 meses+247 %
- Adversarial Robustness in Machine Learning3552 artículos / 12 meses+118 %
- Reinforcement Learning in Robotics2519 artículos / 12 meses+117 %
- Explainable Artificial Intelligence (XAI)2319 artículos / 12 meses+200 %
- Domain Adaptation and Few-Shot Learning2059 artículos / 12 meses+67 %
- Advanced Graph Neural Networks1926 artículos / 12 meses+38 %
