Physical Sciences › Computer Science › Artificial Intelligence
Text and Document Classification Technologies
121 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- China30% · 19 papers
- United States29% · 18 papers
- India7.9% · 5 papers
- Germany7.9% · 5 papers
- United Kingdom7.9% · 5 papers
- Canada6.3% · 4 papers
- France6.3% · 4 papers
- Poland4.8% · 3 papers
Across 63 papers on this subject with at least one lab located. 28 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- GCUL: Ambiguity Identification in Text Emotion Classification via Cluster-Guided Learning
Zhongqi Fan, Tianyou Zhang, Fei Chen · 25 September 2026
Selective classification enables a model to abstain from predictions on uncertain instances, but existing approaches typically reject them through confidence scores, predefined coverage constraints or instance-level distance measures. These approaches may overlook the collective geometric structure …
- MORE-PLR: multi-output regression employed for partial label ranking
Santo M. A. R. Thies, Juan C. Alfaro, Viktor Bengs · 25 September 2026
The partial label ranking problem is a supervised learning scenario that aims to fit a preference model that predicts a bucket order defined over a set of labels for a given input instance. This problem generalizes the well-known label ranking problem, which, in practice, is limited to outputting to…
- ReCalMatch:Reliability-Calibrated Semantic Guidance for Semi-Supervised Fine-Grained Recognition
Yundi Hong, Hongyang He, Zheng Fang, Xuanyu Liu, Victor Sanchez · 25 September 2026
Semi-supervised fine-grained visual recognition is highly vulnerable to overconfident pseudo-label errors: visually similar categories frequently produce high-confidence yet incorrect predictions, and consistency regularization then reinforces these errors throughout training. Existing semi-supervis…
- A Manifold-Aware Topic Modeling Approach via Rank-Based Prototypes
Thiago C\'esar Castilho Almeida, Daniel Carlos Guimar\~aes Pedronette · 25 September 2026
Recent topic models leverage pretrained embeddings, but neural architectures produce latent representations without grounding in specific texts, and clustering-based pipelines assign representative documents only post hoc, relying on absolute distances distorted by hubness and anisotropy in high-dim…
- How to Estimate Whether You Have Found Several Needles in a Haystack: Measuring Calibration in Multi-Label Text Classification
Sophie Henning, Georg Hofmann, Alexander Schulte, Alexander Fraser, Annemarie Friedrich · 23 September 2026
A key factor in deciding whether to trust an automatic prediction is its confidence score, which should be calibrated to match the actual probability of the prediction being correct. Most confidence calibration metrics target binary or multi-class tasks, while multi-label calibration remains largely…
- Follow the Geometry, Not the Model: Cold Start Semi-Supervised Learning
Itai David, Daphna Weinshall · 15 September 2026
Modern semi-supervised learning (SSL) couples pseudo-label generation and classifier training, using the classifier's own confidence to select the pseudo-labels that are then used to update the model. In the cold-start regime, where at most a few labels per class are available, this coupling is ill-…
- LLM-Enhanced Dual-Branch Learning for Large-Scale Multi-Label Text Classification
Hui Ye, Jing Zhang, Xiulong Yang, Rajshekhar Sunderraman · 14 September 2026
Large-scale multi-label text classification assigns a small subset of relevant labels to each document from a vocabulary containing thousands or tens of thousands of candidate labels. Although pretrained language models have improved semantic text representations, most representation-based approache…
- Importance Weighting for Unlabeled-unlabeled Learning under Distribution Shift
Atsutoshi Kumagai, Tomoharu Iwata, Hiroshi Takahashi, Taishi Nishiyama, Kazuki Adachi, Yasuhiro Fujiwara · 11 September 2026
Unlabeled-unlabeled (UU) learning allows us to learn a binary classifier from two sets of unlabeled data with different class-priors. It is a general framework because it includes a wide variety of supervised learning such as positive-unlabeled (PU) learning, noisy label learning, and similarity-bas…
- A statistical approach to bias in zero-shot learning: the lens of handwriting recognition
Clarence Chew, Gim Siang Chia, Sukalpa Chanda, Subhroshekhar Ghosh, Soumendu Sundar Mukherjee · 10 September 2026
Generalized zero-shot learning (GZSL) has emerged as an important paradigm for visual recognition systems that must generalize to classes that were not observed during training. Traditional GZSL techniques are limited by their applicability to a relatively small number of such unseen classes, scalab…
- The Accuracy Paradox: Empirical Diagnostic of Default Decision Thresholds in Multi-Label Enzyme Commission Prediction [With Code]
Bilal Ahmad, Rajed Mehmood · 9 September 2026
Automated prediction of Enzyme Commission (EC) numbers plays a central role in functional annotation and computational drug discovery. However, standard multi-label machine learning pipelines frequently rely on default decision thresholds (t=0.50), assuming balanced prior distributions across target…
- OSR: Output Space Redistribution for Adaptive Label Removal in Classification Models
Minyi Peng, Darian Gunamardi, Ivan Tjuawinata, Yongsen Zheng, Kwok-Yan Lam · 4 September 2026
Label removal occurs frequently in classification systems with evolving taxonomies, where categories must be dynamically updated or eliminated. To accommodate such changes, classification models must adapt accordingly. Existing solutions, broadly categorized as retraining-based and feature-space-adj…
- Position: Unlabeled IS NOT Equal to No Human Supervision in Visual Learning
Dong Lao · 4 September 2026
This position paper argues that the absence of labels does not imply the absence of human supervision in visual learning, and urges the research community to identify sources of supervision more explicitly. Many recent methods in computer vision build upon representations learned from large-scale un…
- No Data Wasted: A Semi-supervised Generative Model for Incomplete Multi-view Data Integration with Missing Labels
Yiyang Shen, Weiran Wang · 3 September 2026
Multi-view learning is widely applied to real-life datasets, but it often suffers from both missing views and missing labels. Prior probabilistic approaches addressed the missing view problem by using a product-of-experts scheme to aggregate representations from present views and achieved superior p…
- From Confusion to Clarity: Confusion-Aware Retrieval and Knowledge Injection for Text Classification
Manish Gupta, Chaitanya Giri, Jayasimha Talur · 2 September 2026
Large language models (LLMs) struggle to classify text into taxonomies with many semantically similar labels, as the distinctions are domain-specific and not captured by pre-training. To handle large label spaces, a common approach retrieves top-$K$ candidate labels by embedding similarity and promp…
- Type-Balanced Contextual Learning for Incremental Named Entity Recognition
Duzhen Zhang, Yahan Yu, Xiuyi Chen, Chenxing Li, Dong Yu · 1 September 2026
Incremental Named Entity Recognition (INER) stands as a pivotal task in information extraction, emphasizing the successive identification of new entity types within unstructured text. Faced with the continuous influx of entity types, INER grapples with two significant challenges: the widespread issu…
- Label Semantic Expansion via Label Guided Neural Topic Modeling
Haojia Zheng, Yuyin Lu, Juntian Huang, Fan Ou, Yanghui Rao, Haoran Xie, Fu Lee Wang · 1 September 2026
Topic models are widely used for content analysis, where users often analyze corpora around predefined labels rather than unordered latent topics. Existing label-aware topic models mainly follow a labels-for-topics perspective, using labels to guide topic learning, while the learned topics are not d…
- Comprehensive Evaluation and Analysis for NSFW Concept Erasure in Text-to-Image Diffusion Models
Die Chen, Zhiwen Li, Cen Chen, Yuexiang Xie, Xiaodan Li, Jinyan Ye, Yingda Chen, Yaliang Li · 31 August 2026
Text-to-image diffusion models have gained widespread application across various domains, demonstrating remarkable creative potential. However, the strong generalization capabilities of diffusion models can inadvertently lead to the generation of not-safe-for-work (NSFW) content, posing significant …
- Self-Explainable Multi-Label Graph Neural Network for Correlated Evidence Attribution
Yingqi Feng, Yufei Tang, Min Shi, Xingquan Zhu · 31 August 2026
Multi-label graph learning intends to capture the intrinsic complexity of real-world applications, where one sample is often related to multiple groups or consists of multiple objects. To date, a handful of multi-label graph learning methods exist, but none of them integrate training-time interpreta…
- PaSta: Noisy Node Classification with Partial Label Learning
Yujing Liu, Yixin Liu, Yu Zheng, Yue Tan, Alan Wee-Chung Liew, Shirui Pan · 27 August 2026
Noisy node classification problem is a fundamental yet challenging task for real-world graph-related web services, where node labels are often corrupted or unreliable due to weak supervision or automatic annotation. However, existing methods typically train models based on one-hot labels, which not …
- Adapting Dense Vision-Language Relationships for Multi-label Classification with Partial Label
Cheng Chen, Yifan Zhao, Jia Li · 25 August 2026
Learning multi-label image classification with incomplete annotations is a challenging task that has been widely studied for its superior trade-off between high efficiency and less labor consumption on large-scale datasets. Predominant methods rely on strong prior assumptions to recover the missing …
- H$^2$EDL: Hyper Evidential Deep Learning for Hierarchical Classification
Yuanye Liu, Xiahai Zhuang · 19 August 2026
Fine-grained recognition often involves hierarchical label spaces, where a model may be confident about a coarse semantic concept while remaining uncertain among its descendant classes. Such structured ambiguity requires uncertainty representations that capture both fine-grained classes and intermed…
- FeatureHospital: A Skill-Driven Multi-Agent Framework for Automated Algorithm Customization in Multi-View Multi-Label Feature Selection
Junxuan Li, Zhiqi Chen, Yuzhou Liu, Peng Zhang, Huaxiao Liu · 18 August 2026
Multi-view multi-label feature selection aims to identify a compact and informative feature subset from heterogeneous views while preserving discriminative information for multiple labels. Existing methods are generally developed from specific modeling perspectives and incorporate mechanisms tailore…
- TEAMMix: Taxonomy Enrichment Augmentation and Minority-augmented Mixing Strategy for LLM-enhanced Weak-Supervised Hierarchical Text Classification
Jian Zhang, Zhuohao Yang, Songlin Lei, Bangli Liu, Ziwei Wang, Xufeng Weng, Gehan Amaratunga, Yu Lin, Hongwei Wang · 12 August 2026
Hierarchical Text Classification (HTC), as a critical text mining task, faces challenges such as complex label hierarchies and class imbalance. Existing methods based on large language models (LLMs) struggle to be efficiently applied to this task due to issues like lengthy prompts and loss of label …
- A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models
Wajdi Ben Saad, Safa Madiouni · 12 August 2026
Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recognition, yet aggregate evaluation often hides large differences between high-resource and low-resource languages. Uniform inference policies are simple to deploy, …
- Recent advances in weakly supervised learning: New supervision paradigms, assumption relaxations, and practical solutions
Wei Wang, Gang Niu, Masashi Sugiyama · 10 August 2026
Deep learning has achieved great success in recent years thanks to the availability of high-quality, well-annotated training data. However, this requirement is often not met in real-world applications. Weakly supervised learning aims to train an accurate model with incomplete, inexact, or inaccurate…
Other topics in Artificial intelligence
The topics the OpenAlex classification attaches to the same theme, most active first.
- Large Language Models7,407 papers / 12 months+247%
- Adversarial Robustness in Machine Learning3,552 papers / 12 months+118%
- Reinforcement Learning in Robotics2,519 papers / 12 months+117%
- Explainable Artificial Intelligence (XAI)2,319 papers / 12 months+200%
- Domain Adaptation and Few-Shot Learning2,059 papers / 12 months+67%
- Advanced Graph Neural Networks1,926 papers / 12 months+38%
