Physical Sciences › Computer Science › Artificial Intelligence
Advanced Text Analysis Techniques
67 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Últimos artículos
- FourierQK: Filter Shape, Admissibility and the Leakage-Coverage Law
Athanasios Zeris · 2 de octubre de 2026
Frequency-collapse attention [Zeris, 2026e] achieves large gains over standard dot-product attention by replacing the Q/K dot product with a bandpass-filtered inner product at a learned frequency. A natural follow-up question is: which filter shape works best, and why? We test five hypotheses about …
- Hyperspherical Semantic Trajectory Analysis: Mapping Technological Diffusion across Academic Preprints, Patent Signals, and Compute Scaling
Muhammad Sukri Bin Ramli · 1 de octubre de 2026
Macroeconomic productivity metrics, such as Total Factor Productivity, register technological breakthroughs with multi-year reporting lags due to administrative survey intervals and national accounting conventions. This paper introduces Hyperspherical Semantic Trajectory Analysis (HSTA), an unsuperv…
- Beyond Token Scale: Chunk-Level Sparse Autoencoders for Reliable Semantic Feature Discovery
Xu Wang, Yifan Yang, TingHao YU, Difan Zou · 30 de septiembre de 2026
Sparse autoencoders (SAEs) expose features that help us understand and steer language models, but faithful reconstruction does not guarantee informative concepts. Token-level objectives reward lexical and formatting details alongside semantic content, all competing for a limited sparse budget. We in…
- High-Level Text Preprocessing for Semantic Similarity Analysis of Discursive Texts: A Framework and Empirical Demonstration
Mehmet Murat Albayrakoglu, Mehmet Nafiz Aydin · 30 de septiembre de 2026
Semantic Textual Similarity (STS) methods assume that a document's lexical content faithfully represents what it asserts. This assumption fails for discursive documents that discuss, compare, critique, and contextualize other positions in the process of articulating their own. The result is semantic…
- Query Expansion and Key Specialization in Transformer Attention Geometry
Vidit Gupta, Siddhesh Nadkarni, Mihik Chaudhari, Vinaya Sawant, Prachi Tawde · 29 de septiembre de 2026
The projection of queries and keys are central to the attention mechanism in Transformer architectures. While they are mathematically symmetric, they play different roles in attention mechanisms. The question of whether there is an effect from their functional distinction on their geometric developm…
- Language as an Independent Information Layer: A Conceptual Model of Communication, Cognition and Decision-Making
Anastasiia Alifanova, Elena Benderskaya · 29 de septiembre de 2026
Based on an analysis of the role of language in thought and communication, this article proposes a new concept for designing corporate knowledge bases. The concept integrates the probabilistic vector space of a corporate vocabulary, reflect-ing industry specifics, subject focus, terminology, and cul…
- A Computational Framework for Modelling Organisation-Level Semantic Identity from Longitudinal Textual Data
Brinda Murali Krishna, Oktay Karaku\c{s}, Can Eyupoglu · 25 de septiembre de 2026
Organisations continuously generate large volumes of textual data that capture how they communicate, evolve and differentiate themselves over time. Although recent advances in natural language processing have substantially improved organisation-level text analytics, existing approaches primarily rep…
- Can Classical Semantic-Extractive Summarization Be Evaluated in Hindi? A Replication Study
Showket Ahmad Khan, Mudasir Mohd, Nasrullah Sheikh, Mohsin Altaf Wani, Abid Hussain Wani, Hilal Ahmad Khanday, Niyaz Ahmad Wani · 25 de septiembre de 2026
We replicate the distributional-semantics extractive summarisation method of Mohd, Jan and Shah (2020) and adapt it to Hindi, substituting a Devanagari-appropriate component at every language-specific step. The system is evaluated on two independent corpora --- the Hindi portion of XL-Sum and FIRE I…
- A Semiotics-Aware Framework for Evaluating Fidelity and Coverage in Natural Language Generation
Lorenzo Zangari, Davide Picca · 23 de septiembre de 2026
When two texts describe the same expression, standard metrics based on lexical overlap or whole-text similarity may fail to detect meaningful differences in how that expression is framed. We propose a framework to evaluate semiotic alignment between texts, where a semiotic profile encompasses both t…
- From Outliers to Topics in Language Models: Anticipating Trends in News Corpora
Evangelia Zve, Benjamin Icard, Alice Breton, Lila Sainero, Gauvain Bourgne, Jean-Gabriel Ganascia · 22 de septiembre de 2026
This paper examines how outliers, often dismissed as noise in topic modeling, can act as weak signals of emerging topics in dynamic news corpora. Using vector embeddings from state-of-the-art language models and a cumulative clustering approach, we track their evolution over time in French and Engli…
- Marginal utility, matrix factorization, and the Key-Value (KV) cache: a unified information-economic framework for sovereign geo-mining inference
Caroline Gans Combe (INSEEC) · 18 de septiembre de 2026
This paper builds a theoretical bridge between the economic notion of marginal utility and two machine-learning constructs, matrix factorization and the Key--Value cache of transformer language models. The singular value spectrum of a rating matrix is shown to be a diminishing marginal utility sched…
- Differentially Private Semantic Plans for Aggregate Insight Generation
Behrooz Razeghi · 16 de septiembre de 2026
\texttt{URANIA} provides end-to-end differential privacy (DP) for summaries of data-dependent clusters. However, its cluster--keyword release does not directly provide collection-wide aggregates for semantic concepts defined independently of the protected corpus. Records may express several concepts…
- Single Document Extractive Summarization using Domination in Hypergraph
Aamir Miyajiwala, Aabha Pingle, Sheetal Sonawane, Surajit Kr. Nath · 16 de septiembre de 2026
Automatic Text Summarization (ATS) in Natural Language Processing has been an important task in Information Retrieval. It compresses a document to create a summary that captures all the relevant and important information conveyed in the document. This study explores Hypergraph for extractive text su…
- Cortex: Content Analysis Support Software, a Resource for Qualitative Research
Ana Julia da Silva Soares, Rafael Coimbra Pinto · 14 de septiembre de 2026
Qualitative research is widely used in the human and social sciences, characterized by a deep understanding of phenomena through the interpretation of meanings and contexts. Among qualitative data analysis methods, content analysis stands out as a consolidated technique, which allows for the systema…
- Automated Detection and Structuring of Social Tipping Point Evidence in Climate related Documents: A Modular AI Framework
Kavindu Perera, Mohammad Abaeiani, Ekaterina Gilman, Lauri Loven, Mourad Oussalah, Tassos Kanellos, Beatrice Gobbo, Dante Adami, Nicol\`o Ferriani, Maximiliano Romero, Pierre Rossel, Marc Bonazountas, Christina Deligianni, Nikos Xyderis, Artur Bogucki, Lampros Argyriou, Prasasthy Balasubramanian · 14 de septiembre de 2026
The climate literature has grown faster than review teams can read it. That gap matters most for a concept like the environmental social tipping point, the threshold at which a small change triggers rapid, self-reinforcing change in a social system. Evidence of this kind of shift is usually containe…
- CMNIE: An Information Extraction Benchmark for Chinese Military News
Yan Yu, Mengna Zhu, Zhenyu Song, Hao Yang, Haiwen Chen, Mao Wang · 11 de septiembre de 2026
Structured extraction from Chinese military news supports intelligence analysis, decision-making, and knowledge base construction. However, existing resources provide limited support for joint informa?tion extraction in this domain, especially when events, event arguments, entities, and relations mu…
- Beyond Top Words: MonoTM for Topic Modeling with Interpretable Monosemantic Features
Una Joh, Bei Yu · 10 de septiembre de 2026
Topic models summarize large text corpora, but top-ranked words often provide only a limited representation of topic semantics. Sparse autoencoders (SAEs) offer a way to move beyond word-level descriptors by extracting interpretable features from dense representations, yet how feature interpretabili…
- Event Interaction in Low-Rank Bottlenecks for Temporal Relation Extraction
Wei Sun, Tingyu Qu, Jesse Davis, Marie-Francine Moens · 9 de septiembre de 2026
Temporal relation extraction determines whether an event occurs before, after, or simultaneously with another event, and therefore relies on accurately modeling how the two events interact. Mainstream systems achieve this by concatenating event spans or using shallow fusion, which works well when al…
- Key Path Identification for Resolving Knowledge Conflicts via SAE-based Steering
Wenbo Zhang, Zhongxiang Sun, Zhiguang Han, Jun Xu · 9 de septiembre de 2026
Sparse autoencoder (SAE)-based steering has been widely used to address knowledge conflicts by guiding LLMs to be more faithful to the contextual knowledge. Existing methods usually perform mass steering, which modifies a large batch of SAE features identified via correlation-based methods. However,…
- Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time
Lisa Bouger, Yannick Teglia, Philippe Loubet Moundi · 7 de septiembre de 2026
We propose an influence score to quantify the contribution of attention heads to classification decisions in Transformer-based models designed for prompt injection detection. The score combines directional influence on the logits with structural contribution within the residual stream, enabling a mu…
- DiffIE: Diffusion-based Open Information Extraction
Konstantin Fedorov, Valentin Malykh · 3 de septiembre de 2026
A single sentence often expresses multiple valid relational triplets, which makes Open Information Extraction (OpenIE) fundamentally a multi-output task. Existing neural systems handle this by autoregressive generation, which is flexible but slow and prone to redundancy, or by fixed-slot prediction,…
- TaxCE : A Framework for Automated Taxonomy Construction and Evaluation at Scale
Sandeep Sricharan Mukku, Albert Aristotle Nanda, Rohit Pyati · 1 de septiembre de 2026
Organizing unstructured feedback text into hierarchical taxonomy is a fundamental challenge in NLP, particularly in domains where feedback arrives at massive scale in varied forms such as reviews, transcripts, and surveys. Existing approaches either produce shallow hierarchies, neglect long-tail top…
- ITL: Interpretable Document Alignment with Structured Reference Frameworks
Ra\'ul Gir\'aldez, Dayrelis Mena, Jes\'us S. Aguilar--Ruiz · 28 de agosto de 2026
Measuring alignment between documents and structured reference frameworks requires identifying conceptual evidence distributed throughout the text and reporting it through measures that are quantitative, interpretable, and traceable. Many commonly used retrieval and classification approaches return …
- KinyaEmbed: Contrastive Sentence Embeddings for Kinyarwanda via Multi-Stage Curriculum Training
Ireddi Rakshitha, Devavarapu Yashwanth, Ntakirutimana Pierre · 28 de agosto de 2026
We present KinyaEmbed, the first dedicated sentence embedding model for Kinyarwanda, a morphologically rich Bantu language spoken by over 12 million people in Rwanda. Existing multilingual embedding models such as LaBSE, mE5-large, and OpenAI text-embedding-3-large perform poorly on Kinyarwanda due …
- Key Point Analysis Needs Structure Recovery: Task Definition, Dataset Diagnosis, and a Structure-Aware Benchmark
Zhiqiang Shi, Oana Cocarascu · 27 de agosto de 2026
Key Point Analysis (KPA) aims to identify a concise set of key points that summarize a collection of arguments together with their prevalence. We argue that KPA is fundamentally a structured prediction problem that requires recovering semantic groupings, generating representative key points, ensurin…
Otros asuntos del tema Inteligencia artificial
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Large Language Models7407 artículos / 12 meses+247 %
- Adversarial Robustness in Machine Learning3552 artículos / 12 meses+118 %
- Reinforcement Learning in Robotics2519 artículos / 12 meses+117 %
- Explainable Artificial Intelligence (XAI)2319 artículos / 12 meses+200 %
- Domain Adaptation and Few-Shot Learning2059 artículos / 12 meses+67 %
- Advanced Graph Neural Networks1926 artículos / 12 meses+38 %
