Physical Sciences › Computer Science › Artificial Intelligence
Authorship Attribution and Profiling
284 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- Vereinigte Staaten40 % · 58 Artikel
- China18 % · 26 Artikel
- Deutschland8,9 % · 13 Artikel
- Vereinigtes Königreich7,5 % · 11 Artikel
- Indien5,5 % · 8 Artikel
- Frankreich4,8 % · 7 Artikel
- Italien4,8 % · 7 Artikel
- Japan4,1 % · 6 Artikel
Über 146 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 46 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- Using LLMs to Detect LLM-Generated Texts: A Cross-Generation Analysis
Haiyue Yuan, Jie Guo, Weidong Qiu, Zheng Huang, Ruizhe Li, Shujun Li · 30. September 2026
Automated detection of LLM-generated texts (LGTs) is critical, yet dedicated detectors often struggle to generalize across domains and models. While general-purpose LLMs offer flexible zero-shot authorship classification with explanatory rationale, their detection behavior, especially regarding self…
- LA-CPD: Local-Evidence-Aware Change-Point Detection for Human-LLM Authorship Segmentation
Qing Yang, Zhenyu Mao, Zixiang Luo, Zezheng Wu, Xinghe Cheng, Qinggang Zhang, Jingwei Zhang, Jiapu Wang · 30. September 2026
As LLM-generated text becomes increasingly human-like, accurately localizing LLM-authored spans in human-LLM co-authored documents is important for attribution and accountability in cases involving copyright infringement, fraud, and other harmful uses of AI-generated content. Sentence-level detector…
- C-HAT-Bench: Benchmarking Chinese AI-Text Detection Beyond Fully Generated Text
Qing Yang, Zixiang Luo, Zhenyu Mao, Zezheng Wu, Xinghe Cheng, Haibo Chen, Qinggang Zhang, Jiapu Wang, Jingwei Zhang · 30. September 2026
Large Language Models (LLMs) increasingly participate in writing by modifying or extending human drafts, causing machine involvement to vary in both form and extent. Yet most Machine-Generated Text (MGT) detectors are evaluated only on fully human-written versus fully AI-generated text. Because huma…
- MixDetect: Word-Level Localization and Quantification of AI Editing
Hongrui Bao, Yubing Ren, Zhendong Pan, Fang Fang, Shi Wang, Yanan Cao · 30. September 2026
Large language models are increasingly used to edit human-written text rather than generate entire texts from scratch. Conventional AI-text detectors mainly distinguish human-written from fully AI-generated text, while recent methods for AI-edited text typically provide only a text-level label or ed…
- Agents Can Use Base Models to Evade AI Detection
Bhuwan Dhingra, Danish Pruthi · 30. September 2026
We show that coding agents equipped with a base language model can successfully assemble responses from its samples to evade detection. Base models have been shown to evade commercial detectors, however, prior "humanization" techniques rely on using these models to paraphrase AI outputs over several…
- Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models
Mir Tafseer Nayeem, Davood Rafiei · 29. September 2026
Names are personal identifiers, but they also carry social meaning and are widely used to evaluate how language models treat different people. Such evaluations typically assume that matched names are comparable model inputs. We show that this assumption often fails at the lexical interface: matched …
- Identifying Scientists on X
Philipp Meier, Katarina Boland, Laura Kallmeyer, Stefan Dietze · 28. September 2026
With the growing importance of science-related discourse on the Web and the erosion of the classical knowledge order, it is important to identify different user groups, such as scientists, automatically. This work proposes an approach for identifying scientists and non- scientists on X/Twitter based…
- FAVoR: Measuring and Mitigating Author-Style Homogenization in Federated Personalized Generation
Lu Han, Jingyao Zhang, Katy Ilonka Gero, Nguyen H. Tran · 28. September 2026
Large language models are increasingly used as personalized writing assistants, but adapting a model across many authors can compromise individual writing style by pulling author-specific signals toward a shared register. Federated parameter-efficient fine-tuning (PEFT) offers a data-local setting f…
- A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID
Pawe{\l} Blicharz, Mi{\l}osz Grunwald · 28. September 2026
AI-generated text detectors achieve high accuracy on standard benchmarks, yet the internal representations that drive these predictions remain poorly understood. We study which neurons in a frozen BERT-base-uncased encoder support AI-text detection, using the RAID benchmark across six generators spa…
- Robust Detection of LLM-Generated Text under Contamination
Jiaxun Li, Saptarshi Chakraborty, Ambuj Tewari · 25. September 2026
We study the detection of LLM-generated text under editing and contamination. Modeling human and machine text as finite-order Markov processes with Huber contamination, we characterize an exact boundary for reliable detection under our assumptions. Detection is impossible when contamination is suffi…
- Style, Not Self: Surface Cues Explain Zero-Shot Code Attribution by Large Language Models
Ehsan Barkhordar, Surendrabikram Thapa · 25. September 2026
If a language model can recognize code it wrote, it may favor that code as a judge, and instances of one model monitoring each other could collude. We test this zero-shot on current commercial models. Five LLMs generate solutions to MBPP, HumanEval, and DS-1000, seven more to MBPP, and models act as…
- Contrastive Learning for Authorship Verification
Peter Kirby · 24. September 2026
Our results show that contrastive learning outperforms a classification-based approach to authorship verification under the tested settings. We identify loss function, batch size, training duration, pre-trained model, input context length, and random text span data augmentation as important factors …
- Beyond Poetry: Can Large Language Models Generate Classical Arabic Maqamat?
AbdulRahman A. Morsy (Department of Computer Science, School of Engineering and Applied Sciences, George Washington University, Washington DC, United States), Aya Zirikly (Department of Computer Science, School of Engineering and Applied Sciences, George Washington University, Washington DC, United States, Center for Speech and Language Processing, Whiting School of Engineering, Johns Hopkins University, Baltimore MD, United States) · 24. September 2026
Large language models (LLMs) have shown strong performance in creative text generation, yet their ability to produce culturally grounded and stylistically constrained literary forms remains underexplored. Prior work has focused largely on modern language varieties and poetry, while classical prose t…
- AraGenre 2026: A Hierarchical Definition-Guided Arabic Genre Classification Shared Task
Mo El-Haj, Saad Ezzini, Shadi Abudalfa, Mustafa Jarrar, Nguyen Minh Chi, Nguyen Minh Quan · 24. September 2026
AraGenre is a shared task on hierarchical, definition-guided Arabic genre classification, motivated by the limited availability of annotated data in Arabic and other low-resource languages. Systems assign each Arabic text segment both a broad communicative genre and a fine-grained specific genre. Th…
- The Challenge of Identifying the Origin of Black-Box Large Language Models
Ziqing Yang, Yixin Wu, Yun Shen, Wei Dai, Michael Backes, Yang Zhang · 23. September 2026
The tremendous commercial potential of large language models (LLMs) has heightened concerns over their unauthorized use. To address this, we focus on the task of identifying the origin of black-box LLMs. We further propose PlugAE, an effective and efficient identification method that proactively lev…
- Identifying Intelligent Processes via Online Sequential Testing
Aritra Das, Debayan Gupta · 23. September 2026
Active sequential hypothesis testing studies how to identify an unknown hypothesis with a given set of sensing actions. We study this in the setting of identifying large language models (LLMs), \textit{i.e.}, if a user is conversing with an LLM drawn from a known set of models, how can they identify…
- Detecting GPT-Assisted Writing Using Interpretable Stylometric Features
Rajesh Kumar, Nabeel Siddiqui, Alexander Fuchsberger · 23. September 2026
Distinguishing GPT-assisted from independently authored student writing has become a critical challenge in academia. This paper evaluates the discriminative capability of interpretable stylometric features extracted solely from submitted text. Using data from 90 participants who wrote both independe…
- Authorship identification under domain shift: a survey of stylistic measures and learned author representations
Haining Wang · 22. September 2026
Authorship identification uses patterns in writing to infer who wrote a text, but those patterns also reflect topic, genre, and register. This survey argues that topic-independence is not a property of a stylistic feature but of the feature together with its encoding, its scoring rule, and the evalu…
- Fairness Beyond Anonymization? Demographic Leakage in German LLM-Generated Resumes
Charlotte Leininger, Helena Veit, Matthias A{\ss}enmacher, Andreas Bender · 22. September 2026
Large language models (LLMs) are increasingly integrated into AI-assisted hiring pipelines, including automated resume generation and screening. Under the EU AI Act, the hiring domain is classified as high-risk, making fairness and transparency critical requirements. Existing work has primarily focu…
- Privacy Personalization Trade offs in LLMs: The Impact of Stylometric Signal Reduction on User-Specific Text Generation
Muhammed Nazmul Arefin, Omar Jamal Hammad · 22. September 2026
Large language models (LLMs) have demonstrated the ability to generate user-specific text with high stylistic fidelity. However, the personal data that enables such personalization frequently embeds demographic, cultural, and stylistic markers that raises concerns about stylometric re- identificatio…
- SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler (Sitefire) · 18. September 2026
Word-level detectors identify unedited AI-generated text almost perfectly, but the literature documents their brittleness under rewording, and a word-level score neither characterizes a text nor identifies which AI model wrote it. We ask whether AI-generated text can be identified one level deeper, …
- Sampling Reveals Style: Unsupervised, Training-Free Discovery of Prompt-Conditional Stylistic Axes in LLM Activations
Ajit Mallavarapu, Ziwei Gu · 18. September 2026
Large language models (LLMs) encode rich stylistic structure in their hidden activations, but discovering which stylistic dimensions are salient for a given prompt typically requires supervised contrastive data. We present a training-free, prompt-conditional alternative: we repeatedly sample complet…
- Fingerprinting Multimodal Large Language Models
Chao Huang, Meng Tong, Kejiang Chen · 18. September 2026
While multimodal large language models (MLLMs) enable a wide range of image-text reasoning tasks, recent incidents indicate that they are vulnerable to illicit deployment and unauthorized distillation. Existing solutions for model provenance are typically confounded by shared language backbones in M…
- Is Luke the Author of a Gospel and the Acts of the Apostles?
Jacques Savoy · 17. September 2026
According to Christian tradition, Luke is credited with authoring a Gospel and the Acts of the Apostles, even if his name does not appear in either book, both originally written in Koine Greek. Several biblical scholars assume that both texts were written by a common author, while others deduce the …
- Beyond the Name: Demographic Leakage in De-Identified R\'esum\'es and Evaluation Artifacts in LLM Bias Audits
Qiangju Chen, Yang Xiao · 16. September 2026
De-identified r\'esum\'e screening assumes that redacting explicit fields prevents ethnocultural inference; however, recent audits attribute residual leakage to declared languages. We investigate whether eliminating language fields resolves this leakage across nine open-weight models and 620 counter…
Weitere Unterthemen aus Künstliche Intelligenz
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Large Language Models7.407 Papiere / 12 Monate+247 %
- Adversarial Robustness in Machine Learning3.552 Papiere / 12 Monate+118 %
- Reinforcement Learning in Robotics2.519 Papiere / 12 Monate+117 %
- Explainable Artificial Intelligence (XAI)2.319 Papiere / 12 Monate+200 %
- Domain Adaptation and Few-Shot Learning2.059 Papiere / 12 Monate+67 %
- Advanced Graph Neural Networks1.926 Papiere / 12 Monate+38 %
