Physical Sciences › Computer Science › Information Systems
Expert finding and Q&A systems
205 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- Vereinigte Staaten40 % · 32 Artikel
- China35 % · 28 Artikel
- Vereinigtes Königreich9,9 % · 8 Artikel
- Italien7,4 % · 6 Artikel
- Deutschland4,9 % · 4 Artikel
- Singapur4,9 % · 4 Artikel
- Israel4,9 % · 4 Artikel
- Kanada4,9 % · 4 Artikel
Über 81 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 28 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- System Attribution in LLM Brand Recommendations: Single Responses Identify the System, Aggregated Brand Profiles Do Not Transfer
Dmitrij \.Zatuchin · 2. Oktober 2026
Audits of AI visibility summarise the brand recommendations of deployed language models into per-system profiles. We test whether such a profile describes the system on one corpus of 6,475 stored responses (6,324 analysable) collected between December 2025 and February 2026 from five deployed endpoi…
- Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation
Noy Sternlicht, Simra Shahid, Peter Jansen, Daniel S. Weld, Pao Siangliulue, Tom Hope · 2. Oktober 2026
Automated ideation systems are often evaluated on the novelty of the ideas they produce, and that judgment is increasingly delegated to large language models. Such judges are typically built ad hoc and validated, if at all, on human-authored papers rather than on the generated ideas they are meant t…
- What Wins a Vote? Formatting, Length, and Lexical Diversity in the French Compar:IA LLM Arena
Simonas Zilinskas, Maayeesha Farzana, Christophe Benavent · 2. Oktober 2026
LLM arenas turn pairwise human preferences into model rankings. Those preferences may reflect how an answer is presented as well as what it says. We take a stylometric approach to 137,293 decisive French-language votes from the July 2026 Compar:IA release; the primary formatting analysis includes 13…
- MatrixReward: Reward from Rubric Matrix for Open-Ended Generation
Zihan Shen, Qi Liu, Zixuan Yang, Yiqun Chen, Chenglong Zhao, Xiaozhao Wang, Lei He · 2. Oktober 2026
Open-ended query generation lacks standard answers, thus necessitating an effective reward mechanism. Pointwise scoring rubrics provide limited information about the relative quality of sample answers under the same prompt; merging multiple rubric judgments into a single score may also mask the diff…
- Pair Difficulty Matters: Rethinking Pairwise LLM-as-a-Judge Evaluation and Consistency
Bruno Brocai, Maria Becker · 1. Oktober 2026
Large Language Model judges are widely used to rank texts and text-generating systems through pairwise comparison, and their reliability is typically assessed via three proxies: position bias, transitivity, and pairwise agreement (self- or human-labeled). Because these proxies drive judge selection …
- Your Benchmark Is Not Saturated: Reviving Multiple-Choice Evaluation with Answer Pooling
Mohamed Eltahir, Abobaker Ahmed, Nawaf Barebood, Hussain Bu Subayt, Tanveer Hussain, Naeemullah Khan · 1. Oktober 2026
Multiple-choice benchmarks are cheap to grade and are running out of room, and the standard remedy, writing harder items, is slow and repeated for every benchmark. A saturated benchmark still holds a harder task. Each question's wrong options are written for that question alone, so a model can score…
- Generating Edit-Inducing Questions for AI Research Manuscripts
Sebastian Joseph, Zichao Wang, Jennifer Healey, Alexa Siu, Junyi Jessy Li, Ani Nenkova · 1. Oktober 2026
We study the ability of LLMs to generate edit-inducing questions whose answer will improve a paper draft. On a dataset of paired submission and camera-ready papers from ICLR and NeurIPS, we compare the helpfulness of questions from GPT models with or without full paper context to that of human revie…
- Community-Driven API and AI Writer Design for Openly Scaling Community Notes
Brad Miller, Jay Baxter, Jiansong Chao, Keith Coleman, Sophie Hilgard, Daniel Ortiz · 1. Oktober 2026
Community Notes is a crowd-sourced approach for adding context to posts on X. Contributors propose and rate notes, forming the inputs to an open-source, open-data algorithm that determines which notes show broadly to users. Since September 2025, Community Notes' AI Note Writer API has provided an op…
- SEPAL: Separated Expert Pairs with Answer-Level Fusion for Reliable LLM Collaboration
Weijie Ren, Yanwen Zhang, Hao Li, Zhuolin Qi, Hengyi Zhang, Naibo Wang · 1. Oktober 2026
Multi-agent collaboration lets large language models (LLMs) improve question answering through deliberation and feedback. Yet shared discussion couples correction with exposure to the same mistakes, which can erode the diversity needed for voting. Self-consistency offers sampling diversity without f…
- False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
Meijia Chen, Hao Li, Zheng Lu, Hongshan Lin, Junbai Tian, Yichen Liu, Zijun Tian, Yufan Zou, Shuhan Sun, Hanxin Chen, Zeyu Zhang, Weizhi Du, Yueting Li, Tianyu Shi, Alaa Khamis · 1. Oktober 2026
Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward imp…
- Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics
Maoqi Liu, Junwei He, Bowen Zhang, Feiran Li, Wentao Ma, Rongyi Lin, Shuhan Zhong, Quan Fang · 1. Oktober 2026
Rubric-based reinforcement learning (Rubric-RL) trains language models where no verifier exists. A judge checks each criterion of a rubric, and the verdicts are aggregated into a reward, most often by a weighted sum. We show that this additive aggregation is the weak point. Under a sum, criteria com…
- ConflictGuide: AutoResearch Improves When Competing Behaviors Are Made Visible
Binqian Xu, Qiran Zou, Xiangbo Shu, Dianbo Liu · 1. Oktober 2026
When designing machine learning models, desirable properties are often in tension: improving one behavior can impair another, so task progress can depend on alleviating the conflict. LLM-based AutoResearch systems, which iteratively edit model code and retain edits based on scalar task-performance f…
- ArchitectureIQ: On the Measure of Training Intuition
Zirui Ren, Shaoyang Guo, Chencheng Tang, Jinxin Wang, Chengyu Xiong, Shanbin Yu, Peihang Li, Yidi Wu, Bangzhe Huang, Qingyu Qu, Leqian Yang, Ziming Liu · 1. Oktober 2026
Top researchers have good intuition, but do language models have as good intuition about model training as top AI researchers? To measure model intuition of LLMs and humans, we introduce the ArchitectureIQ benchmark. Each question presents a synthetic dataset and several training recipes, and the te…
- Multi-LLM Collaborative Alignment via Stackelberg Games
Christina Hahn, Shangbin Feng, Dean Light, Swastik Roy, Hila Gonen, Yulia Tsvetkov · 1. Oktober 2026
A pool of language models can collaborate and improve collectively by learning from one another's responses. These interactions depend on the instructions used during training. Existing methods typically sample instructions uniformly, even though their usefulness may change as the models improve: an…
- Reach Into The CHOIR: Free-List Elicitation Uncovers Distinct Model Voices in LLM Ensembles
Ben Wigler, Maria Tsfasman · 1. Oktober 2026
Open-ended LLM homogeneity can create false plurality when several systems appear to offer independent perspectives while returning the same familiar default. Single-pass answers obscure the distinction between agreement produced by a tightly constrained answer space, prompt-vocabulary echo, and bro…
- Navigating the Changing Landscape of Online Knowledge Consumption and Production in the Age of Generative AI: Evidence from Stack Overflow
Ji Eun Kim, L\'ea Vitale, Libby Hemphill, Yulin Yu · 1. Oktober 2026
Online knowledge communities rely on a division of epistemic labor between users who seek information and those who produce it. Generative AI may blur these roles, but how it reallocates knowledge-seeking and knowledge-producing activities and reshapes the nature and returns of participation remains…
- Rubric Rewards from Item Response Theory
Milad Yazdani, Yaser Souri, Xiren Zhou, Pranit Chawla, Dena Shahriari, Subhojit Som, Xia Song · 30. September 2026
Many language tasks have no single answer that can be checked automatically. Rubrics provide criteria for judging responses to these tasks. For reinforcement learning, the resulting verdicts must be combined into a scalar reward. A common approach sums the points assigned to satisfied criteria. Dist…
- Certified Selective Automation of LLM Agent Evaluation
Chengguang Gan, Yunhao Liang, Qinghao Zhang, Shiwen Ni · 30. September 2026
Evaluating LLM agents still ends with a human reading trajectories, because automatic judges carry no guarantee on how often they are wrong. We ask the operational question: what fraction of agent evaluation can a judge take over, with a certificate that the error rate among auto-decided trajectorie…
- Does CoT-Pass@k Really Check the CoT? A Multilingual Mathematical Audit
Tar{\i}k Tuna Ta\c{s}alt{\i}, Burcu H\"udaverdi, David Semedo · 30. September 2026
Pass@k measures whether a model reaches a correct answer under repeated sampling, but never how: a lucky guess counts the same as sound reasoning. CoT-Pass@k was proposed to close that gap, adding an LLM-as-judge that must assess a solution's reasoning chain before it counts. Its value rests entirel…
- Evaluating and Benchmarking the System One Model Jev
Tobias Deu{\ss}er, Lorenz Sparrenberg, Rafet Sifa · 30. September 2026
Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed options, a position on a rubric, or the probability that a statement is true, with probabilities the vendor describes as calibrated. Such models target…
- PADM\'E: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators
Cheng Chang, Yining Mao, Peng Qi · 30. September 2026
Language models are frequently employed to evaluate other language models. An LM evaluator scoring agentic behaviors across multiple criteria is valuable, provided that its decisions align with human judgment. We call the problem of evaluating this alignment Meta-Evaluation. Tackling it directly is …
- How to Run Statistics over LLM Judges and Trust the Results: Calibrated Inference for Small-Sample AI Evaluation with evalstats
Ian Arawjo · 30. September 2026
Researchers across academia increasingly base significance claims on LLM judge scores and small-sample AI evaluations. Yet without well-calibrated confidence intervals (CIs), hypothesis tests, and judge-bias corrections, such claims are unreliable. We address these issues in several contributions. F…
- Binarization Flattens the Score Space
Jacob Cole · 30. September 2026
Large language model (LLM) judges are often used as rewards to train policies on objectives that deterministic verifiers cannot capture. However, these rewards are often collapsed to pass/fail ({0, 1}), which reports the verdict but not how well a response met each criterion. We model each pass/fail…
- Breaking the Illusion of Review Reliability under Static Evaluation: SCOPE Fuzzing for LLM-based Scientific Reviewers
Zhuo Chen, Hao Zeng, Jiawei Liu, Guoxiu He, Le Cai, Liu Haotan, Li Wenbo, Yong Huang, Wei Lu · 30. September 2026
The rapid growth of submissions and reviewing workload has accelerated the use of large language models (LLMs) in peer review. Prior studies suggest that LLM-based reviewers can penalize content perturbations, such as overclaiming, indicating a certain degree of reliability. Yet these conclusions ar…
- From Granular Revision Operations to Meaningful Revision Units: Evaluating LLMs for Revision Boundary Detection
Yu Tian, Andrew Potter, Katerina Christhilf, Motahareh Darvishpour Ahandani, Jessica Early, Steve Graham, Danielle S. McNamara · 29. September 2026
Revision traces provide valuable evidence about students' writing processes, but their usefulness for learning analytics depends on how individual revisions are represented. Automated draft-comparison methods often produce granular edit operations that can fragment a single purposeful revision into …
Weitere Unterthemen aus Informationssysteme
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Software Engineering Research853 Papiere / 12 Monate+218 %
- Information Retrieval and Search Behavior727 Papiere / 12 Monate+506 %
- Recommender Systems and Techniques600 Papiere / 12 Monate+154 %
- Information and Cyber Security169 Papiere / 12 Monate+1500 %
- Blockchain Technology Applications and Security146 Papiere / 12 Monate+650 %
- Big Data and Digital Economy127 Papiere / 12 Monate+14 %
