Physical Sciences › Computer Science › Information Systems
Information Retrieval and Search Behavior
727 artículos indexados
Los trabajos reunidos bajo este tema exploran cómo los sistemas de información recuperan y organizan datos relevantes, especialmente en el contexto de la inteligencia artificial. Abordan métodos como el Retrieval-Augmented Generation (RAG), donde los modelos combinan búsqueda de información y generación de respuestas, así como técnicas de evaluación y optimización de resultados, tales como el reranking o el anclaje de razonamientos latentes. Estas investigaciones examinan también los comportamientos de los agentes de búsqueda, sus mecanismos de decisión y los desafíos relacionados con la verificación de respuestas, en particular en ámbitos como las finanzas o la programación.
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos45 % · 130 artículos
- China35 % · 101 artículos
- India7,3 % · 21 artículos
- Canadá5,2 % · 15 artículos
- Alemania4,5 % · 13 artículos
- Reino Unido4,5 % · 13 artículos
- Corea del Sur3,8 % · 11 artículos
- Australia3,1 % · 9 artículos
Sobre 288 artículos de este tema con al menos un laboratorio localizado. 44 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- RPTune: Learned Context Curation for LLM Catalog Search
Chuxuan Hu, Hejie Cui, Norman Huang, Shubham Kumar Bharti, Wang-Chiew Tan, Sercan \"O. Ar{\i}k · 2 de octubre de 2026
For small merchant businesses (SMBs) whose catalogs fit within a long-context LLM, full-catalog prompting offers a compelling alternative to multi-stage retrieval designed primarily for large marketplaces with millions of items. However, fitting the full catalog into the context window does not ensu…
- On-Device Commercial Intent Retrieval Under Size, Latency, and Privacy Constraints: A 3 MiB Retrieval System with Typed Egress Boundaries
Hyojung Han · 2 de octubre de 2026
We study commercial intent inference that runs entirely on the user's device, under three constraints frozen before the work began: the downloaded payload under 3 MiB, Tier-0 inference under 20 ms at p95, and no raw text, content embedding, or stable identifier leaving the device. Under them we buil…
- A Matryoshka Hierarchical RAG for Efficient Multi-Hop Question Answering
Gianluca Bonifazi, Christopher Buratti, Michele Marchetti, Federica Parlapiano, Giulia Quaglieri, Davide Traini, Domenico Ursino, Luca Virgili · 2 de octubre de 2026
Retrieval-Augmented Generation (RAG) systems for multi-hop Question Answering (QA) must balance retrieval quality with computational cost. This cost is incurred during indexing time, through the use of expensive Knowledge Graphs (KGs) or Large Language Models (LLMs) to generate summaries, or during …
- AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines
Sumin Lee, Sukmin Cho, Suengjae Lim, Youngjin Kwon · 2 de octubre de 2026
Retrieval-based speculative decoding (SD) drafts tokens by copying continuations from existing text, which suits coding agents that repeatedly reproduce code, logs, and earlier attempts. Yet existing methods fall short in agent pipelines: much of the reusable text is missing from their corpora or st…
- Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning
Meghana Sunil, Shravya V, Shravan Venkatraman, Joe Dhanith PR · 2 de octubre de 2026
Large Language Models (LLMs) have demonstrated remarkable fluency across many tasks but remain limited by their static, parameter bound knowledge and their susceptibility to hallucinating information. Retrieval Augmented Generation (RAG) addresses these issues by incorporating external retrieval int…
- HHR: Hierarchical Hash Retrieval for Efficient LLM Generation
Lianjun Liu, Tiantian Zheng, You Huang, Weiqi Yan, Mingte Qiu, Huazhong Liu, Xiaofeng Zhu, Yunshan Zhong · 2 de octubre de 2026
Efficient long-context inference is essential for large language models (LLMs), yet it poses a severe computational bottleneck. Hash-based retrieval offers an efficient alternative by encoding queries and keys into binary codes and using Hamming distance for key selection. However, this leads to a c…
- What Should an Agent Remember? Disentangling Retention from Retrieval in Bounded-Memory Evaluation
Juli Huang · 2 de octubre de 2026
A persistent agent must decide both what to retain as information arrives and what to surface once a query appears, yet memory evaluations can confound these decisions by comparing methods that differ in both retention and selection. We build a streaming-recall benchmark crossing retention and selec…
- IROH: Insightful Ranking Of Humor using Multi-Stage Hybrid Retrieval with Rationale-Distilled LLM Judges for JOKER 2026 Track Task 1 English
Ana-Maria Luisa Mocanu, Sebastian Mocanu, Ciprian-Octavian Truica, Elena-Simona Apostol · 1 de octubre de 2026
Our team, VANGUARD, presents IROH (Insightful Ranking of Humor), a three-stage retrieval system for JOKER Task 1 English at CLEF 2026, achieving first place on the leaderboard with 0.6347 MAP. Our pipeline combines hybrid sparse-dense retrieval, cross-encoder reranking, and a LoRA-adapted Large Lang…
- Effective Dense Retrieval using Only In-Context Examples
Nour Jedidi, Abdul Basit Ali, Hang Li, Jimmy Lin · 1 de octubre de 2026
Turning decoder-only large language models (LLMs) into strong dense retrievers typically requires some form of retriever training. In this paper, we ask whether LLMs can instead be prompted to produce effective representations for dense retrieval given only a few in-context examples. To answer this,…
- MERGE: Multi-LLM Ensemble for Retrieval via Generative Enrichment
Tzu-I Ho, Yung-Yu Shih, Shang-Yu Su, Dongzhe Wang, Yun-Nung Chen · 1 de octubre de 2026
Large Language Models (LLMs) are increasingly used to enrich user queries in information retrieval (IR) so that a standard retriever such as BM25 can bridge vocabulary gaps with the target corpus. Any single LLM, however, is limited by its training data and architectural biases, and its enrichment b…
- Bridging Semantic Gaps in RAG through Generated Context Knowledge Fusion
Xinkai Du, Chao Lv, Yalin Sun, Quanjie Han, Lei Yao, Maosong Sun · 1 de octubre de 2026
Retrieval-Augmented Generation has established itself as a fundamental framework in natural language processing, seamlessly integrating information retrieval with the generative capabilities of large language models. However, this process is fundamentally constrained by a critical challenge: semanti…
- Traverse: Learning When to Remember, Reset, and Redirect for Long-Horizon Web Search
Jingyuan Ma, Lynx Aster, He Zhang, Siyao Song, Weijie Yuan, Zhe Zhang, Kai Jia, Zhifang Sui · 1 de octubre de 2026
Long-horizon information-seeking agents often accumulate noisy or misleading context, causing early mistakes to persist and making recovery increasingly difficult. We introduce an autonomous search harness in which the agent manages its own search process through three states: Rubric, Answer, and Ve…
- Grounded Revision vs. Prior Injection: Probing Retrieval-Augmented Patent Claim Amendment
Josepha Michiko Leo, Hyun-seok Min, Yehoon Jang, Irvan Zidny, Jin-Woo Chung, Sungchul Choi · 1 de octubre de 2026
Retrieval-augmented generation is widely used in professional writing, yet whether retrieval grounds revision or merely injects templates is rarely tested where "correct" has a definable meaning. Patent claim amendment supplies that signal: the examiner names the attacked limitation and cites prior …
- Beyond the Context Window: An Adaptive Entropy-Based Routing Framework for Hybrid Retrieval and Long-Context Language Models
Isaac Olufadewa, Miracle Adesina, Ezekiel Oladejo, Owen Adeniyi, Fadare Fadekemi, Olamide Oso, Uthman Babatunde, Matthew Olawoyin · 1 de octubre de 2026
Modern large language models now support context windows of more than one million tokens, which has raised the question of whether retrieval-augmented generation (RAG) is still necessary. Pure long-context (LC) processing is expensive and is known to under-attend to information placed in the middle …
- Sieve and Sage: Efficient Distraction Filtering for Reliable RALM Abstention
Jongbin Won, Sung Geun An, Jay-yoon Lee · 1 de octubre de 2026
Just as Socrates recognized the limits of his own knowledge, Retrieval-Augmented Language Models (RALMs) should learn to abstain when the retrieved evidence cannot support a reliable response. Existing approaches largely rely on monolithic LLMs to handle heterogeneous retrieval failures in a single …
- Instruction Retrieval at Inference Time for Small Language Models
Kenan Alkiek, David Jurgens, Vinod Vydiswaran · 1 de octubre de 2026
The facts a language model stores are tied to its parameter count, so small models that fit on edge devices fail on expert problems, which need specialized knowledge and follow multi-step procedures. Fine-tuning for a specific domain or task writes the knowledge into the parameters but must be repea…
- TRACE: Trajectory Selection for Parallel Scaling of Search Agents
Qisheng Zhou, Zhen Xiong, Qiaoyu Tan · 1 de octubre de 2026
Parallel search may generate a correct answer that final-answer voting fails to select. We formulate this consolidation stage as trajectory selection and introduce TRACE (Trajectory Ranking with Aggregated Cross-Rollout Evidence), a lightweight learned selector that ranks completed trajectories usin…
- RAGScope: A Leakage-Controlled, Cost-Aware Evidence-Gating Protocol for RAG Hallucination Triage
Zeming Liu, Qibai Chen, Jingtao Zhang, Hang Lyu · 1 de octubre de 2026
Retrieval-augmented generation (RAG) systems need inexpensive ways to route generated answers: accept low-risk outputs, review uncertain ones, and reserve strong verifiers for the expensive tail. We present RAGScope, a leakage-controlled protocol for evaluating local evidence gates that use only the…
- SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale
Guanqun Yang, Wenlong Zhang, Tian Shi, Ping Wang · 1 de octubre de 2026
Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself…
- Re-ranking and Late Interaction Drive Retrieval Quality: A Controlled Comparison of RAG Strategies for Scientific Question Answering
Bhagyesh Rathi, Eshan Chawla, William B. Andreopoulos · 1 de octubre de 2026
Retrieval-Augmented Generation (RAG) is now the standard way to ground Large Language Models (LLMs) in external knowledge, yet the design space of retrieval pipelines is large and the trade-offs between variants are not well understood, especially on domain-specific corpora at realistic scale. In th…
- TAGGRAPH: Tag-Augmented Graphs for Graph Retrieval of Agent Persistent Histories
Yu-Su Chen, Yu-Jung Liang, Pengtao Xie · 1 de octubre de 2026
Long-term memory lets LLM agents recall past interactions and remain consistent across sessions, but memory systems are hard to compare because they often vary in representation, indexing, retrieval, and evaluation. We present a controlled evaluation framework based on shared 5W-style conversational…
- BELIEFRAG: Making Adaptive RAG State-Aware under Evolving Evidence
Hongji Pu · 1 de octubre de 2026
Adaptive RAG uses signals such as confidence, relevance, support, and retrieval quality to decide when to search or correct evidence. In multi-step retrieval, however, these local signals must be combined into a persistent view of what the current evidence supports, what remains missing, and which a…
- Learning to Route in Visual Space via Multi-Step Embedding Retrieval
Tianyu Chen, Mingyuan Zhou, Jiaxing Wu · 1 de octubre de 2026
LLM agents rely on retrieval tools to access external knowledge, yet visual agentic search remains severely bottlenecked by standard single-step retrievers. In current pipelines, the agent must issue text queries for every intermediate step, struggling when visual clues are difficult to describe or …
- Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems
Maximilian Schall, Sedigheh Eslami, Markus Krimmel, Antoine Chaffin, Louis Milliken, Bo Wang, Denis Bykov · 30 de septiembre de 2026
Evaluating first-stage retrievers in large-scale production RAG requires a benchmark that pairs a large-scale corpus with a large set of agent-reformulated search queries based on real user queries and their conversation threads, and that labels many relevant documents per query. No existing public …
- AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi · 30 de septiembre de 2026
Document rerankers determine what evidence reaches the downstream model in RAG and deep research, yet mainstream rerankers select by relevance matching, and individually relevant documents rarely constitute the complete, complementary, non-redundant set a complex information need demands. Prior work…
Otros asuntos del tema Sistemas de información
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Software Engineering Research853 artículos / 12 meses+218 %
- Recommender Systems and Techniques600 artículos / 12 meses+154 %
- Expert finding and Q&A systems205 artículos / 12 meses+1175 %
- Information and Cyber Security169 artículos / 12 meses+1500 %
- Blockchain Technology Applications and Security146 artículos / 12 meses+650 %
- Big Data and Digital Economy127 artículos / 12 meses+14 %
