Social Sciences › Decision Sciences › Information Systems and Management
Personal Information Management and User Behavior
287 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- Vereinigte Staaten55 % · 64 Artikel
- China41 % · 48 Artikel
- Vereinigtes Königreich11 % · 13 Artikel
- Indien7,7 % · 9 Artikel
- Singapur5,1 % · 6 Artikel
- Deutschland5,1 % · 6 Artikel
- Sonderverwaltungsregion Hongkong4,3 % · 5 Artikel
- Südkorea3,4 % · 4 Artikel
Über 117 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 28 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking
Anjali Kantharuban, Jonas Mueller · 5. Oktober 2026
Recent benchmarks rely on user simulators to evaluate AI agents in multi-turn interaction. While existing simulation techniques demonstrate surface fidelity to human style and behavior, ecologically valid interactive benchmarking also requires alignment in when and how agents fail across simulated a…
- When History Fails to Become Experience: Action Calibration in Language Agents
Jingyu Liu, Zhiwen Wang, Yuxin Jing, Huanyu Zhou, Yong Liu · 5. Oktober 2026
Language agents should draw on prior attempts and environmental feedback to improve subsequent decisions within the same task. However, providing additional interaction history can sometimes reduce task success, suggesting that agents do not consistently use this information effectively. To investig…
- DyadMem: A Long-Term Memory Benchmark of How Agents Work with Users
Yifei Tao, Xinyu Zhong, Henry Hengyuan Zhao, Fanyi Wang, Tengda Guo, Wentao Qiu, Ying Wang, Liujian Tang · 5. Oktober 2026
Long-term agents must remember not only what is true about a user, but also how a particular agent should work with that user as their shared history evolves. Existing benchmarks primarily supervise user facts and preferences or experience reusable across users, leaving this relationship-specific ag…
- Interpreting at Write Time: A Policy Ablation for Multi-Goal Agent Memory
Albert Sadowski, Jaros{\l}aw A. Chudziak · 5. Oktober 2026
A long-running assistant cannot keep everything it has seen, so it summarises. Summarising is not neutral: what is kept is chosen against some notion of what the record is for, and that choice is made once, before anyone knows which of the user's standing goals will ask. Goals rarely disagree about …
- Lost in the Request: How Communication Variation Disrupts Retrieval and Action in Email Agents
Feng Chen, Ritam Dutt, Atnaz Taheri, Alex Williams · 5. Oktober 2026
An email assistant should not complete less work simply because a user phrases the same request differently. Yet most benchmarks test each task with only one canonical request, leaving this form of robustness largely unmeasured. We test whether email assistants remain reliable when the requested inf…
- RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations
Arman Behnam, Sunglyoung Kim, Liangwei Yang · 2. Oktober 2026
A companion that talks with a person for months should come to understand them. It should remember what they said, infer who they are, and know when the past bears on the message in front of it. Testing this requires a real person's record, and such records are private, so benchmarks generate the pe…
- Empty Commitments: When Agents Promise What Their Runtime Cannot Deliver
Jiaqi Tang, Lan Wei, Bingyu Shen, Boyang Li · 2. Oktober 2026
A chatbot that says "I will remind you tomorrow" will not run again until the user writes. We call such a promise an empty commitment: a promise of an action after the current turn that nothing in the agent's tools or runtime can carry out. Unlike a broken promise, its emptiness follows from the age…
- MemFit: Efficient Long-Term Agentic Memory
Mitchell Piehl, Muchao Ye · 2. Oktober 2026
Long-term memory systems for large language models (LLMs) have gained popularity for extending reasoning capabilities across applications. Current memory systems rely on LLM agents to organize and consolidate memory, resulting in costly, inefficient write operations. To address this limitation, we p…
- ReLiveGym: Evaluating Long-Lived Agents over Weeks of Replayed Reality
Xisen Jin, Jingheng Li, Zhenglun Chen, Junyi Du, Xiang Ren · 2. Oktober 2026
As large language model (LLM) agents become widely adopted, they are increasingly deployed for tasks that require persistent monitoring or recurring actions (e.g., market analysis). These agents are expected to operate unattended for days or weeks, act at the right timing, and adapt to the dynamic e…
- Heavy-Tailed Memory Traces in Long-Horizon Language Agents
Xinyuan Song, Zekun Cai · 2. Oktober 2026
Long-horizon language agents increasingly rely on external memory as a frozen world model, yet current memory systems are usually judged only by task success or token cost. We argue that the missing object is the shape of memory use: under finite context and repeated retrieval, agent memory can conc…
- AMU:Admission and Memory Update for Personalized Conversations---Structured Memory with SLM Guided Control
Tao Hwang, Yishi Diao · 1. Oktober 2026
Large language models (LLMs) have become the foundation of personalized assistants, but maintaining persistent user memory across long-term interactions remains challenging. Existing memory systems often focus on storage, retrieval, or consolidation, while memory writing remains less controlled: tra…
- ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context
Jianshu Zhang, Keliang Wu, Chengxuan Qian, Xiyuan Yang, Ce Zhang, Ariel Tian, Anbang Liu, Haoran Lu, Han Liu · 1. Oktober 2026
Embodied agents now take on ever longer tasks. For long tasks, knowing only whether a task finally succeeds or fails says little; the steps along the way matter. Progress Reward Models (PRMs) score how far a task has come at every step, and serve as dense rewards, verifiers and monitors. Yet in long…
- Memory Consolidation Flattens the Temporal Shape of User Facts
Sugam Panthi, Muhaiminul Yeamin, Siyan Luo, Rabab Abdelfattah · 1. Oktober 2026
Long-term memory systems turn conversations into short stored notes. A note can keep a user fact while losing evidence about whether the fact still holds. For example, "I am driving a Peugeot" can become "The user drives a Peugeot," which drops the cue that the activity is ongoing. We call this aspe…
- cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents
Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck, Daniel Fried, Ruslan Salakhutdinov, Jing Yu Koh · 1. Oktober 2026
Computer use agents (CUAs), which use graphical user interfaces (GUIs) to complete tasks on a computer, have recently surpassed human performance on many standard benchmarks, including difficult long-horizon tasks. Their capabilities are undoubtedly impressive, however, a key barrier to the widespre…
- Richard: Voice-First Mobile Interaction for Persistent Tasks
Xinyang Chen · 1. Oktober 2026
Mobile terminals need to provide application and network services while supporting users' control over their attention. We explore voice-first interaction organized around requests and delegated tasks, allowing users to leave a conversation and later inspect, revise, and retrieve the work. We presen…
- Beyond Oracle Communication: Benchmarking Interactive Intent Alignment Under Miscommunication and Evolving User Intent
Zheyuan Zhang, Mengyuan Chao, Ke Xiao, Ziyi Chen, Daoan Zhang, Yan Zhang, Yanfang Ye, Wei Xu · 1. Oktober 2026
Modern LLM agents increasingly tackle complex tasks through interactive, long-horizon exchanges with users, while existing benchmarks generally assume that users always accurately and sufficiently communicate a fixed intent. However, this oracle communication assumption rarely holds in practice: use…
- When Context Changes: Understanding Update Failures in LLMs
Junyu Guo, Yuchen Fang, Shangding Gu, Costas Spanos, James Demmel, Javad Lavaei · 1. Oktober 2026
As preferences, goals, and facts change, LLM agents must use the current state while earlier versions remain in context. Yet they can answer with an old value of the same variable, a failure that we call stale binding. To study when models use outdated information and why, we introduce Controlled In…
- Action Conditioned Bisimulation For GUI Agent Memory
Hongbo Zhang, Liuyang Song, Quanquan Li, Daqian Yang, Yan Wen, Zhengtao Yao · 1. Oktober 2026
An agent that remembers what it did on a web page must decide when two pages count as the same. Memories built on observation similarity merge pages that look alike but behave differently, and GUIs are full of such pages: two tabs of one widget or two rows of one menu answer the same click different…
- LongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual Puzzles
Bingo Zhang, Haochuan Lu, Zongjie Li, Genjian Li, Ari Yu Zhang, Chaozheng Wang · 30. September 2026
GUI agents need long-horizon visual reasoning: they must interpret a changing interface while keeping a multi-step plan viable as earlier actions constrain later ones. Existing benchmarks evaluate grounding, computer use, and game play, but rarely test whether agents stay coherent across long chains…
- Just-In-Time Agent Memory with Runtime Agentic Research
Bingyu Yan, Chaofan Li, Hongjin Qian, Shuqi Lu, Chaozhuo Li, Zheng Liu · 30. September 2026
Memory is critical for AI agents. Many existing agent-memory systems follow an Ahead-of-Time (AOT) design, constructing memory before a specific request arrives. While this reduces online serving cost, such request-agnostic memory construction can discard fine-grained information that later becomes …
- Adapting Context Compression for Long-Horizon Agents with Counterfactual Continuations
Guanghui Min, Liang Wu, Mingjia Shi, Yinhan He, Mayank Darbari, Liangjie Hong, Chen Chen · 30. September 2026
Long-horizon agents require context compression to manage growing interaction histories. Compression quality, however, is ultimately determined by downstream execution. Existing prompt-adaptation methods infer compression errors by comparing full-context and compressed trajectories. Such comparisons…
- Mnemon: Raw Records, Fast Judgments, Slow Thoughts
Guangren Wang · 30. September 2026
Long-term memory lets an LLM assistant use a history it can no longer reread, and most memory systems build it by rewriting conversations into facts, graphs or typed memories at write time. We argue that the work of memory divides, as thinking does, into two systems. Most of it is fast System 1 work…
- When Successful Memories Mislead Embodied Agents:Memory Adaption For Task-Conditioned Execution
Quanquan Li, Hongbo Zhang, Yihe Chi, Liuyang Song, Jingyu Li, Yuxiang Huang, Hongzhen Zhang, Guitao Cao · 30. September 2026
Experience reuse can reduce repeated exploration in embodied agents, but a trajectory that succeeded previously may be unsuitable for the current execution context. Existing memory systems pri marily optimize construction and retrieval; semantic relevance and historical success therefore remain insu…
- UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
Ashish Jain, Armaan Sandhu · 30. September 2026
Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks score only the agent and do not directly measure whether the user c…
- ContextRender: From Execution Dependencies to Agent Context
Savini Kashmira, Jayanaka L. Dantanarayana, Lingjia Tang, Jason Mars · 30. September 2026
LLM agents performing long-horizon tasks accumulate tool results that later steps may need. Passing the full history to every invocation is costly even when it fits within the context window, while reducing it risks omitting needed information. Existing context management methods can overlook how ea…
