Physical Sciences › Computer Science › Information Systems
Web Data Mining and Analysis
110 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- Vereinigte Staaten59 % · 34 Artikel
- China31 % · 18 Artikel
- Vereinigtes Königreich12 % · 7 Artikel
- Indien10 % · 6 Artikel
- Singapur6,9 % · 4 Artikel
- Japan6,9 % · 4 Artikel
- Sonderverwaltungsregion Hongkong5,2 % · 3 Artikel
- Schweden5,2 % · 3 Artikel
Über 58 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 27 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories
Chengguang Gan, Zimeng He, Yoshihiro Tsujii, Ken-ichiro Kobayashi, Hiroki Itoh, Kotaro Funakoshi · 2. Oktober 2026
Web agents are an important application of large language models, yet their evaluation often depends on rule based or language model evaluators that inspect only the final outcome. Human verification of task completion and detailed analysis of failed trajectories remain limited. We audit all 165 Web…
- Constructing Challenging Browser-Use Tasks by Controlled Environment Interventions
Xunjian Yin, Tianchen Guan, Jinao Wang, Weili Cao, Daisy Xinlei Lin, Royce Cheng-Yue, Keagan Long, Kyle Wong, Bhuwan Dhingra, Xiangjun Wang, Shuyan Zhou · 30. September 2026
As browser-use agents improve, benchmarks keep pace by collecting new tasks, websites, and applications, often making tasks longer or more novel. This makes difficulty expensive to refresh and difficult to control: when many aspects change at once, it is unclear what actually makes a task challengin…
- ReScraper: Unified Scraping and Cleaning of Web Data for Effective LLM Pretraining
Zichun Yu, Jiarui Yan, Shlok Sanghvi, Nihar Atri, Chenyan Xiong · 29. September 2026
LLM pretraining corpora are normally cleaned by a stack of hand-written heuristics. A heuristic scraper extracts the main content from HTML, and dozens of rule-based filters then clean it, so corpus quality is capped by the coarseness and accuracy of the rules. In this work, we propose ReScraper, a …
- WebPageBench: Event-Level Verification and Controlled UI-Variant Generation for Web Agents
Anton Emelyanov, Maria Tikhonova, Zaven Martirosian, Sergei Averkiev, Alena Fenogenova · 29. September 2026
We present WebPageBench, an open framework for evaluating web agents in which every task is verified from the interface's own event log. Six instrumented mock sites with brand identifiers removed (a marketplace, a bookstore, a grocery service, rail ticketing, hotel search and a document cabinet) emi…
- AX is the New AEO
Ido Finder, Assaf Elovic, Gad Shalev · 29. September 2026
In 2023, AI models answered from training data and hallucinated when it ran out, and businesses were told to seed that knowledge. Models' training knowledge has since given way to live web search, and the advice followed it there: answer-engine optimization, or AEO, now tells businesses to scatter b…
- Probe to Act: Elevating Browser-Use Agent via Active Visual Probing
Keliang Li, Heng Wang, Chen Hu, Daxin Jiang, Hong Chang, Shiguang Shan · 29. September 2026
Browser-use agents require seamless alignment between structured web metadata and visual information, while preserving relevant context across long interactions. Existing interfaces often rely on either screenshot-level action prediction or static Set-of-Marks overlays, leaving the model to resolve …
- From Trajectories to Grounded Preferences: Process Preference Synthesis via Interaction Element Graphs for Web PRMs
Yangzhe Peng, Xiaoyang Wang, Yiyang Zhao, Lijun Wu, Kun He · 29. September 2026
Comparative Process Reward Models (PRMs) provide critical step-level guidance for autonomous web agents by evaluating state-conditioned preferences between candidate actions. However, existing preference training data synthesized via multi-policy sampling suffers from a severe scarcity of Grounded M…
- Epstein Files Engine: Agentic Search for Investigative Journalism
Duy K. Nguyen, Teresa Mondr\'ia Terol, Dylan Freedman, Zach Seward · 28. September 2026
On Jan. 30, 2026, the U.S. Department of Justice released a mixed-media collection concerning Jeffrey Epstein, including about three million pages of PDFs. We describe the Epstein Files Engine, an A.I. agent The New York Times deployed to investigate the files. The Engine translated reporter questio…
- Improving LLM-based Autonomous Web Agents with Filtering
Zhitong Guo, Jing Yu Koh, Ruiyu Li · 24. September 2026
Autonomous web agents, powered by Large Language Models (LLMs), have garnered significant attention for automating various web-based tasks with multi-step reasoning and decision-making capabilities. An open research question in the development of these agents lies in the format of the webpage input.…
- Guides That Cause Actions: An Offline Study of Guide-Action Mutual Reinforcement in Multimodal Web Agents
Chengguang Gan, Yunhao Liang, QingHao Zhang, Shiwen Ni · 24. September 2026
Web agents are usually evaluated in live environments, where environment state and judge models drift between runs, so the same checkpoint rarely reproduces the same score, making controlled studies of training phenomena impractical. We present WebMRE, an offline benchmark of 541 tasks and 5,293 ste…
- Omni2Web: Benchmarking Audiovisual Website Development
Minghao Han, Zhenghao Xing, Xize Cheng, Yuxuan Wang, Junming Lin, Ling Wang, Yinsong Yan, Yunfei Chu, Qize Yang, Jin Xu · 22. September 2026
Screen-recorded web editing requests contain weak deictic expressions such as ``this'' and ``there,'' whose referents depend on speech, cursor trajectories, page state, and edit history. Such requests require intent recovery beyond the explicit specifications assumed by many existing web-editing ben…
- You Can Tell Who's Asking: What the Web's Questions Are Made Of, and Where They Come From
Calvin Zhou, Vincent McCloskey, Krishna Srinivasan · 22. September 2026
Questions scraped from the web are used across academia and industry as a proxy for what people want to know. Across QA training data, retrieval benchmarks, and content strategy, questions on a page are assumed to reflect human intent. We test this assumption at scale by extracting 13.4B question oc…
- EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data
Yinzhu Quan, Zefang Liu · 18. September 2026
Web agents often revisit the same sites, yet most evaluations discard the procedures learned in earlier successful interactions. We introduce EconSkills, a skill library and evaluation framework that distills verified EconWebArena trajectories into parameterized standard operating procedures for ret…
- Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses
Mahsa Amani, Seungeon Lee, Abhisek Dash, Asmaa El Fraihi, Yunah Jang, Elisabeth Kirsten, Qinyuan Wu, Krishna P. Gummadi, Manish Gupta, Abhilasha Ravichander, Muhammad Bilal Zafar, Soumi Das · 18. September 2026
Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood. We present the first study of Web search across four major conversational platforms (ChatGPT, Claude, Grok, and DeepSeek), combining real-world user interactions (invi…
- Quantifying Organizational Environmental Action from Web Data and Large Language Models
Quinn Reynolds, Daniel Shore, Vianey Leos Barajas, Tanhum Yoreh, Meredith Franklin · 16. September 2026
Quantifying organizational environmental action from publicly available web content remains a challenging environmental data science problem because relevant information can be dispersed across multiple webpages and is primarily communicated through unstructured text. We present a scalable computati…
- AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents
Xinyun Cao, Adriana Szekeres, Fazle Elahi Faisal · 15. September 2026
Web agents can utilize reusable tools to reduce the cost and latency of low-level browser interaction, but automatically discovered tool collections can be large, redundant, and poorly aligned with user demand. We present AutoTailor, a meta-agentic framework for constructing and maintaining a compac…
- Token Efficient Task Execution via Application Behavior Modeling for Web Agents
Alexandru Ianta, Eleni Stroulia · 15. September 2026
The strong performance of AI Agents across an impressive variety of tasks is driving an unprecedented investment in agentic infrastructures, however the cost of processing tokens is fast increasing. Web agents automate the execution of web-application tasks described in natural language, by analyzin…
- terms.txt: A Consent and Compensation Protocol for Agentic Web Access
Rajarshi Chowdhury · 11. September 2026
The open web ran on an unwritten bargain: sites admitted crawlers, and search engines sent visitors back. Public measurements show that bargain breaking under AI crawlers and agents. Automated clients now make up most requests, training dominates Cloudflare-classified crawling, and the largest AI pl…
- Query Brand Entity Linking in E-Commerce Search
Dong Liu, Sreyashi Nag · 10. September 2026
Associating user search queries with the correct brand entity is critical for e-commerce product retrieval, yet remains challenging due to the brevity of queries (three to four words on average), their lack of grammatical structure, and a catalog of hundreds of thousands of distinct brands. We formu…
- Data Scout: Targeted Web Crawling for Domain-Specific Pretraining Corpora
Chirag Garg, Eelaaf Zahid, Farhan Ahmed, Jay Pankaj Gala, Eric Butler, Heiko Ludwig · 9. September 2026
The dominant approach to building domain-specific pretraining corpora is to filter large web archives such as CommonCrawl. This works well for popular domains but breaks down for specialized ones, where relevant content is sparse and often beyond the reach of popularity-driven crawlers. We present D…
- SCAFFOLD: Self-Improving Web Agents via Recursive Parametric Skill Abstraction
Bowei He, Xiaokun Zhang, Meng Ding, Xue Liu · 9. September 2026
Web agents need to navigate visually rich, long-horizon interfaces that change across sites, yet most previous agents still learn each task in isolation and discard the procedural knowledge they accumulate. Recent skill-augmented frameworks take an important first step, but they treat the skill libr…
- Discriminative World Models for Web Agents
Kelvin Li, Dhruv Pendharkar, Anish Pahilajani, Chuyi Shang, Leon Oks, Leonid Karlinsky, Rogerio Feris, Trevor Darrell, Roei Herzig · 3. September 2026
Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resulting web states, and ranking them with a ranker model or a Process Reward Model (PRM). These world models are typically trained via supervised next-state prediction to generate fixed …
- Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision
Sitong Pan, Yipeng Shen, Yilin Lu, Caiwen Ding, Lu Cheng, Qianwen Wang · 3. September 2026
Reliable web-agent monitoring is difficult when model-internal uncertainty signals such as token logits are unavailable. In this work, we study prefix-level risk prediction for web agents using observable trajectory signals: given an evolving prefix, estimate whether the current execution remains on…
- Web Price Extraction: State of the Art and an Adaptive Browserless Implementation
Evgeniia Kositsyna, Jorge Lloret-Gazo · 2. September 2026
Price extraction from websites is a key task for market monitoring, price comparison, and business analytics in e-commerce. Existing approaches can be broadly divided into four groups, and understanding their trade-offs in accuracy and scalability is essential for selecting suitable extraction strat…
- WebWorld: The Browser as a World Model for Self-Improving Web Code
Jiajun Wu, Jian Yang, Yaxin Du, Wei Zhang, Haowen Wang, Junhang Cheng, Yuxuan Zhang, Tuney Zheng, Xianglong Liu, Ming Zhou · 1. September 2026
VLM-driven self-improvement of web code has a structural flaw: the model that proposes the repair is the model that judges it, and visual plausibility under that judge is a poor proxy for whether the page actually works. What the loop is missing is a counterparty the VLM cannot fool, and the browser…
Weitere Unterthemen aus Informationssysteme
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Software Engineering Research853 Papiere / 12 Monate+218 %
- Information Retrieval and Search Behavior727 Papiere / 12 Monate+506 %
- Recommender Systems and Techniques600 Papiere / 12 Monate+154 %
- Expert finding and Q&A systems205 Papiere / 12 Monate+1175 %
- Information and Cyber Security169 Papiere / 12 Monate+1500 %
- Blockchain Technology Applications and Security146 Papiere / 12 Monate+650 %
