Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Video Analysis and Summarization
234 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- China51 % · 31 Artikel
- Vereinigte Staaten36 % · 22 Artikel
- Sonderverwaltungsregion Hongkong9,8 % · 6 Artikel
- Südkorea9,8 % · 6 Artikel
- Singapur8,2 % · 5 Artikel
- Indien6,6 % · 4 Artikel
- Japan4,9 % · 3 Artikel
- Taiwan4,9 % · 3 Artikel
Über 61 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 26 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- Paying for Too Many Tokens? Valid and Cost-Efficient Multimodal LLM Annotation with Simple Heuristics
Zhixi Zhu, Kristina Gligoric · 2. Oktober 2026
Vision-Language Models (VLMs) enable video annotation at scale, but costs accumulate quickly: processing a typical 60-second short-form video at one frame per second requires millions of tokens. To reduce costs, researchers rely on heuristics such as sampling a subset of frames, compressing videos i…
- DramaAgent: Agentic Storytelling Video Generation
Ting Huang, Biao Wu, Ronghao Chen, Zeyu Zhang, Tengfei Cheng, Qizhen Lan, Huacan Wang, Hao Tang · 2. Oktober 2026
Recent diffusion and autoregressive models have substantially improved text-to-video generation, yet producing coherent long-form story videos with consistent characters and aligned audio remains challenging. Existing methods often suffer from narrative drift, unstable character identity, weak cross…
- Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization
Guangyu Yang, Jingbiao Mei, Mingsheng Sun, Jinghong Chen, Yingtong Bu, Pengda Qin, Da Chen, Bill Byrne · 2. Oktober 2026
The rapid growth of video-based social media has increased users' exposure to harmful content, creating a need for reliable automated video safety detection. Although recent Vision-Language Models (VLMs) show strong video understanding capabilities, existing harmful video detection systems face two …
- VETO: Video Efficient Token Optimization for Vision Language Models
Gueter Josmy Faure, Hao Ping Wang, Min-Hung Chen, Winston H. Hsu · 2. Oktober 2026
Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundan…
- Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs
Youngwoo Shin, Yusung Ro, Minseo Kim, Junmo Kim · 2. Oktober 2026
Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal a…
- Video Generation Models: A Survey of Post-Training and Alignment
Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni, Eun Woo Im, Mohammadmahdi Honarmand, Zeyu Wang, Juntong Song, Fei Du, Xilin Jiang, Kexin Zheng, Tianzhi Li, Fei Tao, Pooyan Fazli · 2. Oktober 2026
Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pretrained video models often fail to reliably follow human intent, mainta…
- VideoEvolve: Evolving Agent Harnesses for Video Temporal Grounding
Bingjun Luo, Yuhuan Fan, Jialin Guo, Siqi Li · 2. Oktober 2026
Video temporal grounding aims to localize events in videos from natural-language queries. For agents built around frozen video-language models, the harness determines how queries guide temporal predictions and how those predictions are refined. Manually refining these harnesses requires diagnosing g…
- FOMO: Forget the Concept, Don't Miss Out on the Scene in Selective Video Unlearning
{\L}ukasz Rudnik, Agnieszka Polowczyk, Alicja Polowczyk, Przemys{\l}aw Spurek · 1. Oktober 2026
The rapid advancement of generative video models has enabled the synthesis of increasingly realistic and temporally coherent videos, while also raising concerns about the generation of harmful content. The reliance on large-scale web datasets during training inevitably exposes these models to undesi…
- RESUME: Recurrent State Updates from Motion and Residual Signals for Efficient Video Language Modeling
Can Zhang, Xiaotian Han, Junyuan Shang, Yuchen Ding, Zhenyu Zhang, Shuohuan Wang, Dianhai Yu, Ruirui Li · 1. Oktober 2026
Existing video language models encode sampled RGB frames independently, so a long video must either exhaust the token budget or drop the changes between sampled frames. Codec-aware front-ends read the motion vectors and residuals that encoding produced, but in their deployed form each predictive fra…
- BMASH: Ball-Motion-Aware Soccer Header Spotting
Ahmed Endris Hasen, Muhammad Shahzad Khan, Nikolaos Passalis, Jenni Raitoharju · 1. Oktober 2026
Recent advances in computer vision have made broadcast sports videos increasingly useful for event analysis, performance assessment, and player-safety applications. In soccer, however, header spotting remains a challenging problem due to the subtle and short-lived nature of header events. This paper…
- TSMD: Temporal-Stream Modality Dropout for Robust Video Highlight Detection
Bo-Yuan Cheng, Kuan-Yu Chen, Po-Han Huang, Jeng-Lin Li, Jian-Jiun Ding · 1. Oktober 2026
Existing multimodal video highlight detectors typically assume that visual, audio, and textual streams are continuously available. In practice, however, inputs may suffer from localized frame missingness or complete-stream outage. We formulate this robustness challenge along two dimensions: temporal…
- MEMO: Multi-Level Entity-Aware Memory for Streaming Video Understanding
Yinying Li, Yuqian Fu, Yulin Dai, Jingyu Gong, Tianwen Qian, Xiaoling Wang · 1. Oktober 2026
Streaming video understanding requires models to process unbounded visual streams while preserving rich visual semantics across vast temporal horizons, posing a fundamental challenge for memory modeling. Existing approaches primarily focus on increasing memory capacity, either by compressing histori…
- FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation
Bo Yin, Xiaobin Hu, Jiaqi Zhao, Shuicheng Yan · 1. Oktober 2026
Long-horizon video generation requires models to effectively leverage an increasingly long generation history. As the generated history grows, retaining all previous content becomes increasingly expensive and redundant, making effective historical selection essential. Existing approaches often deter…
- Event-Driven Refresh and Recurrence Memory to Reduce Stale Grounding in Referring Video Object Segmentation
Abu Hanif Muhammad Syarubany, Jaehyun Jang, Siwoo Lim, Seungyeon Ryu, Chang D. Yoo · 1. Oktober 2026
Referring Video Object Segmentation (RVOS) aims to produce a pixel-accurate mask sequence for an object specified by natural language. Sa2VA combines a multimodal large language model with SAM2 for grounded segmentation; however, its inference typically grounds the query from a small fixed set of in…
- No Corners Cut: State-Grounded Transitions for Mid-Stream Prompt Switches in Video Generation
Zejing Rao, Ketong Ren, Xiaoqiang Liu, Yiping Meng, Guoxin Zhang, Fan Tang · 1. Oktober 2026
Streaming video generators allow users to dynamically modulate video synthesis via mid-stream prompt switching. Existing streaming methods can respond to the updated instruction while still cutting corners, prematurely realizing goals or taking heuristic shortcuts that bypass necessary intermediate …
- LongTake: Learning to Sustain Dynamics in Long-Horizon Video Generation
Byoungwoo Park, Jaemoo Choi, Juho Lee, Yongxin Chen · 1. Oktober 2026
World models, game simulators, and long-take video creation require coherent scene evolution and sustained dynamics over extended durations. Autoregressive (AR) video diffusion provides a natural framework for long-horizon generation, yet extended rollouts often become near-static or lose visual qua…
- MOBA-VL: Event-Localized Multi-Turn Reinforcement Learning for Real-Time MOBA Commentary
Shengyun Zhong, Xinkang Zhao, Ziyuan Chu, Linchao Zhu · 1. Oktober 2026
Real-time commentary for Multiplayer Online Battle Arena (MOBA) esports requires a vision-language model (VLM) to narrate a live match second by second, both fluently and accurately. Existing streaming VLMs sound natural but often miss key events such as kills and objectives. To address this limitat…
- VidHarness: Evolving Agent Harnesses for Cost-Efficient Long Video Understanding
Susan Liang, Jianmin Wu, Daxiang Dong · 1. Oktober 2026
Vision-language models (VLMs) can answer questions about hour-long videos, but processing every frame is prohibitively expensive, even though the evidence for a question usually spans only a few seconds. Video agents, i.e., harness programs wrapped around a frozen VLM, address this by observing the …
- PhyProbe: Rethinking Physical Consistency Evaluation in Generated Videos
Max Ku, Jiaojiao Fan, Zekun Hao, Francesco Ferroni, Heng Wang, Wenhu Chen, Ming-Yu Liu, Prithvijit Chattopadhyay · 1. Oktober 2026
Evaluating the physical consistency of generated videos remains a fundamental challenge. Existing approaches rely on off-the-shelf vision-language models, which can often be myopic to physical dynamics, or fine-tuned evaluators trained on human annotations, which overfit to dataset-specific cues and…
- POET: Preference Optimization for Enhanced Text-to-Image Generation
Ruibo Chen, Jiacheng Pan, Heng Huang, Zhenheng Yang · 1. Oktober 2026
Recent advances in text-to-image (T2I) generation have achieved impressive results, yet existing models often struggle with simple or underspecified user prompts due to a distributional gap with their descriptive training captions. This frequently leads to suboptimal image-text alignment, aesthetics…
- ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing
Xinghao Chen, Xiangbo Gao, Jiongze Yu, Yuheng Wu, Zhengzhong Tu · 1. Oktober 2026
Recent video generation is increasingly realistic and controllable, yet video editing remains less developed, particularly for precise local edits that must preserve the original scene dynamics. Video scene text editing replaces text on scene surfaces, such as storefront signs, whiteboards, and prod…
- MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories
Guangzhi Xiong, Xinyuan Zhang, Xiao Yang, Hyokun Yun, Kai Zhang, Shiun-Zu Kuo, Hyeonjeong Ha, Xilun Chen, Kai Sun, Lucas Liang, Guangqiang Dong, Ejaz Ahmed, Ahmed A Aly, Anuj Kumar, Raffay Hamid, Aidong Zhang, Xin Luna Dong · 1. Oktober 2026
Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compact…
- Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
Hui Ren, Lei Fan, Henry Pao, Han Guo, Zeeshan Zia, Ying Chen, Alexander Schwing, Gang Hua · 30. September 2026
Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain…
- Video-RSI: Recursive Self-Improvement of Video Understanding Agents via Harness Evolution
Bingjun Luo, Jialin Guo, Siqi Li · 30. September 2026
Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations. However, execution traces contain only the evidence acquired by the current harness, leaving competing explanations for failure unresolved and limiting the b…
- Watch-Think-Interact: Bootstrapping Long-Horizon Multi-Turn Streaming Video Reasoning with Reinforcement Learning
Ziheng Huang, Yicheng Bao, Xueheng Li, Zhenkun Gao, Bangwei Liu, Kunquan Li, Yuxiang Shen, Bangyan Li, Xuejiao Wang, Changbo Wang, Gaoqi He · 30. September 2026
Streaming video assistance requires models to answer asynchronous questions from an observed prefix under a fixed context budget. Existing approaches model response timing or compress history, but an online state formed before future questions are known can omit visual details before later questions…
Weitere Unterthemen aus Bildverarbeitung und Mustererkennung
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Multimodal Machine Learning Applications8.069 Papiere / 12 Monate+191 %
- Generative Adversarial Networks and Image Synthesis4.992 Papiere / 12 Monate+39 %
- Advanced Neural Network Applications2.354 Papiere / 12 Monate+48 %
- Advanced Vision and Imaging841 Papiere / 12 Monate+78 %
- Human Pose and Action Recognition836 Papiere / 12 Monate+457 %
- Face recognition and analysis482 Papiere / 12 Monate+88 %
