Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Visual Attention and Saliency Detection
265 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- China51 % · 96 Artikel
- Vereinigte Staaten23 % · 44 Artikel
- Deutschland7,4 % · 14 Artikel
- Vereinigtes Königreich5,9 % · 11 Artikel
- Indien5,3 % · 10 Artikel
- Australien4,8 % · 9 Artikel
- Sonderverwaltungsregion Hongkong4,8 % · 9 Artikel
- Südkorea4,8 % · 9 Artikel
Über 188 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 37 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- Consensus-Aware Multi-Source Fusion for Reference-Guided Camouflaged Object Detection
Junyang Xia, Luocheng Zhang, Wenwen Pan, Chifeng Zhu, Yang Yang, Xinchun Liu, Jiajun Ding · 1. Oktober 2026
Reference-guided camouflaged object detection aims to segment a target whose visual appearance closely resembles its surroundings by exploiting auxiliary reference samples. The task remains difficult because reference samples contain inconsistent target cues, while generic visual representations are…
- Beyond End-Task Success: How to Audit Visual Experience Retrieval in Robotics
Eshika Pathak, Leela Krishna · 28. September 2026
Robots that store past experiences must select which one to reuse in a new scene. Most systems select by visual similarity, and most evaluations report only the success of the selected experience. That number does not show whether the selection was good: a rule can score well by repeatedly using one…
- Can Frozen Hyperspherical Features Guide the Selection of Pseudo Masks?
Xinge Guo, Fengyang Xiao, Dingming Zhang, Yuhan Chen, Rihan Zhang, Xingjian Li, Tianyang Wang, Chunming He, Sina Farsiu · 25. September 2026
Foundation segmenters such as SAM return several plausible masks for an unlabeled image, and a student trained on the wrong one inherits its errors. Choosing among them means querying a second large model or fitting a quality head to annotated masks. We show that a candidate can be judged by what it…
- Overlapping Visual Grouping Without Semantic Priors
Teemu Saukkio, Hashem Haghbayan, Juha Plosila · 24. September 2026
Most computer-vision systems organize visual input toward a predefined interpretation, such as semantic categories, prompted regions, learned object-like representations, or a single spatial partition. This work considers an earlier stage of visual organization: the formation of candidate perceptual…
- S2A:Semantic-to-Spatial Alignment for Alignment-Free RGB-T Salient Object Detection
Qiangqiang Zhou, Yang Luo, Yong Chen, Jiawei Xu · 24. September 2026
Alignment-free RGB-T salient object detection (RGB-T SOD) aims to identify salient objects from unregistered RGB and thermal image pairs without costly pre-alignment. However, spatial misalignment breaks pixel-wise correspondence and causes feature contamination during cross-modal fusion. To address…
- S3VD: Semantic-Guidance Spatio-Temporal Scanning for Video Deraining
Kui Jiang, Yiang Chen, Yan Luo, Zhaocheng Yu, Junjun Jiang, Xianming Liu · 21. September 2026
Heavy rainfall severely degrades outdoor videos by corrupting high-frequency details and introducing motion blur, critically undermining the reliability of visual tasks. Recently, State Space Models (SSMs), particularly Mamba, have emerged as efficient alternatives for vision tasks with their linear…
- RA-SOD: Reliability-Aware RGB-T Salient Object Detection under Modality Degradation
Hongbo Gao, Zhengyu Li, Xueru Nie, Dihao Zhu, Lijun Zhao, Yunke Wang, Chang Xu · 14. September 2026
RGB-Thermal (RGB-T) salient object detection leverages complementary cues from visible and thermal modalities to improve robustness in challenging environments. However, in real-world scenarios, the reliability of each modality is inherently unstable: RGB images degrade under low illumination, motio…
- Efficient Semantic Understanding from Digital Foveation
Caterina Caccavella, Vittorio Fra, Andreas Ziegler, Giulia D'Angelo, Yulia Sandamirskaya · 4. September 2026
Dense semantic segmentation allocates computational resources uniformly across the entire image, regardless of scene complexity or task relevance. Inspired by biological vision, we investigate whether semantic understanding can be achieved more efficiently through digital foveated perception. We int…
- When Depth Hurts: Reliability-Aware Geometry Distillation for Depth-Free RGB-D Salient Object Detection
Xuehao Wang, Jiaxin Hua, Runmei Li, Zhenyu Wu, Chenglizhao Chen, Ke Gu, Aimin Hao · 4. September 2026
Depth can resolve appearance ambiguity in RGB-D salient object detection (SOD), yet sensor depth is not uniformly reliable. Missing regions, blurred boundaries, and structural artifacts can propagate through multimodal fusion and make an RGB-D detector less accurate than its RGB-only counterpart. Ex…
- Signal or Noise? Auditing Rotation-Induced Saliency Drift in Medical and Aerial Imaging
Khawaja Murad ul Hassan, Mehran Ebrahimi · 3. September 2026
Post-hoc saliency maps such as Grad-CAM are increasingly used to audit why a deployed vision model made a decision, yet the heatmap drifts when the input is rotated, even when the prediction is unchanged. In domains with no canonical orientation, such as histopathology and aerial imagery, this under…
- Dotting the Eye: An Intent-Driven Image Retouching Agent for Visual Focus Enhancement
Chujie Qin, Zilong Zhang, Zewei Chang, Chunle Guo, Ruixing Wang, Tao Hu, Ming-Ming Cheng, Chongyi Li · 2. September 2026
Image retouching is commonly formulated as enhancing overall visual quality through color adjustment, but in practice, it also serves to emphasize visual focus by guiding viewers' attention toward a specific subject or region. Achieving such focus-oriented retouching is inherently challenging, as it…
- TDFNet: Tri-projection Deformable Fusion Network for Panoramic Salient Object Detection
Qiangqiang Zhou, Jiacong Yu, Jiawei Xu, Yong Chen, Xin Huang, Ping Li · 27. August 2026
Recent years have witnessed the growing potential of panoramic salient object detection in robotic vision, virtual reality, and related applications. However, projecting spherical scenes onto 2D planes inevitably introduces geometric distortions, which fundamentally limit the effectiveness of existi…
- MIMONet: Multi-scale Input and Multi-scale Output Network for Salient Object Detection
Zhaojian Yao, Wei Gao, Tiesong Zhao, Hui Yuan, Sam Kwong · 27. August 2026
The existing methods for saliency detection task focus on the application of multi-level features, aiming to take advantage of the respective strengths of high- and low-level features. However, because the inputs of these models are single-size images, their multi-level features have difficulty in l…
- Saliency-Depth Conditioning for Zero-Shot Segmentation of Communication-Tower Components in Cluttered UAV Imagery
Ali Lesani, Chul Min Yeum, Su-Min Kang · 27. August 2026
Fine-grained segmentation of communication-tower components in UAV imagery is essential for automated inspection, yet task-specific models are hard to develop due to limited instance-level annotations. Zero-shot segmentation models offer a promising alternative, but in cluttered scenes, visually sim…
- SAM3Dual: A 3rd Place Solution to the MOSEv2 Track, 8th LSVOS Challenge
JeongRae Kim, Chaehyun Kim, Changwon Lim · 25. August 2026
We present SAM3Dual, our third-place solution to the MOSEv2 track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge at ECCV 2026. SAM3Dual is a training-free inference extension of pretrained SAM 3 that explicitly separates temporal memory into a short-term branch for recent observa…
- S$^3$AM: A Single-Stream SAM with Reliability-Calibrated Frequency Adapter for Multi-modal Salient Object Detection
Ruichao Hou, Boyue Xu, Tongwei Ren, Dongming Zhou, Gangshan Wu, Jinde Cao · 19. August 2026
Vision foundation models have recently advanced multi-modal salient object detection (MSOD) through parameter-efficient tuning and prompt learning. However, existing Segment Anything Model (SAM)-adapted MSOD methods often rely on dual-stream encoders or auxiliary prompt generators, leading to redund…
- Mask What Matters: Saliency-Guided Video Self-Supervised Learning for Autonomous Driving
Christopher Lang, Alexander Braun, Abhinav Valada · 19. August 2026
Video self-supervised learning through masked spatiotemporal prediction has emerged as a promising paradigm for learning feature representations from unlabeled data. However, existing methods typically rely on random masking, which indiscriminately removes regions irrespective of their semantic or t…
- Distractor-Aware Video Object Segmentation
Andreas Robinson, Abdelrahman Eldesokey, Michael Felsberg · 13. August 2026
Semi-supervised video object segmentation is a challenging task that aims to segment a target throughout a video sequence given an initial mask at the first frame. Discriminative approaches have demonstrated competitive performance on this task at a sensible complexity. These approaches typically fo…
- Hybrid Gated Attention
Zekun Zhou, Ruobing Xie, Lanrui Wang, Weixuan Sun · 13. August 2026
Gated attention is an effective approach to mitigate attention sinks and enhance the representational capacity of attention. To further extend its effectiveness-efficiency Pareto frontier, we propose a Hybrid Gated Attention (HyGA) framework that contains three types of gating strategies. Specifical…
- Is There Really a Camouflaged Object? Towards Realistic Camouflaged Object Detection
Huafeng Chen, Yueming Lyu, Chenyang Si, Wende Tan, Liucheng Guo, Caifeng Shan · 12. August 2026
Camouflaged object detection (COD) aims to segment objects that are visually concealed in their surroundings and has attracted increasing attention in recent years. However, most existing COD methods are developed under a closed-world assumption, where each input image is assumed to contain a camouf…
- LAD-COD: Language-Aligned Dense Perception for Camouflaged Object Detection
Shangye Song, Tianzhi Zhu, Syed Ariff Syed Hesham, Xin He, Yun Liu · 11. August 2026
Camouflaged object detection (COD) aims to segment objects that exhibit high visual similarity to their surroundings, which reduces foreground-background discriminability and weakens boundary evidence across appearance, texture, and structure. Such limitations motivate the use of instruction-conditi…
- Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics
Yuan Gao, Jin Song, Yiyun Fei, Gongzhe Li, Ruigao Yang · 11. August 2026
In recent years, Image Quality Assessment (IQA) for AI-generated images (AIGI) has advanced rapidly; however, existing methods primarily target portraits and artistic images, lacking a systematic evaluation of interior scenes. We introduce Spatial Aesthetics, a paradigm that assesses the aesthetic q…
- Global Attention-Fused Image Cropping with Attention-Guided and Global-Aligned Crop Evaluator
Haotian Yang, Zhile Yang, Kin-Man Lam, Patrick Le Callet, Xin Sun · 6. August 2026
Image cropping aims to improve image aesthetics by preserving important content within an appropriately composed region. However, most existing methods focus primarily on salient regions and therefore have limited sensitivity to the global relationships among the main image components. To address th…
- An Analysis and Implementation of Seam Carving for Content-Aware Image Resizing
Francesco Tosoni · 6. August 2026
Seam carving is a classical content-aware image resizing operator that modifies the width or height of an image by repeatedly removing (or inserting) seams, i.e., 8-connected monotonic paths of pixels of locally minimal importance. Because seams bend around salient content rather than uniformly scal…
- Is It Time for the Renaissance of Salient Object Detection in the Era of MLLMs?
Wenzhuo Zhao, Xiuzhi Li, Zhongkuan Mao, Ronghao Xian, Yao Jiang, Zhao Gao, Keren Fu, Qijun Zhao, Jian Cheng · 3. August 2026
The zero-shot capabilities of multimodal large language models (MLLMs) are pushing salient object detection (SOD) beyond task-specific supervision. To disentangle MLLMs beyond conventional mask-based evaluation, we decompose SOD into localization and segmentation, and re-engineer datasets with phras…
Weitere Unterthemen aus Bildverarbeitung und Mustererkennung
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Multimodal Machine Learning Applications8.069 Papiere / 12 Monate+191 %
- Generative Adversarial Networks and Image Synthesis4.992 Papiere / 12 Monate+39 %
- Advanced Neural Network Applications2.354 Papiere / 12 Monate+48 %
- Advanced Vision and Imaging841 Papiere / 12 Monate+78 %
- Human Pose and Action Recognition836 Papiere / 12 Monate+457 %
- Face recognition and analysis482 Papiere / 12 Monate+88 %
