Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Advanced Neural Network Applications
2 354 papiers indexés
Les recherches regroupées sous ce thème explorent des méthodes pour optimiser et adapter les réseaux de neurones à des tâches spécifiques. Elles abordent des techniques comme le Neural Architecture Search, qui automatise la conception d’architectures, ou le pruning, qui réduit la taille des modèles sans altérer leurs performances. D’autres travaux se concentrent sur des applications concrètes, telles que la détection d’éléments dans des plans architecturaux, la classification d’images ou les attaques ciblées sur des systèmes de segmentation sémantique, tout en intégrant des approches comme le self-supervised learning ou la quantification fine des modèles.
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel - 12 derniers mois
Pays des laboratoires
- Chine38 % · 616 articles
- États-Unis33 % · 539 articles
- Allemagne7,7 % · 126 articles
- Corée du Sud6,4 % · 105 articles
- Royaume-Uni5,5 % · 90 articles
- Inde4,5 % · 73 articles
- Canada4,3 % · 70 articles
- France4 % · 65 articles
Sur 1 631 articles de ce sujet dont au moins un laboratoire est situé. 85 pays représentés.
Il s'agit du pays du laboratoire, jamais de la nationalité des personnes. Un article signé depuis plusieurs pays compte pour chacun d'eux, les parts dépassent donc 100 % au total. La couverture est partielle et le manque n'est pas aléatoire : un chercheur dont l'institution est inconnue publie en général peu, ce qui sur-représente les laboratoires établis.
Derniers papiers
- LAS-CLIP: A Lightweight Adapter Steering Approach for CLIP's Visual Encoder
Anh-Khoa Dinh-Duc, Duc-Tai Dinh, Tam V. Nguyen, Minh-Triet Tran · 5 octobre 2026
CLIP's visual encoder produces only global image representations, limiting its use in region-level tasks. Existing adaptations rely on visual prompting, input masking, or encoder fine-tuning, each compromising pre-trained representations. We propose LAS-CLIP, a Lightweight Adapter Steering approach …
- EmbPASS: Towards Cross-Embodiment Open Panoramic Segmentation
Pujun Guo, Yuanfan Zheng, Fei Teng, Mengfei Duan, Guoqiang Zhao, Yuheng Zhang, Kai Luo, Kailun Yang · 5 octobre 2026
Panoramic images provide a complete 360-degree field of view, enabling comprehensive scene understanding for embodied perception. However, heterogeneous embodied platforms exhibit substantial differences in observation viewpoints and spatial layouts, giving rise to cross-embodiment observation shift…
- ViTok: Improving Dense Semantics in AM-RADIO-Style Multi-Teacher Distillation with PHI-S and Masked Image Modelling
Hailun Xu, Kanchan Sarkar · 5 octobre 2026
We study how to consolidate the current VITOK progress into a single multi-teacher distillation recipe that jointly preserves global recognition and dense semantics. Our starting point is an AM-RADIO-style student distilled from SigLIP2 and DINOv3-L, where SigLIP2 supplies strong global semantics an…
- FUSEye: Training-Light Fisheye Detection with Overlapping Views and Zero-Initialized Adapters
Wenya Su, Kai Luo, Di Wen, Ruiping Liu, Yufan Chen, Junwei Zheng, Kunyu Peng, Kailun Yang · 5 octobre 2026
Fisheye cameras give mobile robots a single-sensor, low-cost view of their surroundings, yet the COCO-pretrained detectors that practitioners routinely reuse fail on them: strong radial distortion warps local image structure, while boundary compression shrinks objects to near-invisible sizes. Full f…
- Physical AI Smart Spaces: A Large-Scale Benchmark for Multi-Camera 3D Perception in Smart Spaces
Yuxing Wang, Yizhou Wang, Anqi Li, Shuo Wang, Sameer Satish Pusegaonkar, Haoquan Liang, Jiajun Li, Shenxin Jiang, Jianhe Yuan, Shangru Li, Tongwei Dai, Zihao Chen, David C. Anastasiu, Sujit Biswas, Xunlei Wu, Zheng Tang · 5 octobre 2026
Physical AI Smart Spaces is, to the best of our knowledge, the first benchmark to simultaneously provide large-scale, multi-class, and multi-camera 3D perception data for indoor smart spaces. It contains over 280 hours of synchronized 1080p footage captured by nearly 1,800 cameras in warehouses, hos…
- Spatial Lifting for Dense Prediction
Mingzhi Xu, Tao Zhou, Yong Li, Yizhe Zhang · 2 octobre 2026
We present Spatial Lifting (SL), a novel methodology for dense prediction tasks. SL operates by lifting standard inputs, such as 2D images, into a higher-dimensional space and subsequently processing them using networks designed for that higher dimension, such as a 3D U-Net. Counterintuitively, this…
- MC-PanDA++: Simpler, Stronger, and More Robust Domain-Adaptive Panoptic Segmentation
Ivan Martinovi\'c, Josip \v{S}ari\'c, Yuki M. Asano, Sini\v{s}a \v{S}egvi\'c · 1 octobre 2026
Unsupervised domain adaptation (UDA) reduces the annotation burden in panoptic segmentation by leveraging a cost-effectively labeled source domain (e.g., synthetic) and an unlabeled target domain to bridge the distribution gap. Existing panoptic UDA methods rely on teacher-student consistency learni…
- DensePed-Lite: Quality-Aware Adaptive Detection for Dense Pedestrians under Occlusion
ZiAn Wang, MingZhe Liu, Chaoyi Guo, ChangChun Li, Fangming Gu · 1 octobre 2026
Pedestrian detection plays a crucial role in computer vision with applications in autonomous driving, surveillance, and public safety. However, real-world dense scenes bring severe challenges, including heavy occlusion, drastic scale variations, and strict real-time requirements. Existing lightweigh…
- COBICount: Separating Object and Background Responses for Remote Sensing Object Counting Without Training on Target Data
Junjing Zheng, Zhiyi Zhou, Ningrui Yang, Hongying Meng · 1 octobre 2026
Remote sensing object counting estimates how many buildings, vehicles, or ships appear in overhead images. Most supervised counters predict a density map, whose sum gives the object count, and assume similar categories, sizes, and backgrounds. Applying them across regions, sensors, or categories oft…
- On the Relaxation of Conditional Independence Assumption for Image Segmentation
Zixun Wang, Ben Dai · 1 octobre 2026
In semantic segmentation, a recent line of RankSEG methods directly optimizes Dice/IoU scores at inference time, improving alignment with evaluation metrics without modifying model training. Despite its theoretical and empirical success, RankSEG relies on the restrictive Conditional Independence Ass…
- SANTA++: Sampling Attention through Representative Keys
Kyle Lee, Christian Z. Pratt, Ruoyu Fang, Heekyung Lee, Avinash Lohitsa, Ryan Modafe, Kerem Y. Camsari · 30 septembre 2026
Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative keys for memory-efficient selection wit…
- When Can Attention Heads Be Statically Defined?
Weixian Waylon Li, Yintao Tai, Marcio Fonseca, Shay B. Cohen · 30 septembre 2026
Some attention heads learn similar patterns across inputs. Reusing these patterns could reduce training cost by avoiding repeated query-key score computation and softmax. Through controlled pretraining comparisons, we identify Selective Attention Freezing (SAF), which selects heads with low attentio…
- From Unity Simulation to Diffusion-Based Augmentation: Quantifying Dataset Balance for Robust Object Detection
Mohamed Benkedadra, Aissa Saoudi, Maxime Gloesener, Sidi Ahmed Mahmoudi, Matei Mancas · 30 septembre 2026
Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety monitoring, data collection is costly, hazardous, and ethically constrained. This paper presents a systematic study comparing two complementary data gen…
- Embedded Bi-Temporal Building Damage Assessment for On-Board Data Reduction
Thomas Goudemant, Benjamin Francesconi, Marjorie Bellizzi, Adrien Dorise · 30 septembre 2026
Rapid assessment of building damage after natural disasters is essential to support emergency response. Earth Observation satellites can acquire relevant imagery shortly after an event, but exploitation is limited by uplink and downlink capacity and by ground-processing latency. We address this with…
- Progressive-View On-Policy Distillation for Regional-to-Global Transfer in Multimodal LLMs
Shanfeng Huang, Zhou Fang, Song Xiao, Hai Du · 29 septembre 2026
Regional-to-global distillation uses crop-conditioned guidance to improve full-image understanding. The challenge is to effectively transfer the teacher's crop-based advantage to the student's full-image inference. We propose progressive-view on-policy distillation (PVD), which shifts the student's …
- CASS: Contribution-Aware Structured Sparsity for Model Merging
Yan Li, Guiping Cao, Meng Xu, Tao Jiang, Yaguang Song, Ming Tao, Yaowei Wang, Dongmei Jiang · 29 septembre 2026
Model merging integrates task-specific fine-tuned models into a single multi-task model, but often suffers from parameter interference caused by conflicting task-vector updates. Existing methods typically mitigate conflicts by pruning task vectors based on weight magnitude or random heuristics, trea…
- AdapToPASS: Ambiguity-aware Adaptive Spherical Transformer for Panoramic Semantic Segmentation
Soumyaratna Debnath, Weiming Zhang, Shriram Damodaran, Dingwen Xiao, Addison Lin Wang · 28 septembre 2026
Spherical Transformers have emerged as a promising framework for panoramic semantic segmentation (PASS) by operating directly on spherical geometry and alleviating projection-induced distortions. However, existing architectures often assume canonical spherical structure and stable viewpoints, which …
- CSCWD: Cross-Scale Channel-wise Knowledge Distillation for Lightweight Tiny Object Detection on Edge Devices
Amir Zamani, Zeinab Ghasemi-Naraghi · 28 septembre 2026
Real-time tiny object detection in aerial imagery is constrained by the weak spatial evidence of very small objects and the loss of high-resolution detail in lightweight detectors. This study presents Cross-Scale Channel-wise Knowledge Distillation (CSCWD), a training-time framework that transfers h…
- FoCal: Frequency-Oriented Cross-Modal Interaction and Spectral Calibration for Aerial Visible-Infrared Object Detection
Ben Liang, Chao Sui, Junqi Bai, Yuan Liu, Chunlai Li, Xiubao Sui, Qian Chen · 25 septembre 2026
In aerial RGB--IR object detection, effectively exploiting complementary information across modalities is critical for robust perception under complex illumination and environmental conditions. Existing multimodal detectors mainly focus on spatial-domain interaction or frequency-specific feature enh…
- RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation
Shaohua Dong, Zexuan Meng, Haiyan Sun, Bing Fan, Cuicui Zhang, Dylan Joseph, Kewei Sha, Yunhe Feng, Heng Fan · 25 septembre 2026
In this paper, we propose RGBD20K, a novel dataset for facilitating the development of more robust and general RGB-D semantic segmentation by encompassing abundant categories and high-quality annotations. RGBD20K possesses several attractive properties: (1) Expanded Semantic Space. In particular, it…
- OD3: Optimization-free Dataset Distillation for Object Detection
Salwa K. Al Khatib, Ahmed ElHagry, Shitong Shao, Zhiqiang Shen · 24 septembre 2026
Training large neural networks on large-scale datasets requires substantial computational resources, particularly for dense prediction tasks such as object detection. Although dataset distillation (DD) has been proposed to alleviate these demands by synthesizing compact datasets from larger ones, mo…
- SMDDFNet: State-space Modeling and Dynamic Dual Fusion Network for Traffic Sign Detection
TianYi Yu, DaJian Zhong, Lilin Wang · 24 septembre 2026
Traffic sign detection is a challenging visual signal processing task for advanced driver assistance, where small objects, scale variation, and occlusion limit conventional detectors with fixed receptive fields. This paper proposes State-space Modeling and Dynamic Dual Fusion Network (SMDDFNet), a d…
- LightMIS: Ultra-Lightweight Medical Image Segmentation Without a Stage-Wise Decoder
Andrei Arhire, Mihaela-Elena Breab\u{a}n, Radu Timofte · 24 septembre 2026
We present LightMIS, a scalable family of ultra-lightweight convolutional networks for 2D binary medical image segmentation without a learned stage-wise decoder. LightMIS aligns the outputs of a five-level encoder to a common resolution using Scale-Aligned Projection blocks, aggregates them once, an…
- PEARL: A Lightweight Prompt-based Feature Interpreter Framework for Real-Time, Anonymous, and Heterogeneous Collaborative Perception
Armin Maleki, Hayder Radha · 24 septembre 2026
Heterogeneity across Collaborative Perception (CP) agents is a major challenge for emerging CP frameworks due to domain gaps from differing sensors, architectures, and training data. Prior works mitigate this challenge by aligning features in a unified space via model retraining or per-agent-type in…
- GTR: Gated Token Recurrence for Efficient Dense Prediction
Zhe Feng, Longfei Liu, Wei Liu, Kai Chen, Jiangjiang Kong, Wei Zhou, Yifeng Qian, Dexiong Chen, Xuanlong Yu, Xi Shen · 23 septembre 2026
Self-attention-based vision backbones perform well on dense prediction, but the quadratic computational cost of global softmax attention limits their efficiency as image resolution increases. We introduce Gated Token Recurrence (GTR), a softmax-free recurrent vision backbone that combines gated line…
Autres sujets du thème Vision par ordinateur et reconnaissance de formes
Les sujets rattachés au même thème par la classification OpenAlex, les plus actifs d'abord.
- Multimodal Machine Learning Applications8 069 papiers / 12 mois+191 %
- Generative Adversarial Networks and Image Synthesis4 992 papiers / 12 mois+39 %
- Advanced Vision and Imaging841 papiers / 12 mois+78 %
- Human Pose and Action Recognition836 papiers / 12 mois+457 %
- Face recognition and analysis482 papiers / 12 mois+88 %
- Image Enhancement Techniques463 papiers / 12 mois+20 %
