Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Advanced Neural Network Applications
2354 artículos indexados
Las investigaciones agrupadas bajo este tema exploran métodos para optimizar y adaptar las redes neuronales a tareas específicas. Abordan técnicas como el Neural Architecture Search, que automatiza el diseño de arquitecturas, o el pruning, que reduce el tamaño de los modelos sin alterar su rendimiento. Otros trabajos se centran en aplicaciones concretas, como la detección de elementos en planos arquitectónicos, la clasificación de imágenes o los ataques dirigidos a sistemas de segmentación semántica, al tiempo que integran enfoques como el self-supervised learning o la cuantización fina de los modelos.
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- China38 % · 616 artículos
- Estados Unidos33 % · 539 artículos
- Alemania7,7 % · 126 artículos
- Corea del Sur6,4 % · 105 artículos
- Reino Unido5,5 % · 90 artículos
- India4,5 % · 73 artículos
- Canadá4,3 % · 70 artículos
- Francia4 % · 65 artículos
Sobre 1631 artículos de este tema con al menos un laboratorio localizado. 85 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- LAS-CLIP: A Lightweight Adapter Steering Approach for CLIP's Visual Encoder
Anh-Khoa Dinh-Duc, Duc-Tai Dinh, Tam V. Nguyen, Minh-Triet Tran · 5 de octubre de 2026
CLIP's visual encoder produces only global image representations, limiting its use in region-level tasks. Existing adaptations rely on visual prompting, input masking, or encoder fine-tuning, each compromising pre-trained representations. We propose LAS-CLIP, a Lightweight Adapter Steering approach …
- EmbPASS: Towards Cross-Embodiment Open Panoramic Segmentation
Pujun Guo, Yuanfan Zheng, Fei Teng, Mengfei Duan, Guoqiang Zhao, Yuheng Zhang, Kai Luo, Kailun Yang · 5 de octubre de 2026
Panoramic images provide a complete 360-degree field of view, enabling comprehensive scene understanding for embodied perception. However, heterogeneous embodied platforms exhibit substantial differences in observation viewpoints and spatial layouts, giving rise to cross-embodiment observation shift…
- ViTok: Improving Dense Semantics in AM-RADIO-Style Multi-Teacher Distillation with PHI-S and Masked Image Modelling
Hailun Xu, Kanchan Sarkar · 5 de octubre de 2026
We study how to consolidate the current VITOK progress into a single multi-teacher distillation recipe that jointly preserves global recognition and dense semantics. Our starting point is an AM-RADIO-style student distilled from SigLIP2 and DINOv3-L, where SigLIP2 supplies strong global semantics an…
- FUSEye: Training-Light Fisheye Detection with Overlapping Views and Zero-Initialized Adapters
Wenya Su, Kai Luo, Di Wen, Ruiping Liu, Yufan Chen, Junwei Zheng, Kunyu Peng, Kailun Yang · 5 de octubre de 2026
Fisheye cameras give mobile robots a single-sensor, low-cost view of their surroundings, yet the COCO-pretrained detectors that practitioners routinely reuse fail on them: strong radial distortion warps local image structure, while boundary compression shrinks objects to near-invisible sizes. Full f…
- Physical AI Smart Spaces: A Large-Scale Benchmark for Multi-Camera 3D Perception in Smart Spaces
Yuxing Wang, Yizhou Wang, Anqi Li, Shuo Wang, Sameer Satish Pusegaonkar, Haoquan Liang, Jiajun Li, Shenxin Jiang, Jianhe Yuan, Shangru Li, Tongwei Dai, Zihao Chen, David C. Anastasiu, Sujit Biswas, Xunlei Wu, Zheng Tang · 5 de octubre de 2026
Physical AI Smart Spaces is, to the best of our knowledge, the first benchmark to simultaneously provide large-scale, multi-class, and multi-camera 3D perception data for indoor smart spaces. It contains over 280 hours of synchronized 1080p footage captured by nearly 1,800 cameras in warehouses, hos…
- Spatial Lifting for Dense Prediction
Mingzhi Xu, Tao Zhou, Yong Li, Yizhe Zhang · 2 de octubre de 2026
We present Spatial Lifting (SL), a novel methodology for dense prediction tasks. SL operates by lifting standard inputs, such as 2D images, into a higher-dimensional space and subsequently processing them using networks designed for that higher dimension, such as a 3D U-Net. Counterintuitively, this…
- MC-PanDA++: Simpler, Stronger, and More Robust Domain-Adaptive Panoptic Segmentation
Ivan Martinovi\'c, Josip \v{S}ari\'c, Yuki M. Asano, Sini\v{s}a \v{S}egvi\'c · 1 de octubre de 2026
Unsupervised domain adaptation (UDA) reduces the annotation burden in panoptic segmentation by leveraging a cost-effectively labeled source domain (e.g., synthetic) and an unlabeled target domain to bridge the distribution gap. Existing panoptic UDA methods rely on teacher-student consistency learni…
- DensePed-Lite: Quality-Aware Adaptive Detection for Dense Pedestrians under Occlusion
ZiAn Wang, MingZhe Liu, Chaoyi Guo, ChangChun Li, Fangming Gu · 1 de octubre de 2026
Pedestrian detection plays a crucial role in computer vision with applications in autonomous driving, surveillance, and public safety. However, real-world dense scenes bring severe challenges, including heavy occlusion, drastic scale variations, and strict real-time requirements. Existing lightweigh…
- COBICount: Separating Object and Background Responses for Remote Sensing Object Counting Without Training on Target Data
Junjing Zheng, Zhiyi Zhou, Ningrui Yang, Hongying Meng · 1 de octubre de 2026
Remote sensing object counting estimates how many buildings, vehicles, or ships appear in overhead images. Most supervised counters predict a density map, whose sum gives the object count, and assume similar categories, sizes, and backgrounds. Applying them across regions, sensors, or categories oft…
- On the Relaxation of Conditional Independence Assumption for Image Segmentation
Zixun Wang, Ben Dai · 1 de octubre de 2026
In semantic segmentation, a recent line of RankSEG methods directly optimizes Dice/IoU scores at inference time, improving alignment with evaluation metrics without modifying model training. Despite its theoretical and empirical success, RankSEG relies on the restrictive Conditional Independence Ass…
- SANTA++: Sampling Attention through Representative Keys
Kyle Lee, Christian Z. Pratt, Ruoyu Fang, Heekyung Lee, Avinash Lohitsa, Ryan Modafe, Kerem Y. Camsari · 30 de septiembre de 2026
Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative keys for memory-efficient selection wit…
- When Can Attention Heads Be Statically Defined?
Weixian Waylon Li, Yintao Tai, Marcio Fonseca, Shay B. Cohen · 30 de septiembre de 2026
Some attention heads learn similar patterns across inputs. Reusing these patterns could reduce training cost by avoiding repeated query-key score computation and softmax. Through controlled pretraining comparisons, we identify Selective Attention Freezing (SAF), which selects heads with low attentio…
- From Unity Simulation to Diffusion-Based Augmentation: Quantifying Dataset Balance for Robust Object Detection
Mohamed Benkedadra, Aissa Saoudi, Maxime Gloesener, Sidi Ahmed Mahmoudi, Matei Mancas · 30 de septiembre de 2026
Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety monitoring, data collection is costly, hazardous, and ethically constrained. This paper presents a systematic study comparing two complementary data gen…
- Embedded Bi-Temporal Building Damage Assessment for On-Board Data Reduction
Thomas Goudemant, Benjamin Francesconi, Marjorie Bellizzi, Adrien Dorise · 30 de septiembre de 2026
Rapid assessment of building damage after natural disasters is essential to support emergency response. Earth Observation satellites can acquire relevant imagery shortly after an event, but exploitation is limited by uplink and downlink capacity and by ground-processing latency. We address this with…
- Progressive-View On-Policy Distillation for Regional-to-Global Transfer in Multimodal LLMs
Shanfeng Huang, Zhou Fang, Song Xiao, Hai Du · 29 de septiembre de 2026
Regional-to-global distillation uses crop-conditioned guidance to improve full-image understanding. The challenge is to effectively transfer the teacher's crop-based advantage to the student's full-image inference. We propose progressive-view on-policy distillation (PVD), which shifts the student's …
- CASS: Contribution-Aware Structured Sparsity for Model Merging
Yan Li, Guiping Cao, Meng Xu, Tao Jiang, Yaguang Song, Ming Tao, Yaowei Wang, Dongmei Jiang · 29 de septiembre de 2026
Model merging integrates task-specific fine-tuned models into a single multi-task model, but often suffers from parameter interference caused by conflicting task-vector updates. Existing methods typically mitigate conflicts by pruning task vectors based on weight magnitude or random heuristics, trea…
- AdapToPASS: Ambiguity-aware Adaptive Spherical Transformer for Panoramic Semantic Segmentation
Soumyaratna Debnath, Weiming Zhang, Shriram Damodaran, Dingwen Xiao, Addison Lin Wang · 28 de septiembre de 2026
Spherical Transformers have emerged as a promising framework for panoramic semantic segmentation (PASS) by operating directly on spherical geometry and alleviating projection-induced distortions. However, existing architectures often assume canonical spherical structure and stable viewpoints, which …
- CSCWD: Cross-Scale Channel-wise Knowledge Distillation for Lightweight Tiny Object Detection on Edge Devices
Amir Zamani, Zeinab Ghasemi-Naraghi · 28 de septiembre de 2026
Real-time tiny object detection in aerial imagery is constrained by the weak spatial evidence of very small objects and the loss of high-resolution detail in lightweight detectors. This study presents Cross-Scale Channel-wise Knowledge Distillation (CSCWD), a training-time framework that transfers h…
- FoCal: Frequency-Oriented Cross-Modal Interaction and Spectral Calibration for Aerial Visible-Infrared Object Detection
Ben Liang, Chao Sui, Junqi Bai, Yuan Liu, Chunlai Li, Xiubao Sui, Qian Chen · 25 de septiembre de 2026
In aerial RGB--IR object detection, effectively exploiting complementary information across modalities is critical for robust perception under complex illumination and environmental conditions. Existing multimodal detectors mainly focus on spatial-domain interaction or frequency-specific feature enh…
- RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation
Shaohua Dong, Zexuan Meng, Haiyan Sun, Bing Fan, Cuicui Zhang, Dylan Joseph, Kewei Sha, Yunhe Feng, Heng Fan · 25 de septiembre de 2026
In this paper, we propose RGBD20K, a novel dataset for facilitating the development of more robust and general RGB-D semantic segmentation by encompassing abundant categories and high-quality annotations. RGBD20K possesses several attractive properties: (1) Expanded Semantic Space. In particular, it…
- OD3: Optimization-free Dataset Distillation for Object Detection
Salwa K. Al Khatib, Ahmed ElHagry, Shitong Shao, Zhiqiang Shen · 24 de septiembre de 2026
Training large neural networks on large-scale datasets requires substantial computational resources, particularly for dense prediction tasks such as object detection. Although dataset distillation (DD) has been proposed to alleviate these demands by synthesizing compact datasets from larger ones, mo…
- SMDDFNet: State-space Modeling and Dynamic Dual Fusion Network for Traffic Sign Detection
TianYi Yu, DaJian Zhong, Lilin Wang · 24 de septiembre de 2026
Traffic sign detection is a challenging visual signal processing task for advanced driver assistance, where small objects, scale variation, and occlusion limit conventional detectors with fixed receptive fields. This paper proposes State-space Modeling and Dynamic Dual Fusion Network (SMDDFNet), a d…
- LightMIS: Ultra-Lightweight Medical Image Segmentation Without a Stage-Wise Decoder
Andrei Arhire, Mihaela-Elena Breab\u{a}n, Radu Timofte · 24 de septiembre de 2026
We present LightMIS, a scalable family of ultra-lightweight convolutional networks for 2D binary medical image segmentation without a learned stage-wise decoder. LightMIS aligns the outputs of a five-level encoder to a common resolution using Scale-Aligned Projection blocks, aggregates them once, an…
- PEARL: A Lightweight Prompt-based Feature Interpreter Framework for Real-Time, Anonymous, and Heterogeneous Collaborative Perception
Armin Maleki, Hayder Radha · 24 de septiembre de 2026
Heterogeneity across Collaborative Perception (CP) agents is a major challenge for emerging CP frameworks due to domain gaps from differing sensors, architectures, and training data. Prior works mitigate this challenge by aligning features in a unified space via model retraining or per-agent-type in…
- GTR: Gated Token Recurrence for Efficient Dense Prediction
Zhe Feng, Longfei Liu, Wei Liu, Kai Chen, Jiangjiang Kong, Wei Zhou, Yifeng Qian, Dexiong Chen, Xuanlong Yu, Xi Shen · 23 de septiembre de 2026
Self-attention-based vision backbones perform well on dense prediction, but the quadratic computational cost of global softmax attention limits their efficiency as image resolution increases. We introduce Gated Token Recurrence (GTR), a softmax-free recurrent vision backbone that combines gated line…
Otros asuntos del tema Visión por computador y reconocimiento de formas
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Multimodal Machine Learning Applications8069 artículos / 12 meses+191 %
- Generative Adversarial Networks and Image Synthesis4992 artículos / 12 meses+39 %
- Advanced Vision and Imaging841 artículos / 12 meses+78 %
- Human Pose and Action Recognition836 artículos / 12 meses+457 %
- Face recognition and analysis482 artículos / 12 meses+88 %
- Image Enhancement Techniques463 artículos / 12 meses+20 %
