Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Advanced Neural Network Applications
2.354 indexierte Paper
Die unter diesem Thema zusammengefassten Forschungen untersuchen Methoden zur Optimierung und Anpassung neuronaler Netze an spezifische Aufgaben. Sie behandeln Techniken wie Neural Architecture Search, das die Gestaltung von Architekturen automatisiert, oder Pruning, das die Größe von Modellen reduziert, ohne deren Leistung zu beeinträchtigen. Andere Arbeiten konzentrieren sich auf konkrete Anwendungen, wie die Detektion von Elementen in Architekturplänen, die Bildklassifikation oder gezielte Angriffe auf Systeme zur semantischen Segmentierung, wobei Ansätze wie Self-Supervised Learning oder die Feinquantisierung von Modellen integriert werden.
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- China38 % · 616 Artikel
- Vereinigte Staaten33 % · 539 Artikel
- Deutschland7,7 % · 126 Artikel
- Südkorea6,4 % · 105 Artikel
- Vereinigtes Königreich5,5 % · 90 Artikel
- Indien4,5 % · 73 Artikel
- Kanada4,3 % · 70 Artikel
- Frankreich4 % · 65 Artikel
Über 1.631 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 85 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- Spatial Lifting for Dense Prediction
Mingzhi Xu, Tao Zhou, Yong Li, Yizhe Zhang · 2. Oktober 2026
We present Spatial Lifting (SL), a novel methodology for dense prediction tasks. SL operates by lifting standard inputs, such as 2D images, into a higher-dimensional space and subsequently processing them using networks designed for that higher dimension, such as a 3D U-Net. Counterintuitively, this…
- MC-PanDA++: Simpler, Stronger, and More Robust Domain-Adaptive Panoptic Segmentation
Ivan Martinovi\'c, Josip \v{S}ari\'c, Yuki M. Asano, Sini\v{s}a \v{S}egvi\'c · 1. Oktober 2026
Unsupervised domain adaptation (UDA) reduces the annotation burden in panoptic segmentation by leveraging a cost-effectively labeled source domain (e.g., synthetic) and an unlabeled target domain to bridge the distribution gap. Existing panoptic UDA methods rely on teacher-student consistency learni…
- DensePed-Lite: Quality-Aware Adaptive Detection for Dense Pedestrians under Occlusion
ZiAn Wang, MingZhe Liu, Chaoyi Guo, ChangChun Li, Fangming Gu · 1. Oktober 2026
Pedestrian detection plays a crucial role in computer vision with applications in autonomous driving, surveillance, and public safety. However, real-world dense scenes bring severe challenges, including heavy occlusion, drastic scale variations, and strict real-time requirements. Existing lightweigh…
- COBICount: Separating Object and Background Responses for Remote Sensing Object Counting Without Training on Target Data
Junjing Zheng, Zhiyi Zhou, Ningrui Yang, Hongying Meng · 1. Oktober 2026
Remote sensing object counting estimates how many buildings, vehicles, or ships appear in overhead images. Most supervised counters predict a density map, whose sum gives the object count, and assume similar categories, sizes, and backgrounds. Applying them across regions, sensors, or categories oft…
- On the Relaxation of Conditional Independence Assumption for Image Segmentation
Zixun Wang, Ben Dai · 1. Oktober 2026
In semantic segmentation, a recent line of RankSEG methods directly optimizes Dice/IoU scores at inference time, improving alignment with evaluation metrics without modifying model training. Despite its theoretical and empirical success, RankSEG relies on the restrictive Conditional Independence Ass…
- SANTA++: Sampling Attention through Representative Keys
Kyle Lee, Christian Z. Pratt, Ruoyu Fang, Heekyung Lee, Avinash Lohitsa, Ryan Modafe, Kerem Y. Camsari · 30. September 2026
Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative keys for memory-efficient selection wit…
- When Can Attention Heads Be Statically Defined?
Weixian Waylon Li, Yintao Tai, Marcio Fonseca, Shay B. Cohen · 30. September 2026
Some attention heads learn similar patterns across inputs. Reusing these patterns could reduce training cost by avoiding repeated query-key score computation and softmax. Through controlled pretraining comparisons, we identify Selective Attention Freezing (SAF), which selects heads with low attentio…
- From Unity Simulation to Diffusion-Based Augmentation: Quantifying Dataset Balance for Robust Object Detection
Mohamed Benkedadra, Aissa Saoudi, Maxime Gloesener, Sidi Ahmed Mahmoudi, Matei Mancas · 30. September 2026
Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety monitoring, data collection is costly, hazardous, and ethically constrained. This paper presents a systematic study comparing two complementary data gen…
- Embedded Bi-Temporal Building Damage Assessment for On-Board Data Reduction
Thomas Goudemant, Benjamin Francesconi, Marjorie Bellizzi, Adrien Dorise · 30. September 2026
Rapid assessment of building damage after natural disasters is essential to support emergency response. Earth Observation satellites can acquire relevant imagery shortly after an event, but exploitation is limited by uplink and downlink capacity and by ground-processing latency. We address this with…
- Progressive-View On-Policy Distillation for Regional-to-Global Transfer in Multimodal LLMs
Shanfeng Huang, Zhou Fang, Song Xiao, Hai Du · 29. September 2026
Regional-to-global distillation uses crop-conditioned guidance to improve full-image understanding. The challenge is to effectively transfer the teacher's crop-based advantage to the student's full-image inference. We propose progressive-view on-policy distillation (PVD), which shifts the student's …
- CASS: Contribution-Aware Structured Sparsity for Model Merging
Yan Li, Guiping Cao, Meng Xu, Tao Jiang, Yaguang Song, Ming Tao, Yaowei Wang, Dongmei Jiang · 29. September 2026
Model merging integrates task-specific fine-tuned models into a single multi-task model, but often suffers from parameter interference caused by conflicting task-vector updates. Existing methods typically mitigate conflicts by pruning task vectors based on weight magnitude or random heuristics, trea…
- AdapToPASS: Ambiguity-aware Adaptive Spherical Transformer for Panoramic Semantic Segmentation
Soumyaratna Debnath, Weiming Zhang, Shriram Damodaran, Dingwen Xiao, Addison Lin Wang · 28. September 2026
Spherical Transformers have emerged as a promising framework for panoramic semantic segmentation (PASS) by operating directly on spherical geometry and alleviating projection-induced distortions. However, existing architectures often assume canonical spherical structure and stable viewpoints, which …
- CSCWD: Cross-Scale Channel-wise Knowledge Distillation for Lightweight Tiny Object Detection on Edge Devices
Amir Zamani, Zeinab Ghasemi-Naraghi · 28. September 2026
Real-time tiny object detection in aerial imagery is constrained by the weak spatial evidence of very small objects and the loss of high-resolution detail in lightweight detectors. This study presents Cross-Scale Channel-wise Knowledge Distillation (CSCWD), a training-time framework that transfers h…
- FoCal: Frequency-Oriented Cross-Modal Interaction and Spectral Calibration for Aerial Visible-Infrared Object Detection
Ben Liang, Chao Sui, Junqi Bai, Yuan Liu, Chunlai Li, Xiubao Sui, Qian Chen · 25. September 2026
In aerial RGB--IR object detection, effectively exploiting complementary information across modalities is critical for robust perception under complex illumination and environmental conditions. Existing multimodal detectors mainly focus on spatial-domain interaction or frequency-specific feature enh…
- RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation
Shaohua Dong, Zexuan Meng, Haiyan Sun, Bing Fan, Cuicui Zhang, Dylan Joseph, Kewei Sha, Yunhe Feng, Heng Fan · 25. September 2026
In this paper, we propose RGBD20K, a novel dataset for facilitating the development of more robust and general RGB-D semantic segmentation by encompassing abundant categories and high-quality annotations. RGBD20K possesses several attractive properties: (1) Expanded Semantic Space. In particular, it…
- OD3: Optimization-free Dataset Distillation for Object Detection
Salwa K. Al Khatib, Ahmed ElHagry, Shitong Shao, Zhiqiang Shen · 24. September 2026
Training large neural networks on large-scale datasets requires substantial computational resources, particularly for dense prediction tasks such as object detection. Although dataset distillation (DD) has been proposed to alleviate these demands by synthesizing compact datasets from larger ones, mo…
- SMDDFNet: State-space Modeling and Dynamic Dual Fusion Network for Traffic Sign Detection
TianYi Yu, DaJian Zhong, Lilin Wang · 24. September 2026
Traffic sign detection is a challenging visual signal processing task for advanced driver assistance, where small objects, scale variation, and occlusion limit conventional detectors with fixed receptive fields. This paper proposes State-space Modeling and Dynamic Dual Fusion Network (SMDDFNet), a d…
- LightMIS: Ultra-Lightweight Medical Image Segmentation Without a Stage-Wise Decoder
Andrei Arhire, Mihaela-Elena Breab\u{a}n, Radu Timofte · 24. September 2026
We present LightMIS, a scalable family of ultra-lightweight convolutional networks for 2D binary medical image segmentation without a learned stage-wise decoder. LightMIS aligns the outputs of a five-level encoder to a common resolution using Scale-Aligned Projection blocks, aggregates them once, an…
- PEARL: A Lightweight Prompt-based Feature Interpreter Framework for Real-Time, Anonymous, and Heterogeneous Collaborative Perception
Armin Maleki, Hayder Radha · 24. September 2026
Heterogeneity across Collaborative Perception (CP) agents is a major challenge for emerging CP frameworks due to domain gaps from differing sensors, architectures, and training data. Prior works mitigate this challenge by aligning features in a unified space via model retraining or per-agent-type in…
- GTR: Gated Token Recurrence for Efficient Dense Prediction
Zhe Feng, Longfei Liu, Wei Liu, Kai Chen, Jiangjiang Kong, Wei Zhou, Yifeng Qian, Dexiong Chen, Xuanlong Yu, Xi Shen · 23. September 2026
Self-attention-based vision backbones perform well on dense prediction, but the quadratic computational cost of global softmax attention limits their efficiency as image resolution increases. We introduce Gated Token Recurrence (GTR), a softmax-free recurrent vision backbone that combines gated line…
- DTKDP: A Dual Teacher Knowledge Distillation and Pruning Framework for Lightweight Oriented SAR Ship Detection
Yuming Li, Fan Zhang, Alin M. Achim · 22. September 2026
Two-stage oriented detectors achieve high localization accuracy in synthetic aperture radar (SAR) ship detection, but their large backbones, feature pyramids, proposal modules, and heavy region of interest (RoI) heads hinder deployment. Existing lightweight SAR ship detectors typically use one-stage…
- Ev-YOLO: Uncertainty-Aware Object Detection via a Unified Evidential Formulation
Simon Barbarit-Gaboriau (LITIS - STI, INSA Rouen Normandie), Hind Laghmara (LITIS - STI), R\'emi Boutteau (LITIS - STI), Samia Ainouz (LITIS, LITIS - STI) · 22. September 2026
Reliable uncertainty estimation is essential for deploying object detectors in autonomous systems operating in uncertain environments. Evidential Deep Learning (EDL) provides a principled framework for uncertainty-aware classification by representing network outputs as evidence and interpreting pred…
- LiAuto-MindViT: A Hybrid Vision Backbone with Adaptive Bidirectional Mamba
Lifu Mu, Shuai Chen, Wen Zheng, Haoyi Sun, Xueyang Fu, Pengfei Yu, Ning Mao, Tao Wei, Zhou Pan · 22. September 2026
While Mamba-based models have shown strong potential for long sequence modeling, adapting them to vision is challenging due to the requirement of local neighborhood correlations and multi-directional spatial contexts for visual understanding. In this paper, we present LiAuto-MindViT, a novel hybrid …
- SRPR-Net: Semantic and Relational Prompt Refinement for Automated SAM-based Instance Segmentation
Lufei Liu, Guojie Li, Suncheng Xiang, Fan Zhang · 22. September 2026
Instance segmentation is a fundamental computer vision task with diverse real-world applications. Recently, prompt-driven foundation models have shown promising generalization. However, automated prompting remains limited by insufficient semantic guidance and inter-instance modeling. To address this…
- Lightweight Pedestrian Head-Orientation Recognition Network for Safe Pedestrian-Vehicle Interaction
Yuanzhe Li, Yidi Huang, Xiaotong Chang, Hounian Liu · 22. September 2026
Pedestrian head orientation recognition plays an important role in autonomous driving by providing valuable cues for understanding pedestrian attention and anticipating potential crossing behavior. However, reliable recognition in real-world traffic scenes remains challenging because pedestrian head…
Weitere Unterthemen aus Bildverarbeitung und Mustererkennung
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Multimodal Machine Learning Applications8.069 Papiere / 12 Monate+191 %
- Generative Adversarial Networks and Image Synthesis4.992 Papiere / 12 Monate+39 %
- Advanced Vision and Imaging841 Papiere / 12 Monate+78 %
- Human Pose and Action Recognition836 Papiere / 12 Monate+457 %
- Face recognition and analysis482 Papiere / 12 Monate+88 %
- Image Enhancement Techniques463 Papiere / 12 Monate+20 %
