Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Image Processing and 3D Reconstruction
45 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel - 12 derniers mois
Derniers papiers
- Tracing the Evolution of Oracle Bone Characters Across Three Millennia
Tianhao Fu, Xinxin Xu, Spike Wang, Cunyi Kang, Jian Cao, Xixin Cao · 30 septembre 2026
Of the approximately 4,500 Oracle Bone Inscription (OBI) characters discovered from the Shang dynasty, only about 1,600 have been deciphered. Many computational approaches compare OBI with glyphs from one historical period at a time. However, during the evolution of Chinese characters, significant s…
- Think Before You Restore: Risk-Aware Manchu Manuscript Restoration with Stroke-Guided Attention
Mingqiu Liang, Dongdong Wang, Siyang Lu, Ting Huang, Yingjun Qi · 30 septembre 2026
Full-page blind restoration of historical Manchu manuscripts is challenging due to scarce annotations, unknown degradation regions, and fragile connected strokes. Generic restoration models may improve visual quality but often modify intact content, leading to over-restoration. We propose SAGE-Resto…
- Quantifying the Occult: A Comparative Study of Hindu and Buddhist Deities Using Machine Learning Methods
Ankit Bhattacharjee · 24 septembre 2026
This study introduces a dual-matrix computational architecture to mathematically quantify the morphological and theological divergence of 196 Hindu and Vajrayana Buddhist esoteric deities. Physical morphology is evaluated via a discrete Gower distance matrix enhanced by a novel "Cardinality Weightin…
- Canonical locks that encode part-whole hierarchies
Rajat Modi, Yogesh Singh Rawat · 23 septembre 2026
One of the challenges in representational learning is how to encode part-whole hierarchies in a neural net. Prior works rely on flattening tree-like structures into string-like sequences and training a sequence-to-sequence model via autoregression. While such a representation works for parse-trees i…
- A Lightweight Convolutional Neural Network for Real-Time Recognition of Hand-Drawn Geometric Shapes
Shahir Abdullah · 22 septembre 2026
Recognizing hand-drawn geometric shapes is a foundational sub-problem of sketch recognition, with applications in education, human-computer interaction, and diagram digitization. This paper presents the design, implementation, and evaluation of a desktop application that recognizes four basic hand-d…
- Sharp Reconstruction Bounds for Autoencoders Using the Same Forward Map
Patricia Medina, Hy P. G. Lam · 18 septembre 2026
We study reconstruction in autoencoders that apply the same forward map before and after setting the observed coordinates to zero. For equal odd input and hidden dimensions $d\geq 3$, among orientation-preserving diffeomorphisms whose Jacobian singular values lie in $[m,M]$, we show that the least u…
- Beyond Landmark Extraction: A Framework for Robust Geometric Feature Construction in Structured Image Classification
Saravana Mauree, Sakshi Arya · 2 septembre 2026
Much of the literature on structured image recognition has disproportionately focused on the comparison of classification algorithms. Rather than investigating which classifier performs best, this paper instead asks: what should a classifier know before it ever makes a prediction? In structured visi…
- X-SG$^2$S: Safe and Generalizable Gaussian Splatting with X-dimensional Watermarks
Zihang Cheng, Wentao Bao, Huiping Zhuang, Chun Li, Xin Meng, Ziqian Zeng, Cen Chen, Ming Li, F. Richard Yu · 2 septembre 2026
3D Gaussian Splatting (3DGS) has been widely used in 3D reconstruction and 3D generation. However, the rapid adoption of 3D Gaussian Splatting raises growing concerns about information leakage and unauthorized use, urging the exploration of effective watermarking techniques. However, existing method…
- When Should a Network Emit Geometry, and When Should It Detect It? Readout, Reconciliation, and Representation in Floorplan Vectorization
He Zhang · 27 août 2026
A network trained to recover the walls, openings, and rooms of a rasterized floorplan can produce its output in two ways: by emitting the geometry as an autoregressive coordinate sequence, or by detecting it on dense junction and centerline heatmaps and assembling a graph. We compare the two readout…
- Decoding the Past: An Uncertainty-Aware Deep Learning Framework for Sex Attribution in Prehistoric Hand Stencils
Karel Becerra, Boris Mederos, Dean Snow, Ram\'on A. Mollineda · 17 août 2026
Determining the biological sex of the individuals who created Upper Paleolithic hand stencils remains a challenging problem due to the absence of ground truth, population differences between contemporary and prehistoric groups, and the uncertainty introduced by image degradation. Traditional morphom…
- AmalthAI: An Open-Source Computer Vision Platform for Cultural Heritage
Christos Chatzisavvas, Stelios Alvanos, Efstratios Politis, Panagiotis Rigas, Thomas Pappas, Ioannis Giannoukos, Nikolaos Mitianoudis, Agata Ulanowska, Katarzyna Żebrowska, Nazarij Buławka, Christina Margariti, George Pavlidis, Chairi Kiourt, Anestis Koutsoudis, Vassilis Katsouros, George Ioannakis · 14 août 2026
Computer vision (CV) and machine learning (ML) offer new tools for cultural heritage (CH) artifact analysis, but the CV/ML pipeline remains largely inaccessible to CH domain experts, who lack the background to configure, train, or assess models. We present AmalthAI, an open-source CV platform that b…
- JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
Shawn Li, Wei Yang, Jike Zhong, Jiate Li, Jiawei Yang, You Qin, Ryan Rossi, Franck Dernoncourt, Roger Zimmermann, Yue Wang, Zhengzhong Tu, Vicente Ordonez, Mohit Bansal, Yue Zhao · 31 juillet 2026
Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions. We introduce \textit{\ours{}}, a benchmark with tab-and-blank interlocking pieces where geomet…
- torchsom: The Reference PyTorch Library for Self-Organizing Maps
Louis Berthier, Ahmed Shokry, Maxime Moreaud, Guillaume Ramelet, Eric Moulines · 24 juillet 2026
This paper introduces torchsom, an open-source Python library that provides a reference implementation of the Self-Organizing Map (SOM) in PyTorch. This package offers three main features: (i) dimensionality reduction, (ii) clustering, and (iii) friendly data visualization. It relies on a PyTorch ba…
- From Preimage Search To Source-Grounded Feature Inversion
Kaixiang Shu · 15 juillet 2026
Interpreting a neural network requires understanding what its internal features extract from a particular input. Feature inversion seeks to express a selected feature in the input domain, but canonical iterative methods search for an input whose re-encoded representation matches the target. Because …
- A Vision Based System for Guided and Collaborative Reconstruction of Fragmented Documents
Oliver Krumpek, Diana Leo · 7 juillet 2026
This paper presents the development and evaluation of a collaborative system for real-time reconstruction of fragmented paper documents in the context of cultural heritage preservation. The developed system includes a collaborative robot, or cobot, that can fully manage the positioning of paper frag…
- Structural Pattern Mining in Inka Khipus: Unsupervised Clustering, Provenance Classification, and a Computational Validation of the Santa Valley Match
Maria Contreras · 2 juillet 2026
Khipus -- knotted cord devices -- were the primary recording medium of the Inka Empire (c. 1400-1532 CE), yet their system remains undeciphered. We present a reproducible machine-learning pipeline applied to the Open Khipu Repository (OKR), a public database of 619 khipus comprising 54,403 cords and…
- Enhancing Oracle Bone Inscription Recognition via Multi-Scale Layer Attention
Chaowen Yan, Kaishen Wang, Yong Wang, Jianlong Xiong, Tao He · 2 juillet 2026
Oracle Bone Inscriptions (OBIs) recognition plays a crucial role in understanding ancient Chinese culture. However, accurately recognizing OBIs remains highly challenging due to their complex, irregular, and often degraded shapes. Traditional methods rely on expert knowledge and manual analysis, whi…
- Automated sign detection across the Electronic Babylonian Library: A large-scale dataset and end-to-end cuneiform OCR pipeline
Wentao Che, Esteban Garcés Arias, Asim Niaz, Andreas Bender, Enrique Jiménez · 23 juin 2026
Learning to read cuneiform tablets is an extremely demanding task; consequently, of the roughly half million excavated tablets, only a small fraction has been analysed by Assyriologists. Computer vision offers a promising avenue for decipherment but requires large, densely annotated datasets. To add…
- GAPartManip: A Large-scale Part-centric Dataset for Material-Agnostic Articulated Object Manipulation
Wenbo Cui, Chengyang Zhao, Songlin Wei, Jiazhao Zhang, Haoran Geng, Yaran Chen, Haoran Li, He Wang · 23 juin 2026
Effectively manipulating articulated objects in household scenarios is a crucial step toward achieving general embodied artificial intelligence. Mainstream research in 3D vision has primarily focused on manipulation through depth perception and pose detection. However, in real-world environments, th…
- Variational autoencoders with latent high-dimensional steady geometric flows for dynamics
Andrew Gracyk · 17 juin 2026
We develop Riemannian approaches to variational autoencoders (VAEs) for PDE-type ambient data with regularizing geometric latent dynamics, which we refer to as VAE-DLM, or VAEs with dynamical latent manifolds. We redevelop the VAE framework such that manifold geometries, subject to our geometric flo…
- Best Arm Identification with Minimal Regret
Junwen Yang, Vincent Y. F. Tan, Tianyuan Jin · 16 juin 2026
Motivated by real-world applications that necessitate responsible experimentation, we introduce the problem of best arm identification (BAI) with minimal regret. This variant of the multi-armed bandit problem elegantly amalgamates two of its most ubiquitous objectives: regret minimization and BAI. M…
- Unifying Low Dimensional Spectra in Deep Learning
Connall Garrod, Jonathan P. Keating · 28 mai 2026
Low dimensional structures appear ubiquitously in the eigenspectra of deep learning matrices in classification networks trained in the overparameterized regime. While theoretical advances have aimed to explain this phenomenology, they typically succeed only in capturing subsets of the full behavior …
- Learning Permutation from Structure Without Supervision
Ran Eisenberg, Ofir Lindenbaum · 26 mai 2026
Many learning problems require uncovering a hidden ordering that reveals structure in unordered data, such as monotonicity in sorting or spatial continuity in jigsaw reconstruction. In these settings, permutations can be learned as latent operators by optimizing objectives defined directly on the re…
- Machine learning applied to emerald gemstone grading: framework proposal and creation of a public dataset
FB Pena, D Crabi, Sandro C Izidoro, Érick O Rodrigues, G Bernardes · 25 mai 2026
The grading of gemstones is currently a manual procedure performed by gemologists. A popular approach uses reference stones, where those are visually inspected by specialists that decide which one of the available reference stone is the most similar to the inspected stone. This procedure is very sub…
- Mixup Barcodes: Quantifying Geometric-Topological Interactions between Point Clouds
Hubert Wagner, Nickolas Arustamyan, Matthew Wheeler, Peter Bubenik · 19 mai 2026
We combine standard persistent homology with image persistent homology to define a novel way of characterizing shapes and interactions between them. In particular, we introduce: (1) a mixup barcode, which captures geometric-topological interactions (mixup) between two point sets in arbitrary dimensi…
Autres sujets du thème Vision par ordinateur et reconnaissance de formes
Les sujets rattachés au même thème par la classification OpenAlex, les plus actifs d'abord.
- Multimodal Machine Learning Applications7 825 papiers / 12 mois+585 %
- Generative Adversarial Networks and Image Synthesis4 914 papiers / 12 mois+413 %
- Advanced Neural Network Applications2 314 papiers / 12 mois+419 %
- Advanced Vision and Imaging809 papiers / 12 mois+261 %
- Human Pose and Action Recognition797 papiers / 12 mois+957 %
- Face recognition and analysis467 papiers / 12 mois+338 %
