Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Handwritten Text Recognition Techniques
397 artículos indexados
Las técnicas de reconocimiento de texto manuscrito se centran en extraer e interpretar caracteres o estructuras escritas a mano, ya sean documentos históricos, fórmulas científicas o soportes visuales complejos. Combinan enfoques de Computer Vision y procesamiento automático del lenguaje para analizar datos reales o sintéticos, evaluar la robustez de los modelos frente a perturbaciones o adaptar métodos a sistemas de escritura específicos como logogramas o cuneiformes. Estos trabajos exploran tanto la mejora del rendimiento en corpus especializados como la integración de contextos multimodales, como el diseño de página o las anotaciones, para refinar la comprensión automática de los documentos.
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- China29 % · 73 artículos
- Estados Unidos26 % · 64 artículos
- Francia11 % · 27 artículos
- India7,2 % · 18 artículos
- Japón6,4 % · 16 artículos
- Alemania6 % · 15 artículos
- Reino Unido4 % · 10 artículos
- RAE de Hong Kong (China)4 % · 10 artículos
Sobre 250 artículos de este tema con al menos un laboratorio localizado. 60 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- STEPS: Scene Text Editing with Preserved Style Using Diffusion and Contrastive Style Encoding
Nicolas Thiebaut, Nameer Hirschkind, Xiao Yu, Kyle Spence · 1 de octubre de 2026
We introduce Scene Text Editing with Preserved Style (STEPS), a novel diffusion model architecture for quality text replacement in images. Scene Text Editing (STE), also known as Visual Text Editing, consists of changing the textual content in an image while conserving the original style, e.g. font,…
- Which papyrus HTR is good enough? Character-error-rate tolerance of four papyrological tasks on Greek texts
Anton Repushko, Elena Chepel · 1 de octubre de 2026
Purpose: Most Greek papyri remain unpublished and undigitised; a handwritten text recognition (HTR) pipeline that transcribes them automatically would let scholars discover documents and literary works that have so far gone unread. Recognition systems for Ancient Greek papyri are in statu nascendi, …
- Developing an OCR model for Extracting Information from Invoices with Korean Language
Xiem HoangVan, Phu TranQuang, Minh DinhBao, Tien VuHuu · 1 de octubre de 2026
Invoices are commercial documents that contain various pieces of information, including the purchased items, time, and total money. Making the extraction of important information crucial. The stored information serves different purposes. Korean language is the native language of about 80 million peo…
- PolyOCR-Venus: Unified OCR Foundation Models for Text-Centric Visual Intelligence
GuangJian Team, Kaili Huang, Yongshuo Zhang, Bingtao Fu, Changjiang Jiang, Chenfan Qu, Chenfeng Zhang, Fangming Cui, Gaoyang Zhang, Jiangwei Xie, Jianshu Li, Jing Huang, Jingwen Bai, Mingqi Fang, Tao Fang, Weihong Zhang, Wenbo Du, Xiongfei Bai, Xuekang Zhu, Yinan Xia, Zhenming Wang, Jian Liu, Jingjing Liu, Xiang Qi, Weiqiang Wang · 30 de septiembre de 2026
Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, localize, and reason over textual information in complex visual environments. However, existing OCR systems often excel at only some tasks and struggle to …
- Age-Adaptive Handwriting Reconstruction from an IMU-Based Digital Pen through Shared Representations and Domain-Specific Heads
Florent Imbert (LUT), Yann Soullard (IRISA, UR2, SHADOC), Eric Anquetil (INSA Rennes, IRISA, SHADOC), Hui Han (LUT) · 29 de septiembre de 2026
Digital pens are widely used to capture handwriting on digital devices, enabling precise trace recording and enhancing human-computer interaction. However, most are bundled with tablets and lack cross-brand compatibility. Recent digital pens equipped with kinematic sensors have emerged, designed for…
- PalmLeaf-VQA: A Multi-Script Visual Question Answering Benchmark for Historical Palm-Leaf Manuscript Understanding Across Diverse Regions
Nimol Thuon, Jun Du, Panhapin Theang · 29 de septiembre de 2026
Historical manuscripts remain largely absent from modern vision-language benchmarks, leaving open how well multimodal large language models (MLLMs) handle culturally diverse, degraded, and non-Latin document images. We introduce \textbf{PalmLeaf-VQA}, a multi-script visual question answering benchma…
- Source-preserving alignment for robust evidence localization in scientific PDFS
Zihao Liu, Wei Yang, Zixiao Dong, Chenshu Li, Longzhang Liu, Tao Tan, Hong Xie · 29 de septiembre de 2026
Scientific information-extraction systems often return a claim with an evidence string, which users must locate in the original PDF. This is challenging because the extracted evidence and PDF text layer are different representations: line wrapping, Unicode variants, superscripts, citation markers, a…
- Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers
Mert \.Incidelen, Yamen Kashkash, Asya Berker, Murat Aydo\u{g}an · 28 de septiembre de 2026
Vision-language models (VLMs), despite their success in optical character recognition (OCR) tasks, are vulnerable to typographic attacks and have a fragile structure for images with multiple text layers. In this study, the DecoyBench dataset was created using the Decoy Font method. The dataset consi…
- The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models
Christoph Walser, Mauricio Fadel Argerich, Jonathan F\"urst · 28 de septiembre de 2026
Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum. We study this trade-off under a constraint that rules out (closed) cloud services: privacy-sensitive documents processed on-premise by small …
- MEVL-STP: Multi-Encoder and Vision Language Model for Arbitrarily Shaped Scene Text Spotting
Aman Anand, Partha Pratim Roy, Shivakumara Palaiahnakote · 25 de septiembre de 2026
Scene text spotting remains challenging for arbitrarily shaped text instances such as curved signs and dense multi-oriented characters in natural images, where tightly coupled architectures propagate localization errors directly into recognition failures. We present a two-stage pipeline that combine…
- Diffusion Drafts, AR Verifies: Accelerating Document OCR with Self-Speculative Decoding
Dohyun Kim, Sungjun Han, Hyungguk Kim, Yusik Kim, Jamin Shin, Paul Hongsuck Seo, Hongjoon Ahn · 23 de septiembre de 2026
Autoregressive OCR vision-language models accurately convert document images into text and structured markup, but require one sequential decoding step per output token, limiting inference speed. Unlike open-ended text generation, OCR outputs are strongly grounded in the input image, making diffusion…
- Layout-Guided Masking for GROBID: Lightweight Structural Gains in Large-Scale Scientific PDF Ingestion
Luca Foppiano, Sana Khamassi, Vipul Gupta · 23 de septiembre de 2026
Transforming scholarly PDFs into machine-readable fulltext remains a bottleneck for large-scale information systems. Recent vision-based parsers improve accuracy, but need GPUs and may introduce noise into the extracted text. GROBID, a modular font-stream parser running on CPU, is the de-facto stand…
- All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen · 22 de septiembre de 2026
Multilingual scene text recognition (STR) remains challenging due to the scarcity of training data for most languages and the difficulty of serving diverse scripts within a single model. Existing solutions either deploy one recognizer per language, inflating cost and introducing error accumulation, …
- Mind the Gaps: A Curated Benchmark for Form Field Detection
Iheb Brini, Omar Moured, Hamza Gbada, Elisa Barney · 22 de septiembre de 2026
Form Field Detection (FFD) is a fundamental component of document understanding systems, enabling applications ranging from large-scale industrial digitization to accessible form interaction for automated analysis. Unlike conventional object detection tasks, FFD is inherently challenging because fie…
- From Retrieval to Recognition:How Vision--Language Models Become OCR Specialists
Yuanxiang Huangfu, Hanmeng Zhong, Linqing Chen, Jeffrey Tiong Jee Hui · 21 de septiembre de 2026
Does a general vision--language model acquire specialized OCR ability by developing a new reading circuit or by reusing an existing mechanism? We address this question in the setting of full-sequence OCR, rather than local-answer retrieval. Using an evidence-grounded protocol with held-out causal in…
- Evaluating In-Context Learning and Retrieval Strategies for Devanagari Post-OCR Correction
Abhishek Bhandari, Gaurav Harit · 21 de septiembre de 2026
In-context learning using Large Language Models (LLMs) offers a compelling path to training-free post-OCR correction, yet its effectiveness for Devanagari script remains entirely unexplored. We present the first systematic evaluation of LLMs (3B-32B) for post-OCR correction in Hindi and Marathi, com…
- WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing
Hao Yu, Kang Liu, Linnan Zhao, Jiabo Zhan, Chong Sun, Chen Li, Jing Lyu · 18 de septiembre de 2026
Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are biased toward common document types and clean digital pages, while expanding coverage alone does not specify how to address …
- A Free Lunch? Adapting PP-OCRv6 for Historical Text Recognition
Benjamin Kiessling (ALMAnaCH) · 18 de septiembre de 2026
Despite impressive reported scores, large vision-language models have seen limited practical uptake in historical automatic text recognition because of their computational cost, dependence on large-scale pretraining, and hallucination. Historical ATR therefore continues to rely largely on compact CR…
- LinePilot Digitizer: Line-Plot Recovery with Manual and Automatic Calibration
Fengbo Ma, Rayan Akhtar, Aakash H. Joshi, Xiaoting Li, Haijian Sun, Zhen Xiang, Xianyan Chen, Yiping Zhao · 18 de septiembre de 2026
Recovering numerical series from line plots requires accurate axis calibration and reliable curve extraction. We present LinePilot Digitizer (LinePilot), which combines continuous color-based curve recovery with three calibration modes: LinePilot (standard), LinePilot (enhanced), and LinePilot (OCR)…
- Perception, Layout, and Validation: Calibrated Confidence for Reliable Straight-Through Processing of Financial Documents
Yichao Jin, Yushuo Wang, Yuxuan Han, Kwan Ching Yee Sonia, Weiyang Song, Chiu Jin-Chun Kent, Wong Chong Hwee, Wong Tiong Kiat, Kenneth Zhu Ke, Jingyuan Zhao · 18 de septiembre de 2026
Straight-through processing (STP) on extracted key-value fields from financial documents without human review requires a calibrated probability together with a bounded guarantee on the residual error of the auto-approved tier. The emergence of modern Vision Language Models (VLMs) provides an out-of-…
- FROD: Feature Matching Residual Denoising Oracle Bone Decipher
Yanbin Hou, Biao Xiong, Guojun Xu, Jianwen Xiang, Cheng Tan, Yanchao Yang, Junwei Zhou · 16 de septiembre de 2026
Oracle bone script (OBS), one of the earliest Chinese writing systems, plays an important role in the study of Chinese etymology. Traditional decipherment relies heavily on domain experts who analyze characters through semantic context and structural evolution. To assist this labor-intensive process…
- Measuring Annotation Efficiency for Handwritten Devanagari Recognition: Sample-Complexity Curves for Four Pretraining Regimes
Manglesh Kumar Pandey, Sumit Kumar Banshal · 16 de septiembre de 2026
To train handwritten text recognition systems we need word images and their corresponding transcriptions, and these transcriptions are produced manually. For a script that can be read by only a small number of specialists, this manual transcription is a limitation, because the trained models are sup…
- Beyond Accuracy: Robustness, Cost, and Governance Trade-offs for Vision-Language Models in Templated Document Extraction
Kushal Patel, Pushkal Shrivastava, Mackenzie Lees, Qirui Lu, Bhargobjyoti Saikia, Liying Li, Junlin Jiang · 15 de septiembre de 2026
Vision-language models (VLMs) are increasingly used to extract structured fields from business documents, yet most evaluations report accuracy on clean benchmarks and offer little guidance to practitioners choosing an approach for a given task complexity. We address this gap with a measurement-groun…
- Write on Paper and Get the Online Digital Trace:\newline A New Era for Handwriting
Florent Imbert, Yann Soullard, Eric Anquetil, Tanja Harbaum, Alexey Serdyuk, Fabian Kress, Tim Hamann, Peter Kampf · 14 de septiembre de 2026
Capturing the digital trace of handwriting usually requires a specific stylus and a compatible substrate, be it a capacitive touchscreen, an ElectroMagnetic Resonance (EMR) tablet as used in Wacom systems or special paper. While writing on regular paper offers rich haptics, no latency and is well kn…
- ExpertHTR: Unified Handwritten Text Recognition with Multi-Task Learning and Sparse Mixture-of-Experts
Dang Hoai Nam, Nguyen Duy Hieu, Quang Huu Hieu, Vo Nguyen Le Duy · 14 de septiembre de 2026
Handwritten text recognition resources are often small and distributed across collections that differ in language, script, document structure, and annotation format, making joint page-level training difficult. We propose ExpertHTR, a unified vision-language framework that addresses this problem thro…
Otros asuntos del tema Visión por computador y reconocimiento de formas
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Multimodal Machine Learning Applications8069 artículos / 12 meses+191 %
- Generative Adversarial Networks and Image Synthesis4992 artículos / 12 meses+39 %
- Advanced Neural Network Applications2354 artículos / 12 meses+48 %
- Advanced Vision and Imaging841 artículos / 12 meses+78 %
- Human Pose and Action Recognition836 artículos / 12 meses+457 %
- Face recognition and analysis482 artículos / 12 meses+88 %
