Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Image Retrieval and Classification Techniques
91 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Últimos artículos
- DEPICT: Scoring Text-to-Image Alignment by Answer Agreement
Vasco Ramos, Sandra Godinho Silva, Joao Magalhaes, Ricardo Rei, Pedro Henrique Martins · 5 de octubre de 2026
Image-text alignment is a core problem in computer vision with applications in caption evaluation, hallucination detection, data curation, and the benchmarking of text-to-image (T2I) generators. As T2I models improve, benchmarking has become demanding, requiring metrics capable of finding a series o…
- Where to Look Is Not How to Fix: Pre-Denoising Diagnostics and Modality-Dependent Control in Diffusion Composition
Fangzheng Wu, Brian Summa · 5 de octubre de 2026
Understanding compositional failures in text-to-image diffusion requires identifying both where stress is detectable and how intervention changes the output. We study these questions through a controlled anchor--stress protocol that jointly evaluates text-encoder diagnostics and denoiser interventio…
- NegT2IBench: When Negation Changes the Picture. A Polarity Benchmark for Text-to-Image Models
Omar Elfatairy, Maria A. Bravo, Jessica Bader, Zeynep Akata · 5 de octubre de 2026
Text-to-image (T2I) models are judged by benchmarks that measure whether requested content appears, but these benchmarks largely overlook the complementary ability to satisfy negated constraints, for example, generating "a non-red cup." Measuring negation raises challenges not faced by affirmation-b…
- Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards
Yuanhao Ban, I-Hung Hsu, Anastasios Angelopoulos, Wei-Lin Chiang, Ion Stoica, Cho-Jui Hsieh · 5 de octubre de 2026
Recent text-to-image generation models have achieved remarkable visual quality, but improving them through post-training remains challenging because no single reward signal captures the full range of human preference. In this work, we develop a simple and effective post-training recipe for open-doma…
- Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis
Peilin Yang, Xiaoyu Liu, Jian Sun, Qinghua Tao · 5 de octubre de 2026
Vision-language models such as CLIP exhibit strong semantic generalization, but remain limited in fine-grained visual perception. A recent work named KUEA presents a natural remedy by finetuning the image encoder under the supervision of the vision-centric DINOv2 to align their kernel matrices eleme…
- Traversing the Satisfaction-Diversity Frontier in Text-to-Image Diffusion
Kevin Zhai, Siva Rajesh Kasa, Soumya Roy, Sumit Negi, Mubarak Shah · 5 de octubre de 2026
Text-to-image generation enables users to explore several images generated from the same prompt. For these generated images to be useful, each one must reflect the user's preferences, measured by a learned reward, and differ visually from the others to maintain diversity. Existing methods are limite…
- FairSSL: Fair Multimodal Self-Supervised Learning
Jiaee Cheong, Abtin Mogharabin, Paul Liang, Hatice Gunes, Sinan Kalkan · 2 de octubre de 2026
Prevalent multimodal self-supervised learning (SSL) methods rely on the redundancy assumption: that different views share substantial task-relevant information. We argue that this assumption fails in complex, real-world settings characterized by heterogeneity (e.g., variable-length healthcare or beh…
- Walking the Embedding Space: Datastore Extraction from Multimodal RAG
Maria Carmen Jica, Ali Satvaty, Suzan Verberne, Fatih Turkmen · 2 de octubre de 2026
Multimodal Retrieval-Augmented Generation (MRAG) has emerged as a reliable and cost-effective technique of grounding the generative capabilities of Multimodal Large Language Models (MLLMs) into relevant, up-to-date, external knowledge. Despite presenting several benefits, such as reducing hallucinat…
- A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions
Giyeong Oh, Junghun Park, Yuhan Bae, Youngjae Yu · 2 de octubre de 2026
Recaptioned image-text corpora are now standard for text-to-image (T2I) training, with vision--language model (VLM) captioners replacing sparse alt-text by dense descriptions. A recaptioned corpus is a supervision distribution induced by a documented captioning policy ($\pi$), captioner ($V_c$), and…
- Personalized Image Generation with Reasoning and Reflection
Bo Ni, Ngoc N. Tran, Qinwen Ge, Franck Dernoncourt, Seunghyun Yoon, Samyadeep Basu, Sungchul Kim, Puneet Mathur, Nedim Lipka, Tong Yu, Yu Wang, Ryan A. Rossi, Tyler Derr · 2 de octubre de 2026
Personalized image generation has remained narrowly focused on conditional synthesis from curated visual exemplars, rather than capturing who a user is. In practice, however, a user's personal context is much richer, comprising reviews, posts, images, captions, and metadata accumulated over time. A …
- Frozen Scenes, Shifting Winners: Configuration Fragility in Text-to-3D Evaluation
Anson Y. Lam, Shuqing Li, Michael R. Lyu · 2 de octubre de 2026
Can a text-to-3D leaderboard change when every generated scene stays fixed? We audit this question for rendered-image evaluation, where camera settings and caption wording become part of the measurement protocol. Across 300 frozen scenes from six generators, we vary eight render and caption factors …
- SAGE: Salient Factor Discovery and Generation with Visual Foundation Representations
Shuang Liang, Lejun Liao, Shiyuan Zhang, Max C. Zhang, Xiaolong Luo, Han Wang, Stefano Anzellotti, Yuan Yuan · 1 de octubre de 2026
Given a target dataset, such as faces with eyeglasses, and a background dataset, such as faces without, contrastive analysis separates \textit{salient} factors specific to the target from \textit{common} content shared by both. We aim for salient representations that capture target-specific detail i…
- HiRAE: Hierarchical Representation Autoencoding with Residual Budgets
Xuanyu Zhu, Yan Bai, Yang Shi, Yihang Lou, Yuanxing Zhang, Tengfei Liu, Jing Jin, Yuan Zhou · 30 de septiembre de 2026
Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for reconstruction can produce a latent distribut…
- Preserve-and-Compose Training for Composed Image Retrieval
Sehyun Kwon · 28 de septiembre de 2026
Composed image retrieval (CIR) aims to retrieve images that satisfy a user-specified modification while preserving relevant visual content from a reference image. Collecting target images for this purpose is costly, motivating zero-shot CIR methods that instead use target captions as supervision. Ho…
- Evaluation of OpenAI o1: Opportunities and Challenges of AGI
Tianyang Zhong, Zhengliang Liu, Yi Pan, Yutong Zhang, Zeyu Zhang, Yifan Zhou, Shizhe Liang, Zihao Wu, Yanjun Lyu, Peng Shu, Xiaowei Yu, Chao Cao, Hanqi Jiang, Hanxu Chen, Yiwei Li, Junhao Chen, Huawen Hu, Yiheng Liu, Huaqin Zhao, Shaochen Xu, Haixing Dai, Lin Zhao, Ruidong Zhang, Wei Zhao, Zhenyuan Yang, Jingyuan Chen, Peilong Wang, Wei Ruan, Hui Wang, Huan Zhao, Jing Zhang, Yiming Ren, Shihuan Qin, Tong Chen, Jiaxi Li, Arif Hassan Zidan, Afrar Jahin, Minheng Chen, Sichen Xia, Jason Holmes, Yan Zhuang, Jiaqi Wang, Bochen Xu, Weiran Xia, Jichao Yu, Kaibo Tang, Yaxuan Yang, Bolun Sun, Lifeng Chen, Tao Yang, Guoyu Lu, Xianqiao Wang, Lilong Chai, He Li, Jin Lu, Xin Zhang, Bao Ge, Xintao Hu, Lian Zhang, Hua Zhou, Lu Zhang, Shu Zhang, Zhen Xiang, Yudan Ren, Jun Liu, Xi Jiang, Yu Bao, Wei Zhang, Xiang Li, Gang Li, Wei Liu, Dinggang Shen, Andrea Sikora, Xiaoming Zhai, Dajiang Zhu, Tuo Zhang, Tianming Liu · 25 de septiembre de 2026
This comprehensive study evaluates the performance of OpenAI's o1-preview large language model across a diverse array of complex reasoning tasks, spanning multiple domains, including computer science, mathematics, natural sciences, medicine, linguistics, and social sciences. Through rigorous testing…
- Beyond Mean Foils: Auditing Worst-Foil Specificity in Frozen CLIP Region Explanations
Kaixin Liu, Zhipeng Ye, Feng Jiang, Zhenghao Wang, Qihang Wu · 24 de septiembre de 2026
A region can overlap a target object yet contribute more to another class. We test regions selected by Cluster-based Concept Importance (CCI) in frozen CLIP. Across COCO and VOC with two checkpoints, 41.08-64.78% of regions that pass overlap and mean-contrast checks fail against the strongest compet…
- MINER: Multi-crop INference-time Enhancement for Rare-Object Retrieval with Frozen Dual Encoders
Abdulmalik Alquwayfili, Faisal AlMeshal, Jumanah Almajnouni, Huda Abdulhadi Alamri, Muhammad Kamran J Khan · 24 de septiembre de 2026
Text-to-image retrieval with frozen dual encoders degrades when the query names a small, visually subordinate object in a cluttered scene: a single global image embedding underrepresents the localized visual evidence. We present MINER, a training-free inference framework that augments a frozen dual …
- Graded-Relevance Composed Multimodal Retrieval for E-commerce Visual Search at Scale
Anubhav Gupta, Hrushikesh Mohapatra, Prijith Chandra, Asish Mohapatra, Anuj Garg, Arvind Maan, Sudip Datta, Venkat Bulusu, Sitesh Kumar Jalan · 22 de septiembre de 2026
Visual search on large e-commerce catalogs must serve both "similarity" queries that ask for items resembling an uploaded image and "modifier" queries that comprise an image and text describing a desired modification (e.g. a color change or style swap). The latter is the setting known as composed im…
- Incentive Noise and Structural Prior Infusion for Multi-modal Object Re-Identification
Weixiang Zhou, Yuhao Wang, Xingguo Xu, Weizhen Zhou, Zhixun Su, Jinshan Pan, Cong Wang · 22 de septiembre de 2026
Multi-modal object Re-Identification (ReID) benefits from complementary information across heterogeneous imaging modalities. To further enrich semantic representation, text descriptions have recently been incorporated as an additional modality. However, recent vision-language approaches often treat …
- IMPLICIT-Bench: Measuring Implicit Bias in Text-to-Image Models under Neutral Prompts
Yue Dai, Ziyang Liu, Marc Cheong, Caren Han · 22 de septiembre de 2026
Text-to-image (T2I) models are typically evaluated for bias using slot-based templates such as ``a photo of a [profession]''. Such templates probe only \emph{explicit} demographic attributes (e.g., gender, skin tone) in isolation. They overlook a broader \emph{implicit} bias that arises in natural p…
- InterHier: Learning Interconnected Hierarchical Semantics for Open-Vocabulary Object Detection
Yeong-Jin Kim, Ho-Joong Kim, Seong-Whan Lee · 22 de septiembre de 2026
In this paper, we investigate the limitations of fixed, hand-crafted connectors in hierarchical semantic representations for open-vocabulary object detection. Existing methods establish semantic relationships between base categories and unseen novel categories by placing a fixed connector between ad…
- RA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes
Abhirama Subramanyam Penamakuri, Shreya Shukla, Anand Mishra · 15 de septiembre de 2026
Fashion Image Captioning (FIC) plays a vital role in enhancing user experience and product search in e-commerce platforms. Unlike natural scene image captioning, FIC requires fine-grained visual reasoning and knowledge of domain-specific terminology to capture subtle attributes such as neckline and …
- PinDCO: Whole-Page Aware Dynamic Creative Optimization at Scale
Yu Hao, Yuchun Li, Peimeng Sui, Meilin Liu, Tianyuan Cui, Hao Li, Zicong Zhou, Akanksha Baid · 14 de septiembre de 2026
Recent advances in generative AI have substantially accelerated the creation of high-quality ad creatives, dramatically expanding the number of candidate variants per campaign. This shift increases the need for scalable dynamic creative optimization (DCO) systems that can match creatives to the most…
- Generative Retrieval for Unsupervised Text-Based Person Search
Mang Ye, Yucheng Ji, Yang Bai, Min Cao, Siyuan Chai, Bo Du, Min Zhang · 14 de septiembre de 2026
Text-based person search (TBPS) aims to retrieve images of a target person from a large image gallery based on a given natural language description. Most existing methods rely on supervised learning with manually annotated image-text pairs. In this paper, we explore unsupervised TBPS, with only unla…
- RiVaT-Fuse: Reliability-Calibrated Variational Tensor Fusion for Multimodal Prediction under Modality Uncertainty
Yingfan Xu, Tieming Liu, Ye Liang, Taiping Liu · 11 de septiembre de 2026
Image-metadata prediction requires fusing heterogeneous evidence whose reliability can vary across samples and latent factors. Existing representation-level fusion methods typically choose an aggregation architecture, such as concatenation, gating, conditional modulation, or attention, without expli…
Otros asuntos del tema Visión por computador y reconocimiento de formas
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Multimodal Machine Learning Applications8069 artículos / 12 meses+191 %
- Generative Adversarial Networks and Image Synthesis4992 artículos / 12 meses+39 %
- Advanced Neural Network Applications2354 artículos / 12 meses+48 %
- Advanced Vision and Imaging841 artículos / 12 meses+78 %
- Human Pose and Action Recognition836 artículos / 12 meses+457 %
- Face recognition and analysis482 artículos / 12 meses+88 %
