Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Image Retrieval and Classification Techniques
91 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Neueste Paper
- FairSSL: Fair Multimodal Self-Supervised Learning
Jiaee Cheong, Abtin Mogharabin, Paul Liang, Hatice Gunes, Sinan Kalkan · 2. Oktober 2026
Prevalent multimodal self-supervised learning (SSL) methods rely on the redundancy assumption: that different views share substantial task-relevant information. We argue that this assumption fails in complex, real-world settings characterized by heterogeneity (e.g., variable-length healthcare or beh…
- Walking the Embedding Space: Datastore Extraction from Multimodal RAG
Maria Carmen Jica, Ali Satvaty, Suzan Verberne, Fatih Turkmen · 2. Oktober 2026
Multimodal Retrieval-Augmented Generation (MRAG) has emerged as a reliable and cost-effective technique of grounding the generative capabilities of Multimodal Large Language Models (MLLMs) into relevant, up-to-date, external knowledge. Despite presenting several benefits, such as reducing hallucinat…
- A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions
Giyeong Oh, Junghun Park, Yuhan Bae, Youngjae Yu · 2. Oktober 2026
Recaptioned image-text corpora are now standard for text-to-image (T2I) training, with vision--language model (VLM) captioners replacing sparse alt-text by dense descriptions. A recaptioned corpus is a supervision distribution induced by a documented captioning policy ($\pi$), captioner ($V_c$), and…
- Personalized Image Generation with Reasoning and Reflection
Bo Ni, Ngoc N. Tran, Qinwen Ge, Franck Dernoncourt, Seunghyun Yoon, Samyadeep Basu, Sungchul Kim, Puneet Mathur, Nedim Lipka, Tong Yu, Yu Wang, Ryan A. Rossi, Tyler Derr · 2. Oktober 2026
Personalized image generation has remained narrowly focused on conditional synthesis from curated visual exemplars, rather than capturing who a user is. In practice, however, a user's personal context is much richer, comprising reviews, posts, images, captions, and metadata accumulated over time. A …
- Frozen Scenes, Shifting Winners: Configuration Fragility in Text-to-3D Evaluation
Anson Y. Lam, Shuqing Li, Michael R. Lyu · 2. Oktober 2026
Can a text-to-3D leaderboard change when every generated scene stays fixed? We audit this question for rendered-image evaluation, where camera settings and caption wording become part of the measurement protocol. Across 300 frozen scenes from six generators, we vary eight render and caption factors …
- SAGE: Salient Factor Discovery and Generation with Visual Foundation Representations
Shuang Liang, Lejun Liao, Shiyuan Zhang, Max C. Zhang, Xiaolong Luo, Han Wang, Stefano Anzellotti, Yuan Yuan · 1. Oktober 2026
Given a target dataset, such as faces with eyeglasses, and a background dataset, such as faces without, contrastive analysis separates \textit{salient} factors specific to the target from \textit{common} content shared by both. We aim for salient representations that capture target-specific detail i…
- HiRAE: Hierarchical Representation Autoencoding with Residual Budgets
Xuanyu Zhu, Yan Bai, Yang Shi, Yihang Lou, Yuanxing Zhang, Tengfei Liu, Jing Jin, Yuan Zhou · 30. September 2026
Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for reconstruction can produce a latent distribut…
- Preserve-and-Compose Training for Composed Image Retrieval
Sehyun Kwon · 28. September 2026
Composed image retrieval (CIR) aims to retrieve images that satisfy a user-specified modification while preserving relevant visual content from a reference image. Collecting target images for this purpose is costly, motivating zero-shot CIR methods that instead use target captions as supervision. Ho…
- Evaluation of OpenAI o1: Opportunities and Challenges of AGI
Tianyang Zhong, Zhengliang Liu, Yi Pan, Yutong Zhang, Zeyu Zhang, Yifan Zhou, Shizhe Liang, Zihao Wu, Yanjun Lyu, Peng Shu, Xiaowei Yu, Chao Cao, Hanqi Jiang, Hanxu Chen, Yiwei Li, Junhao Chen, Huawen Hu, Yiheng Liu, Huaqin Zhao, Shaochen Xu, Haixing Dai, Lin Zhao, Ruidong Zhang, Wei Zhao, Zhenyuan Yang, Jingyuan Chen, Peilong Wang, Wei Ruan, Hui Wang, Huan Zhao, Jing Zhang, Yiming Ren, Shihuan Qin, Tong Chen, Jiaxi Li, Arif Hassan Zidan, Afrar Jahin, Minheng Chen, Sichen Xia, Jason Holmes, Yan Zhuang, Jiaqi Wang, Bochen Xu, Weiran Xia, Jichao Yu, Kaibo Tang, Yaxuan Yang, Bolun Sun, Lifeng Chen, Tao Yang, Guoyu Lu, Xianqiao Wang, Lilong Chai, He Li, Jin Lu, Xin Zhang, Bao Ge, Xintao Hu, Lian Zhang, Hua Zhou, Lu Zhang, Shu Zhang, Zhen Xiang, Yudan Ren, Jun Liu, Xi Jiang, Yu Bao, Wei Zhang, Xiang Li, Gang Li, Wei Liu, Dinggang Shen, Andrea Sikora, Xiaoming Zhai, Dajiang Zhu, Tuo Zhang, Tianming Liu · 25. September 2026
This comprehensive study evaluates the performance of OpenAI's o1-preview large language model across a diverse array of complex reasoning tasks, spanning multiple domains, including computer science, mathematics, natural sciences, medicine, linguistics, and social sciences. Through rigorous testing…
- Beyond Mean Foils: Auditing Worst-Foil Specificity in Frozen CLIP Region Explanations
Kaixin Liu, Zhipeng Ye, Feng Jiang, Zhenghao Wang, Qihang Wu · 24. September 2026
A region can overlap a target object yet contribute more to another class. We test regions selected by Cluster-based Concept Importance (CCI) in frozen CLIP. Across COCO and VOC with two checkpoints, 41.08-64.78% of regions that pass overlap and mean-contrast checks fail against the strongest compet…
- MINER: Multi-crop INference-time Enhancement for Rare-Object Retrieval with Frozen Dual Encoders
Abdulmalik Alquwayfili, Faisal AlMeshal, Jumanah Almajnouni, Huda Abdulhadi Alamri, Muhammad Kamran J Khan · 24. September 2026
Text-to-image retrieval with frozen dual encoders degrades when the query names a small, visually subordinate object in a cluttered scene: a single global image embedding underrepresents the localized visual evidence. We present MINER, a training-free inference framework that augments a frozen dual …
- Graded-Relevance Composed Multimodal Retrieval for E-commerce Visual Search at Scale
Anubhav Gupta, Hrushikesh Mohapatra, Prijith Chandra, Asish Mohapatra, Anuj Garg, Arvind Maan, Sudip Datta, Venkat Bulusu, Sitesh Kumar Jalan · 22. September 2026
Visual search on large e-commerce catalogs must serve both "similarity" queries that ask for items resembling an uploaded image and "modifier" queries that comprise an image and text describing a desired modification (e.g. a color change or style swap). The latter is the setting known as composed im…
- Incentive Noise and Structural Prior Infusion for Multi-modal Object Re-Identification
Weixiang Zhou, Yuhao Wang, Xingguo Xu, Weizhen Zhou, Zhixun Su, Jinshan Pan, Cong Wang · 22. September 2026
Multi-modal object Re-Identification (ReID) benefits from complementary information across heterogeneous imaging modalities. To further enrich semantic representation, text descriptions have recently been incorporated as an additional modality. However, recent vision-language approaches often treat …
- IMPLICIT-Bench: Measuring Implicit Bias in Text-to-Image Models under Neutral Prompts
Yue Dai, Ziyang Liu, Marc Cheong, Caren Han · 22. September 2026
Text-to-image (T2I) models are typically evaluated for bias using slot-based templates such as ``a photo of a [profession]''. Such templates probe only \emph{explicit} demographic attributes (e.g., gender, skin tone) in isolation. They overlook a broader \emph{implicit} bias that arises in natural p…
- InterHier: Learning Interconnected Hierarchical Semantics for Open-Vocabulary Object Detection
Yeong-Jin Kim, Ho-Joong Kim, Seong-Whan Lee · 22. September 2026
In this paper, we investigate the limitations of fixed, hand-crafted connectors in hierarchical semantic representations for open-vocabulary object detection. Existing methods establish semantic relationships between base categories and unseen novel categories by placing a fixed connector between ad…
- RA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes
Abhirama Subramanyam Penamakuri, Shreya Shukla, Anand Mishra · 15. September 2026
Fashion Image Captioning (FIC) plays a vital role in enhancing user experience and product search in e-commerce platforms. Unlike natural scene image captioning, FIC requires fine-grained visual reasoning and knowledge of domain-specific terminology to capture subtle attributes such as neckline and …
- PinDCO: Whole-Page Aware Dynamic Creative Optimization at Scale
Yu Hao, Yuchun Li, Peimeng Sui, Meilin Liu, Tianyuan Cui, Hao Li, Zicong Zhou, Akanksha Baid · 14. September 2026
Recent advances in generative AI have substantially accelerated the creation of high-quality ad creatives, dramatically expanding the number of candidate variants per campaign. This shift increases the need for scalable dynamic creative optimization (DCO) systems that can match creatives to the most…
- Generative Retrieval for Unsupervised Text-Based Person Search
Mang Ye, Yucheng Ji, Yang Bai, Min Cao, Siyuan Chai, Bo Du, Min Zhang · 14. September 2026
Text-based person search (TBPS) aims to retrieve images of a target person from a large image gallery based on a given natural language description. Most existing methods rely on supervised learning with manually annotated image-text pairs. In this paper, we explore unsupervised TBPS, with only unla…
- RiVaT-Fuse: Reliability-Calibrated Variational Tensor Fusion for Multimodal Prediction under Modality Uncertainty
Yingfan Xu, Tieming Liu, Ye Liang, Taiping Liu · 11. September 2026
Image-metadata prediction requires fusing heterogeneous evidence whose reliability can vary across samples and latent factors. Existing representation-level fusion methods typically choose an aggregation architecture, such as concatenation, gating, conditional modulation, or attention, without expli…
- Test-time Prompt Refinement for Text-to-Image Models
Mohammad Abdul Hafeez Khan, Yash Jain, Siddhartha Bhattacharyya, Vibhav Vineet · 10. September 2026
Text-to-image (T2I) generation models have made significant strides but still struggle with prompt sensitivity: even minor changes in prompt wording can yield inconsistent or inaccurate outputs. To address this challenge, we introduce a closed-loop, test-time prompt refinement framework that require…
- Bigger Text Encoders Can Hurt CLIP Zero-Shot Performance
Samir Char, Carles Domingo-Enrich, Randall Balestriero · 9. September 2026
Contrastive Language-Image Pretraining (CLIP) is a building block of many machine learning applications. Scaling laws have guided resource allocation for large-scale training, yet prior work treats total CLIP model size as a single variable, without exploring how the capacity split between encoders …
- VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognition
Jiangang Zhu, Zheng Wang, Bin Zhu, Yi-Ping Phoebe Chen, Jingjing Chen · 7. September 2026
Multi-expert models have become the dominant paradigm for long-tailed learning, largely attributed to their presumed ability to benefit from expert diversity. However, we revisit this central assumption and reveal that diversity induced by logit adjustment or explicit regularizers does not guarantee…
- Whose record is this? Diagnosing and authorizing record use in personalized multimodal models
Xinyu Mao, Junsi Li, Chenyang Liu, Haoji Zhang, Ming Sun · 7. September 2026
Contextualized visual personalization can retrieve a true record yet apply it to the wrong visual subject. We formalize when a record may condition an answer as \emph{record authorization}: subject presence ($P$), record-edge validity ($E$), and answer support ($S$) must all hold. We call violations…
- FoRIS: Progressive Foreground Refinement for Training-Free In-Context Segmentation
Ming Hu, Jianfu Yin, Mingyu Dou, Miaomiao Zhang, Yao Wang, Cong Hu, Bingliang Hu, Quan Wang · 4. September 2026
In-Context Segmentation (ICS) aims to precisely segment arbitrary semantic concepts, such as objects or parts, given one or a few annotated visual exemplars. In this paper, we revisit ICS from a more classical segmentation perspective, viewing it as a coarse-to-fine progressive refinement process. R…
- DICS: Exploring Data Intrinsic Consistency for Visual Instruction Selection
Yuyang Hong, Jinhui Guo, Jiaqi Gu, Lubin Fan, Ruixiang Wang, Kun Ding, Yue Wu, Shiming Xiang, Jieping Ye · 2. September 2026
Visual instruction tuning is crucial for advancing the vision-language alignment and instruction-following capabilities of Vision-Language Models (VLMs). However, identifying optimal subsets under a fixed ratio constraint from rapidly expanding datasets remains a significant bottleneck. While existi…
Weitere Unterthemen aus Bildverarbeitung und Mustererkennung
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Multimodal Machine Learning Applications8.069 Papiere / 12 Monate+191 %
- Generative Adversarial Networks and Image Synthesis4.992 Papiere / 12 Monate+39 %
- Advanced Neural Network Applications2.354 Papiere / 12 Monate+48 %
- Advanced Vision and Imaging841 Papiere / 12 Monate+78 %
- Human Pose and Action Recognition836 Papiere / 12 Monate+457 %
- Face recognition and analysis482 Papiere / 12 Monate+88 %
