Social Sciences › Decision Sciences › Management Science and Operations Research
Data Quality and Management
254 papiers indexés
Les travaux rassemblés sous ce sujet explorent les méthodes pour améliorer la qualité et la gestion des données dans le domaine de l’intelligence artificielle, en particulier lorsqu’elles sont structurées sous forme de tableaux. Ils abordent des techniques comme les Tabular Foundation Models, les diffusion transformers ou les approches bayésiennes pour générer, évaluer ou adapter des données tout en préservant leur cohérence relationnelle, leur équité ou leur utilité opérationnelle. L’accent est mis sur des défis tels que la reconstruction de métadonnées, la découverte de données à partir de valeurs brutes, ou encore l’optimisation de leur exploitation par des agents autonomes ou des systèmes distribués.
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel - 12 derniers mois
Pays des laboratoires
- États-Unis42 % · 54 articles
- Chine31 % · 40 articles
- Allemagne12 % · 16 articles
- Royaume-Uni7,8 % · 10 articles
- France5,4 % · 7 articles
- R.A.S. chinoise de Hong Kong5,4 % · 7 articles
- Canada4,7 % · 6 articles
- Italie3,9 % · 5 articles
Sur 129 articles de ce sujet dont au moins un laboratoire est situé. 33 pays représentés.
Il s'agit du pays du laboratoire, jamais de la nationalité des personnes. Un article signé depuis plusieurs pays compte pour chacun d'eux, les parts dépassent donc 100 % au total. La couverture est partielle et le manque n'est pas aléatoire : un chercheur dont l'institution est inconnue publie en général peu, ce qui sur-représente les laboratoires établis.
Derniers papiers
- Billiger.de Products: A Bilingual Entity Matching Benchmark
Aaron Steiner, Ksenia Elagin, Ralph Peeters, Johannes Knopp, Christian Bizer · 1 octobre 2026
Existing product matching benchmarks primarily contain English-language product data and are often dominated by a single product category, such as electronics. This paper introduces Billiger.de Products, a bilingual German and English entity matching benchmark covering thirteen consumer product cate…
- Adapt Semantics, Not Structure: Few-Instance Schema Calibration for Scientific PDF Extraction
Zixiao Dong, Wei Yang, Zihao Liu, Chenshu Li, Longzhang Liu, Tao Tan, Hong Xie · 30 septembre 2026
A well-designed extraction schema is not necessarily ready for reliable LLM execution. When only limited verified extractions are available, manually tuning hundreds of field definitions through trial and error is costly. We frame this problem as few-instance schema calibration: adapting the operati…
- LoopICL: Looping a single transformer block to solve tabular tasks
Amir Rezaei Balef, Katharina Eggensperger · 30 septembre 2026
Tabular foundation models using in-context learning have recently surpassed gradient-boosted trees on predictive tabular tasks. However, recent mechanistic insights suggest that parameters in these models are largely redundant. We introduce LoopICL, a looped transformer whose core design decouples p…
- Support-Set Target Leakage in Relational Foundation Models during In-Context Learning: Model Dependence and Evaluation Reliability
Roshan Reddy Upendra, Alexandre Dorais, Joe Meyer, Andrew Pouret, Anastasios Lambrianos Stappas, Dinesh Katupputhur Ramprasath, Viswanath Ganapathy, Tom Palczewski, Minghua Li · 30 septembre 2026
Relational in-context learning (ICL) uses labeled support examples and their linked relational context to predict labels for new queries. This creates a failure mode when target-derived features are present in the support context but unavailable for the query. We study this setting as support-set ta…
- Support-Set Target Leakage in Relational Foundation Models during In-Context Learning: Impact, Detection, and Mitigation
Roshan Reddy Upendra, Alexandre Dorais, Joe Meyer, Andrew Pouret, Anastasios Lambrianos Stappas, Dinesh Katupputhur Ramprasath, Tom Palczewski, Minghua Li · 30 septembre 2026
Relational in-context learning (ICL) conditions predictions on the labeled support examples and their linked tables, creating a failure mode when the support set contains target-derived features that are unavailable for the query. We formulate this problem as support-set target leakage, distinct fro…
- InterTab: Interleaved Visual-Structure Alignment for Multi-Modal Table Reasoning
Hanqian Li, Sirui Huang, Chen Ling, Jungang Li, Yu Huang, Kening Zheng, Yonghua Hei, Xiangrong He, Shiyi Wang, Pengcheng Zhu, Dongnan Liu, Wei Zhou, Linjian Mo, Nai Ding, Xuming Hu · 29 septembre 2026
Table images preserve structural information that are often lost in text serialization, and reasoning over them requires locating relevant rows, columns, and cells step by step. Current multimodal large language models (MLLMs) encode the whole image once before reasoning, so they cannot pick up row-…
- TableSeek: Structure-Preserving Agentic Evidence Seeking over Heterogeneous Table Corpora
Jiaming Tian, Liyao Li, Wentao Ye, Haobo Wang, Lihua Yu, Zujie Ren, Gang Chen, Junbo Zhao · 29 septembre 2026
Open-domain table retrieval seeks tables that contain sufficient evidence for answering a question or verifying a claim. Yet semantic relevance is often misleading: topically similar tables may lack the required facts, while answer-bearing evidence is often confined to a few cells whose meaning depe…
- Benchmarking Attention for Tabular Foundation Models
Maximilian Schambach, Clemens Biehl, Sam Thelin · 28 septembre 2026
Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention involves longer sequences while column attent…
- Support-Compiled Feature Folding: More Evidence at Lower Memory Across Tabular Foundation Models
Tian Zhou, Beverly Jin, Xue Wang, Linxiao Yang, Wenwei Wang, Bingqing Peng, Mengni Ye, Jinjie Gu, Liang Sun · 24 septembre 2026
Tabular foundation models face a feature-side scaling dilemma: full-width pairwise mixing grows quadratically with the number of columns, whereas feature selection saves memory by discarding evidence. We introduce Support-Compiled Feature Folding (SCFF), a training-free inference framework that reso…
- What Do Tabular Foundation Models Compute In Context? In-Situ Representation Refinement through Attention-Gated Updates
Tian Zhou, Beverly Jin, Linxiao Yang, Xue Wang, Wenwei Wang, Bingqing Peng, Mengni Ye, Jinjie Gu, Liang Sun · 24 septembre 2026
What reusable computation should a tabular foundation model learn when every table defines a new supervised task? We develop in-situ representation refinement: support labels guide updates to the episode's representations, and these updates transfer to unlabeled queries without changing model parame…
- Automated Extraction of Records of Processing Activities (RoPA) Using Hybrid RAG and Locally Deployed Large Language Models
To Duy Hinh, Nguyen Le Quoc Anh, Phan Van Tri, Khuong Nguyen-An · 24 septembre 2026
Vietnam's Personal Data Protection Law (Law No. 91/2025/QH15) and Decree No. 356/2025/ND-CP, effective January 1, 2026, require organizations to establish and maintain Records of Processing Activities (RoPA). Manual RoPA preparation is labor-intensive, while cloud-hosted large language models (LLMs)…
- Discovery-Driven Integration of Disjoint Tables via Text
Md Ataur Rahman, Dimitris Sacharidis, Oscar Romero, Sergi Nadal · 23 septembre 2026
Integrating heterogeneous datasets within data lakes is a critical challenge, particularly for semantically related tables that lack the explicit attributes needed to be joined. We study Discovery-Driven Integration, where the relevant sources and their missing relational structure must be discovere…
- A JEPA Recipe for Tabular Foundation Models
Mingyu Jeon, Suwan Cho, Jae Young Suh · 23 septembre 2026
Tabular foundation models learn to predict cell values in context, whereas world-model self-supervision asks for prediction in representation space (LeCun, 2022; Assran et al., 2023). On a tabular foundation-model prior, the latent term of a joint-embedding predictive architecture (JEPA) collapsed i…
- Learned Enterprise Data Comprehension: Compression and Routing for Data Agents
Ethan Torres, Eric Mills · 23 septembre 2026
Structured-data agents in enterprise settings must reason over complex data environments whose relevant evidence is distributed across schemas, relationships, policies, and recurring business roles. Modern agentic systems often address this burden through reusable markdown-style memory or skill file…
- Causilo Technical Report
Minyong Cho, Minho Jeong, Dooho Lee, Jinmo Lee, Jaemin Yoo · 22 septembre 2026
We introduce Causilo, a tabular foundation model (TFM) that combines frontier predictive performance with exceptionally fast inference. On TabArena, Causilo achieves 1785.4 Elo, at a median inference time of 0.10 seconds per 1K test samples. It outperforms TabPFN-3.5-Fast with 31.6% less inference t…
- From Tables to Quantified Statements: Evaluating LLM Inference Generation through Executable Verification
Mai Mohamed Eida, Gunjan Anand, Ayush Singh, Aleksandre Maskharashvili · 22 septembre 2026
LLMs can generate fluent descriptions from tables, but their outputs may remain logically unsupported by the structured data. We introduce STAT-TO-TEXT, a controlled task in which LLMs generate quantified natural language inferences from statistical tables using quantified constructions such as all,…
- Beyond the Stitching Assumption: A Unified Framework for Multimodal Synthetic Data Evaluation via Semantic Quantization
Yefeng Yuan, Zhan Shi, Liang Cheng, Yuhong Liu · 22 septembre 2026
Multimodal synthetic datasets combine structured attributes with free text, but are often evaluated separately. Such metrics can remain high after tabular--text pairings are disrupted. We present a projection-based evaluator for tabular--text synthetic data. A fixed sentence encoder maps text to emb…
- PrismAlign: Prior-Steered Multi-View VLM Alignment for Hallucination-Robust Table OCR
Guangyi Liu, Qianjun Huang, Boyu Hou · 21 septembre 2026
Table extraction suffers from frequent structural errors and semantic hallucinations. We propose PrismAlign, a multi-VLM framework aligning diverse visual perspectives to resolve ambiguity. It integrates priors of table logic to assess output plausibility, decoupling structural alignment from cell c…
- BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence
Chuxuan Hu, Yeye He, Penny Zhou, Wee Hyong Tok, Daniel Kang, Surajit Chaudhuri · 21 septembre 2026
Business intelligence (BI) is a cornerstone of enterprise decision-making and is widely used by enterprise users in software such as Power BI and Tableau. In traditional BI workflows, users need to prepare data by (1) identifying relevant tables, (2) performing data transformations, and (3) building…
- A Policy Profile for Croissant: Refusal as a Property of the Dataset
Alexander Chernov · 18 septembre 2026
Croissant is the de facto machine-readable descriptor for ML datasets: JSON-LD over schema.org. Since version 1.1 it also carries data use conditions, recommending DUO and ODRL for them. What no version specifies is how any of them is evaluated: no decision procedure, no bound on evaluation cost, no…
- Generating a Consistent Enterprise: Synthesis and Reference-Free Evaluation of Multi-System Business Data
Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary, Omer Niv · 18 septembre 2026
Synthetic relational data is normally produced by a model trained on a real dataset, and its quality is measured as the distance to that dataset. This paper describes a generator that has no real dataset at either end. Given an industry, a company size, a business model, a set of business applicatio…
- Is It Still Worth Training a Classical Model in the Era of LLMs? A Crossover Benchmark on Tabular Data
Kaihua Ding · 18 septembre 2026
Large language models can label a tabular row from a plain-English description with no training - a capability now shipping in mainstream spreadsheet tools such as Microsoft Copilot in Excel and Anthropic's Claude for Excel - raising a practical question for the many business prediction problems whe…
- Semantic Layer Induction from Raw Telemetry via Hierarchical LLM and RAG Abstraction
Yuanzhe Jia, Ali Anaissi · 18 septembre 2026
Modern applications generate massive volumes of raw telemetry data, but translating those noisy, heterogeneous event streams into actionable business insights remains a fundamental challenge. Data engineers and analysts expend substantial effort reconciling semantic discrepancies, hand-crafting pars…
- From "Who Is This User?" to "What Does This Purchase Mean?": A Deployed Pipeline for Semantic User Profiling at Bank Scale
Ryota Mitsuhashi, Tetsuro Morimura, Hirotake Ito · 18 septembre 2026
Per-user LLM inference on transaction histories binds the inference budget linearly to user count, which becomes prohibitive at applied scale. We re-cast attribute inference from per-user to per-transaction-pattern. The pipeline runs in three phases: Resolve abstracts item names with optional web gr…
- MetaRTL: Meta-path Attention Enhanced Relational Table Learning
Ken Zhong, Weichen Li, Zheng Wang · 18 septembre 2026
Relational table learning has gained increasing attention with the widespread use of relational databases. Existing methods typically rely on deep GNN or HGNN stacks, leading to high computational costs and limited performance on large real-world databases. We propose MetaRTL, a two-stage framework …
Autres sujets du thème Recherche opérationnelle et sciences de gestion
Les sujets rattachés au même thème par la classification OpenAlex, les plus actifs d'abord.
- Advanced Bandit Algorithms Research697 papiers / 12 mois+31 %
- Stock Market Forecasting Methods391 papiers / 12 mois+420 %
- Forecasting Techniques and Applications300 papiers / 12 mois+700 %
- Auction Theory and Applications98 papiers / 12 mois+100 %
- Risk and Portfolio Optimization98 papiers / 12 mois+233 %
- Game Theory and Applications75 papiers / 12 mois+200 %
