Social Sciences › Decision Sciences › Management Science and Operations Research
Data Quality and Management
254 papers indexed
The works gathered under this topic explore methods to enhance the quality and management of data in the field of artificial intelligence, particularly when structured in tabular form. They address techniques such as Tabular Foundation Models, diffusion transformers, or Bayesian approaches to generate, evaluate, or adapt data while preserving its relational consistency, fairness, or operational utility. The focus is on challenges such as metadata reconstruction, data discovery from raw values, or optimizing their exploitation by autonomous agents or distributed systems.
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States42% · 54 papers
- China31% · 40 papers
- Germany12% · 16 papers
- United Kingdom7.8% · 10 papers
- France5.4% · 7 papers
- Hong Kong SAR China5.4% · 7 papers
- Canada4.7% · 6 papers
- Italy3.9% · 5 papers
Across 129 papers on this subject with at least one lab located. 33 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- JoinGR: Learning to Traverse Join Graphs for Table Retrieval
Sandipan De, Abhijit Chakraborty, Sambaran Bandyopadhyay, Vivek Gupta · 2 October 2026
Retrieving the right tables is a prerequisite for Text-to-SQL over realistic databases. Dense table retrievers rank schema elements independently, but this ignores a key source of evidence: some required tables are not mentioned in the question and become identifiable only through their join relatio…
- TabJoinBench: A Benchmark for Joinable Table Discovery
Sandipan De, Jin Wang, Vivek Gupta · 2 October 2026
Join discovery aims to identify tables from large data repositories that can augment a query table with complementary information, enabling downstream tasks such as data exploration, feature engineering, and business intelligence. Although numerous join discovery methods have been proposed, existing…
- BudgetSchemaBench: A Budget-Swept Diagnostic for Schema Context in Text-to-SQL
Chen Shen · 2 October 2026
Data agents over structured sources must fit database schema into the model's context window. Large catalogs can span many databases and thousands of columns, so cost constraints may require choosing between table coverage and serialization detail well before the context window is full. We introduce…
- Enterprise Representation Simplification (ERS): Reducing Representational Complexity for Enterprise AI
Terry Dorsey, Kevin Huggins · 2 October 2026
Enterprise information is represented through artifacts shaped by applications, projects, technologies, organizational boundaries, and local requirements. These structures accumulate over time, creating representational complexity that must be maintained by the enterprise and interpreted by informat…
- Billiger.de Products: A Bilingual Entity Matching Benchmark
Aaron Steiner, Ksenia Elagin, Ralph Peeters, Johannes Knopp, Christian Bizer · 1 October 2026
Existing product matching benchmarks primarily contain English-language product data and are often dominated by a single product category, such as electronics. This paper introduces Billiger.de Products, a bilingual German and English entity matching benchmark covering thirteen consumer product cate…
- Adapt Semantics, Not Structure: Few-Instance Schema Calibration for Scientific PDF Extraction
Zixiao Dong, Wei Yang, Zihao Liu, Chenshu Li, Longzhang Liu, Tao Tan, Hong Xie · 30 September 2026
A well-designed extraction schema is not necessarily ready for reliable LLM execution. When only limited verified extractions are available, manually tuning hundreds of field definitions through trial and error is costly. We frame this problem as few-instance schema calibration: adapting the operati…
- LoopICL: Looping a single transformer block to solve tabular tasks
Amir Rezaei Balef, Katharina Eggensperger · 30 September 2026
Tabular foundation models using in-context learning have recently surpassed gradient-boosted trees on predictive tabular tasks. However, recent mechanistic insights suggest that parameters in these models are largely redundant. We introduce LoopICL, a looped transformer whose core design decouples p…
- Support-Set Target Leakage in Relational Foundation Models during In-Context Learning: Model Dependence and Evaluation Reliability
Roshan Reddy Upendra, Alexandre Dorais, Joe Meyer, Andrew Pouret, Anastasios Lambrianos Stappas, Dinesh Katupputhur Ramprasath, Viswanath Ganapathy, Tom Palczewski, Minghua Li · 30 September 2026
Relational in-context learning (ICL) uses labeled support examples and their linked relational context to predict labels for new queries. This creates a failure mode when target-derived features are present in the support context but unavailable for the query. We study this setting as support-set ta…
- Support-Set Target Leakage in Relational Foundation Models during In-Context Learning: Impact, Detection, and Mitigation
Roshan Reddy Upendra, Alexandre Dorais, Joe Meyer, Andrew Pouret, Anastasios Lambrianos Stappas, Dinesh Katupputhur Ramprasath, Tom Palczewski, Minghua Li · 30 September 2026
Relational in-context learning (ICL) conditions predictions on the labeled support examples and their linked tables, creating a failure mode when the support set contains target-derived features that are unavailable for the query. We formulate this problem as support-set target leakage, distinct fro…
- InterTab: Interleaved Visual-Structure Alignment for Multi-Modal Table Reasoning
Hanqian Li, Sirui Huang, Chen Ling, Jungang Li, Yu Huang, Kening Zheng, Yonghua Hei, Xiangrong He, Shiyi Wang, Pengcheng Zhu, Dongnan Liu, Wei Zhou, Linjian Mo, Nai Ding, Xuming Hu · 29 September 2026
Table images preserve structural information that are often lost in text serialization, and reasoning over them requires locating relevant rows, columns, and cells step by step. Current multimodal large language models (MLLMs) encode the whole image once before reasoning, so they cannot pick up row-…
- TableSeek: Structure-Preserving Agentic Evidence Seeking over Heterogeneous Table Corpora
Jiaming Tian, Liyao Li, Wentao Ye, Haobo Wang, Lihua Yu, Zujie Ren, Gang Chen, Junbo Zhao · 29 September 2026
Open-domain table retrieval seeks tables that contain sufficient evidence for answering a question or verifying a claim. Yet semantic relevance is often misleading: topically similar tables may lack the required facts, while answer-bearing evidence is often confined to a few cells whose meaning depe…
- Benchmarking Attention for Tabular Foundation Models
Maximilian Schambach, Clemens Biehl, Sam Thelin · 28 September 2026
Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention involves longer sequences while column attent…
- Support-Compiled Feature Folding: More Evidence at Lower Memory Across Tabular Foundation Models
Tian Zhou, Beverly Jin, Xue Wang, Linxiao Yang, Wenwei Wang, Bingqing Peng, Mengni Ye, Jinjie Gu, Liang Sun · 24 September 2026
Tabular foundation models face a feature-side scaling dilemma: full-width pairwise mixing grows quadratically with the number of columns, whereas feature selection saves memory by discarding evidence. We introduce Support-Compiled Feature Folding (SCFF), a training-free inference framework that reso…
- What Do Tabular Foundation Models Compute In Context? In-Situ Representation Refinement through Attention-Gated Updates
Tian Zhou, Beverly Jin, Linxiao Yang, Xue Wang, Wenwei Wang, Bingqing Peng, Mengni Ye, Jinjie Gu, Liang Sun · 24 September 2026
What reusable computation should a tabular foundation model learn when every table defines a new supervised task? We develop in-situ representation refinement: support labels guide updates to the episode's representations, and these updates transfer to unlabeled queries without changing model parame…
- Automated Extraction of Records of Processing Activities (RoPA) Using Hybrid RAG and Locally Deployed Large Language Models
To Duy Hinh, Nguyen Le Quoc Anh, Phan Van Tri, Khuong Nguyen-An · 24 September 2026
Vietnam's Personal Data Protection Law (Law No. 91/2025/QH15) and Decree No. 356/2025/ND-CP, effective January 1, 2026, require organizations to establish and maintain Records of Processing Activities (RoPA). Manual RoPA preparation is labor-intensive, while cloud-hosted large language models (LLMs)…
- Discovery-Driven Integration of Disjoint Tables via Text
Md Ataur Rahman, Dimitris Sacharidis, Oscar Romero, Sergi Nadal · 23 September 2026
Integrating heterogeneous datasets within data lakes is a critical challenge, particularly for semantically related tables that lack the explicit attributes needed to be joined. We study Discovery-Driven Integration, where the relevant sources and their missing relational structure must be discovere…
- A JEPA Recipe for Tabular Foundation Models
Mingyu Jeon, Suwan Cho, Jae Young Suh · 23 September 2026
Tabular foundation models learn to predict cell values in context, whereas world-model self-supervision asks for prediction in representation space (LeCun, 2022; Assran et al., 2023). On a tabular foundation-model prior, the latent term of a joint-embedding predictive architecture (JEPA) collapsed i…
- Learned Enterprise Data Comprehension: Compression and Routing for Data Agents
Ethan Torres, Eric Mills · 23 September 2026
Structured-data agents in enterprise settings must reason over complex data environments whose relevant evidence is distributed across schemas, relationships, policies, and recurring business roles. Modern agentic systems often address this burden through reusable markdown-style memory or skill file…
- Causilo Technical Report
Minyong Cho, Minho Jeong, Dooho Lee, Jinmo Lee, Jaemin Yoo · 22 September 2026
We introduce Causilo, a tabular foundation model (TFM) that combines frontier predictive performance with exceptionally fast inference. On TabArena, Causilo achieves 1785.4 Elo, at a median inference time of 0.10 seconds per 1K test samples. It outperforms TabPFN-3.5-Fast with 31.6% less inference t…
- From Tables to Quantified Statements: Evaluating LLM Inference Generation through Executable Verification
Mai Mohamed Eida, Gunjan Anand, Ayush Singh, Aleksandre Maskharashvili · 22 September 2026
LLMs can generate fluent descriptions from tables, but their outputs may remain logically unsupported by the structured data. We introduce STAT-TO-TEXT, a controlled task in which LLMs generate quantified natural language inferences from statistical tables using quantified constructions such as all,…
- Beyond the Stitching Assumption: A Unified Framework for Multimodal Synthetic Data Evaluation via Semantic Quantization
Yefeng Yuan, Zhan Shi, Liang Cheng, Yuhong Liu · 22 September 2026
Multimodal synthetic datasets combine structured attributes with free text, but are often evaluated separately. Such metrics can remain high after tabular--text pairings are disrupted. We present a projection-based evaluator for tabular--text synthetic data. A fixed sentence encoder maps text to emb…
- PrismAlign: Prior-Steered Multi-View VLM Alignment for Hallucination-Robust Table OCR
Guangyi Liu, Qianjun Huang, Boyu Hou · 21 September 2026
Table extraction suffers from frequent structural errors and semantic hallucinations. We propose PrismAlign, a multi-VLM framework aligning diverse visual perspectives to resolve ambiguity. It integrates priors of table logic to assess output plausibility, decoupling structural alignment from cell c…
- BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence
Chuxuan Hu, Yeye He, Penny Zhou, Wee Hyong Tok, Daniel Kang, Surajit Chaudhuri · 21 September 2026
Business intelligence (BI) is a cornerstone of enterprise decision-making and is widely used by enterprise users in software such as Power BI and Tableau. In traditional BI workflows, users need to prepare data by (1) identifying relevant tables, (2) performing data transformations, and (3) building…
- A Policy Profile for Croissant: Refusal as a Property of the Dataset
Alexander Chernov · 18 September 2026
Croissant is the de facto machine-readable descriptor for ML datasets: JSON-LD over schema.org. Since version 1.1 it also carries data use conditions, recommending DUO and ODRL for them. What no version specifies is how any of them is evaluated: no decision procedure, no bound on evaluation cost, no…
- Generating a Consistent Enterprise: Synthesis and Reference-Free Evaluation of Multi-System Business Data
Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary, Omer Niv · 18 September 2026
Synthetic relational data is normally produced by a model trained on a real dataset, and its quality is measured as the distance to that dataset. This paper describes a generator that has no real dataset at either end. Given an industry, a company size, a business model, a set of business applicatio…
Other topics in Management science and operations research
The topics the OpenAlex classification attaches to the same theme, most active first.
- Advanced Bandit Algorithms Research697 papers / 12 months+31%
- Stock Market Forecasting Methods391 papers / 12 months+420%
- Forecasting Techniques and Applications300 papers / 12 months+700%
- Auction Theory and Applications98 papers / 12 months+100%
- Risk and Portfolio Optimization98 papers / 12 months+233%
- Game Theory and Applications75 papers / 12 months+200%
