Physical Sciences › Mathematics › Statistics and Probability
Statistical Methods and Inference
195 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume — last 12 months
Latest papers
- Predicting kernel regression learning curves from only raw data statistics
Dhruva Karkada, Joseph Turnbull, Yuxi Liu, James B. Simon · 12 March 2026
We study kernel regression with common rotation-invariant kernels on real datasets including CIFAR-5m, SVHN, and ImageNet. We give a theoretical framework that predicts learning curves (test risk vs. sample size) from only two measurements: the empirical data covariance matrix and an empirical polyn…
- Kernel Debiased Plug-in Estimation based on the Universal Least Favorable Submodel
Haiyi Chen, Yang Liu, Ivana Malenica · 11 March 2026
We propose ULFS-KDPE, a kernel debiased plug-in estimator based on the universal least favorable submodel, for estimating pathwise differentiable parameters in nonparametric models. The method constructs a data-adaptive debiasing flow in a reproducing kernel Hilbert space (RKHS), producing a plug-in…
- Valid Feature-Level Inference for Tabular Foundation Models via the Conditional Randomization Test
Mohamed Salem · 10 March 2026
Modern machine learning models are highly expressive but notoriously difficult to analyze statistically. In particular, while black-box predictors can achieve strong empirical performance, they rarely provide valid hypothesis tests or p-values for assessing whether individual features contain inform…
- Overlap-Adaptive Regularization for Conditional Average Treatment Effect Estimation
Valentyn Melnychuk, Dennis Frauen, Jonas Schweisthal, Stefan Feuerriegel · 10 March 2026
The conditional average treatment effect (CATE) is widely used in personalized medicine to inform therapeutic decisions. However, state-of-the-art methods for CATE estimation (so-called meta-learners) often perform poorly in the presence of low overlap. In this work, we introduce a new approach to t…
- SPPCSO: Adaptive Penalized Estimation Method for High-Dimensional Correlated Data
Ying Hu, Hu Yang · 9 March 2026
With the rise of high-dimensional correlated data, multicollinearity poses a significant challenge to model stability, often leading to unstable estimation and reduced predictive accuracy. This work proposes the Single-Parametric Principal Component Selection Operator (SPPCSO), an innovative penaliz…
- How important are the genes to explain the outcome - the asymmetric Shapley value as an honest importance metric for high-dimensional features
Mark A. van de Wiel, Jeroen Goedhart, Martin Jullum, Kjersti Aas · 6 March 2026
In clinical prediction settings the importance of a high-dimensional feature like genomics is often assessed by evaluating the change in predictive performance when adding it to a set of traditional clinical variables. This approach is questionable, because it does not account for collinearity nor k…
- Impossibility of Depth Reduction in Explainable Clustering
Chengyuan Deng, Surya Teja Gavva, Karthik C. S., Parth Patel, Adarsh Srinivasan · 3 March 2026
Over the last few years Explainable Clustering has gathered a lot of attention. Dasgupta et al. [ICML'20] initiated the study of explainable $k$-means and $k$-median clustering problems where the explanation is captured by a threshold decision tree which partitions the space at each node using axis …
- A Large-Scale Neutral Comparison Study of Survival Models on Low-Dimensional Data
Lukas Burk, John Zobolas, Bernd Bischl, Andreas Bender, Marvin N. Wright, Raphael Sonabend · 3 March 2026
This work presents the first large-scale neutral benchmark experiment focused on single-event, right-censored, low-dimensional survival data. Benchmark experiments are essential in methodological research to scientifically compare new and existing model classes through proper empirical evaluation. E…
- Adaptive Estimation and Inference in Conditional Moment Models via the Discrepancy Principle
Jiyuan Tan, Vasilis Syrgkanis · 3 March 2026
We study adaptive estimation and inference in ill-posed linear inverse problems defined by conditional moment restrictions. Existing regularized estimators such as Regularized DeepIV (RDIV) require prior knowledge of the smoothness of the nuisance function, typically encoded by a beta source conditi…
- Beyond False Discovery Rate: A Stepdown Group SLOPE Approach for Grouped Variable Selection
Xuelin Zhang, Jingxuan Liang, Xinyue Liu, Hong Chen, Biqin Song · 3 March 2026
High-dimensional feature selection is routinely required to balance statistical power with strict control of multiple-error metrics such as the k-Family-Wise Error Rate (k-FWER) and the False Discovery Proportion (FDP), yet some existing frameworks are confined to the narrower goal of controlling th…
- Throwing Vines at the Wall: Structure Learning via Random Search
Thibault Vatter, Thomas Nagler · 27 February 2026
Vine copulas offer flexible multivariate dependence modeling and have become widely used in machine learning. Yet, structure learning remains a key challenge. Early heuristics, such as Dissmann's greedy algorithm, are still considered the gold standard but are often suboptimal. We propose random sea…
- Kernel Integrated $R^2$: A Measure of Dependence
Pouya Roudaki, Shakeel Gavioli-Akilagun, Florian Kalinke, Mona Azadkia, Zolt\'an Szab\'o · 27 February 2026
We introduce kernel integrated $R^2$, a new measure of statistical dependence that combines the local normalization principle of the recently introduced integrated $R^2$ with the flexibility of reproducing kernel Hilbert spaces (RKHSs). The proposed measure extends integrated $R^2$ from scalar respo…
- Learning and Naming Subgroups with Exceptional Survival Characteristics
Mhd Jawad Al Rahwanji, Sascha Xu, Nils Philipp Walter, Jilles Vreeken · 26 February 2026
In many applications, it is important to identify subpopulations that survive longer or shorter than the rest of the population. In medicine, for example, it allows determining which patients benefit from treatment, and in predictive maintenance, which components are more likely to fail. Existing me…
- Efficient Inference after Directionally Stable Adaptive Experiments
Zikai Shen, Houssam Zenati, Nathan Kallus, Arthur Gretton, Koulik Khamaru, Aur\'elien Bibaut · 26 February 2026
We study inference on scalar-valued pathwise differentiable targets after adaptive data collection, such as a bandit algorithm. We introduce a novel target-specific condition, directional stability, which is strictly weaker than previously imposed target-agnostic stability conditions. Under directio…
- Overparameterized Multiple Linear Regression as Hyper-Curve Fitting
E. Atza, N. Budko · 26 February 2026
This work demonstrates that applying a fixed-effect multiple linear regression (MLR) model to an overparameterized dataset is mathematically equivalent to fitting a hyper-curve parameterized by a single scalar. This reformulation shifts the focus from global coefficients to individual predictors, al…
- Conformal Risk Control for Non-Monotonic Losses
Anastasios N. Angelopoulos · 24 February 2026
Conformal risk control is an extension of conformal prediction for controlling risk functions beyond miscoverage. The original algorithm controls the expected value of a loss that is monotonic in a one-dimensional parameter. Here, we present risk control guarantees for generic algorithms applied to …
- MIBoost: A Gradient Boosting Algorithm for Variable Selection After Multiple Imputation
Robert Kuchen · 24 February 2026
Statistical learning methods for automated variable selection, such as LASSO, elastic nets, or gradient boosting, have become increasingly popular tools for building powerful prediction models. Yet, in practice, analyses are often complicated by missing data. The most widely used approach to address…
- On the Generalization and Robustness in Conditional Value-at-Risk
Dinesh Karthik Mulumudi, Piyushi Manupriya, Gholamali Aminian, Anant Raj · 23 February 2026
Conditional Value-at-Risk (CVaR) is a widely used risk-sensitive objective for learning under rare but high-impact losses, yet its statistical behavior under heavy-tailed data remains poorly understood. Unlike expectation-based risk, CVaR depends on an endogenous, data-dependent quantile, which coup…
- Boosting methods for interval-censored data with regression and classification
Yuan Bian, Grace Y. Yi, Wenqing He · 19 February 2026
Boosting has garnered significant interest across both machine learning and statistical communities. Traditional boosting algorithms, designed for fully observed random samples, often struggle with real-world problems, particularly with interval-censored data. This type of data is common in survival…
- Testing For Distribution Shifts with Conditional Conformal Test Martingales
Shalev Shaer, Yarin Bar, Drew Prinster, Yaniv Romano · 17 February 2026
We propose a sequential test for detecting arbitrary distribution shifts that allows conformal test martingales (CTMs) to work under a fixed, reference-conditional setting. Existing CTM detectors construct test martingales by continually growing a reference set with each incoming sample, using it to…
- Linear Regression with Unknown Truncation Beyond Gaussian Features
Alexandros Kouridakis, Anay Mehrotra, Alkis Kalavasis, Constantine Caramanis · 16 February 2026
In truncated linear regression, samples $(x,y)$ are shown only when the outcome $y$ falls inside a certain survival set $S^\star$ and the goal is to estimate the unknown $d$-dimensional regressor $w^\star$. This problem has a long history of study in Statistics and Machine Learning going back to the…
- Preventing Model Collapse Under Overparametrization: Optimal Mixing Ratios for Interpolation Learning and Ridge Regression
Anvit Garg, Sohom Bhattacharya, Pragya Sur · 13 February 2026
Model collapse occurs when generative models degrade after repeatedly training on their own synthetic outputs. We study this effect in overparameterized linear regression in a setting where each iteration mixes fresh real labels with synthetic labels drawn from the model fitted in the previous itera…
- Highly Adaptive Principal Component Regression
Mingxun Wang, Alejandro Schuler, Mark van der Laan, Carlos Garc\'ia Meixide · 12 February 2026
The Highly Adaptive Lasso (HAL) is a nonparametric regression method that achieves almost dimension-free convergence rates under minimal smoothness assumptions, but its implementation can be computationally prohibitive in high dimensions due to the large basis matrix it requires. The Highly Adaptive…
- Is Memorization Helpful or Harmful? Prior Information Sets the Threshold
Chen Cheng, Rina Foygel Barber · 11 February 2026
We examine the connection between training error and generalization error for arbitrary estimating procedures, working in an overparameterized linear model under general priors in a Bayesian setup. We find determining factors inherent to the prior distribution $\pi$, giving explicit conditions under…
- The Relative Instability of Model Comparison with Cross-validation
Alexandre Bayle, Lucas Janson, Lester Mackey · 10 February 2026
Cross-validation (CV) is known to provide asymptotically exact tests and confidence intervals for model improvement but only when the model comparison is relatively stable. Surprisingly, we prove that even simple, individually stable models can generate relatively unstable comparisons, calling into …
