Physical Sciences › Mathematics › Statistics and Probability
Statistical Methods and Inference
195 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume — last 12 months
Latest papers
- Maximum Mean Discrepancy with Unequal Sample Sizes via Generalized U-Statistics
Aaron Wei, Milad Jalali, Danica J. Sutherland · 17 December 2025
Existing two-sample testing techniques, particularly those based on choosing a kernel for the Maximum Mean Discrepancy (MMD), often assume equal sample sizes from the two distributions. Applying these methods in practice can require discarding valuable data, unnecessarily reducing test power. We add…
- Flexible Deep Neural Networks for Partially Linear Survival Data
Asaf Ben Arie, Malka Gorfine · 12 December 2025
We propose a flexible deep neural network (DNN) framework for modeling survival data within a partially linear regression structure. The approach preserves interpretability through a parametric linear component for covariates of primary interest, while a nonparametric DNN component captures complex …
- Closed-form $\ell_r$ norm scaling with data for overparameterized linear regression and diagonal linear networks under $\ell_p$ bias
Shuofeng Zhang, Ard Louis · 9 December 2025
For overparameterized linear regression with isotropic Gaussian design and minimum-$\ell_p$ interpolator $p\in(1,2]$, we give a unified, high-probability characterization for the scaling of the family of parameter norms $ \\{ \lVert \widehat{w_p} \rVert_r \\}_{r \in [1,p]} $ with sample size. We s…
- Transductive Conformal Inference for Full Ranking
Jean-Baptiste Fermanian (UM, IMAG, IROKO), Pierre Humbert (SU, LPSM), Gilles Blanchard (LMO, DATASHAPE) · 4 December 2025
We introduce a method based on Conformal Prediction (CP) to quantify the uncertainty of full ranking algorithms. We focus on a specific scenario where $n+m$ items are to be ranked by some ``black box'' algorithm. It is assumed that the relative (ground truth) ranking of $n$ of them is known. The obj…
- Novelty detection on path space
Ioannis Gasteratos, Antoine Jacquier, Maud Lemercier, Terry Lyons, Cristopher Salvi · 4 December 2025
We frame novelty detection on path space as a hypothesis testing problem with signature-based test statistics. Using transportation-cost inequalities of Gasteratos and Jacquier (2023), we obtain tail bounds for false positive rates that extend beyond Gaussian measures to laws of RDE solutions with s…
- Gaussian and Non-Gaussian Universality of Data Augmentation
Kevin Han Huang, Peter Orbanz, Morgane Austern · 3 December 2025
We provide universality results that quantify how data augmentation affects the variance and limiting distribution of estimates through simple surrogates, and analyze several specific models in detail. The results confirm some observations made in machine learning practice, but also lead to unexpect…
- Restricted Block Permutation for Two-Sample Testing
Jungwoo Ho · 2 December 2025
We study a structured permutation scheme for two-sample testing that restricts permutations to single cross-swaps between block-selected representatives. Our analysis yields three main results. First, we provide an exact validity construction that applies to any fixed restricted permutation set. Sec…
- Accelerating Diffusion Models with Parallel Sampling: Inference at Sub-Linear Time Complexity
Haoxuan Chen, Yinuo Ren, Lexing Ying, Grant M. Rotskoff · 1 December 2025
Diffusion models have become a leading method for generative modeling of both image and scientific data. As these models are costly to train and \emph{evaluate}, reducing the inference cost for diffusion models remains a major goal. Inspired by the recent empirical success in accelerating diffusion …
- Asymptotic Theory and Phase Transitions for Variable Importance in Quantile Regression Forests
Tomoshige Nakamura, Hiroshi Shiraishi · 1 December 2025
Quantile Regression Forests (QRF) are widely used for non-parametric conditional quantile estimation, yet statistical inference for variable importance measures remains challenging due to the non-smoothness of the loss function and the complex bias-variance trade-off. In this paper, we develop a asy…
- A PLS-Integrated LASSO Method with Application in Index Tracking
Shiqin Tang, Yining Dong, S. Joe Qin · 1 December 2025
In traditional multivariate data analysis, dimension reduction and regression have been treated as distinct endeavors. Established techniques such as principal component regression (PCR) and partial least squares (PLS) regression traditionally compute latent components as intermediary steps -- altho…
- Establishing Linear Surrogate Regret Bounds for Convex Smooth Losses via Convolutional Fenchel-Young Losses
Yuzhou Cao, Han Bao, Lei Feng, Bo An · 27 November 2025
Surrogate regret bounds, also known as excess risk bounds, bridge the gap between the convergence rates of surrogate and target losses. The regret transfer is lossless if the surrogate regret bound is linear. While convex smooth surrogate losses are appealing in particular due to the efficient estim…
- Estimation in high-dimensional linear regression: Post-Double-Autometrics as an alternative to Post-Double-Lasso
Sullivan Hu\'e, S\'ebastien Laurent, Ulrich Aiounou, Emmanuel Flachaire · 27 November 2025
Post-Double-Lasso is becoming the most popular method for estimating linear regression models with many covariates when the purpose is to obtain an accurate estimate of a parameter of interest, such as an average treatment effect. However, this method can suffer from substantial omitted variable bia…
- A Geometric Unification of Distributionally Robust Covariance Estimators: Shrinking the Spectrum by Inflating the Ambiguity Set
Man-Chung Yue, Yves Rychener, Daniel Kuhn, Viet Anh Nguyen · 25 November 2025
The state-of-the-art methods for estimating high-dimensional covariance matrices all shrink the eigenvalues of the sample covariance matrix towards a data-insensitive shrinkage target. The underlying shrinkage transformation is either chosen heuristically - without compelling theoretical justificati…
- Malliavin Calculus for Score-based Diffusion Models
Ehsan Mirafzali, Utkarsh Gupta, Patrick Wyrod, Frank Proske, Daniele Venturi, Razvan Marinescu · 25 November 2025
We introduce a new framework based on Malliavin calculus to derive exact analytical expressions for the score function $\nabla \log p_t(x)$, i.e., the gradient of the log-density associated with the solution to stochastic differential equations (SDEs). Our approach combines classical integration-by-…
- Transforming Conditional Density Estimation Into a Single Nonparametric Regression Task
Alexander G. Reisach, Olivier Collier, Alex Luedtke, Antoine Chambaz · 25 November 2025
We propose a way of transforming the problem of conditional density estimation into a single nonparametric regression task via the introduction of auxiliary samples. This allows leveraging regression methods that work well in high dimensions, such as neural networks and decision trees. Our main theo…
- Exponential Lasso: robust sparse penalization under heavy-tailed noise and outliers with exponential-type loss
The Tien Mai · 20 November 2025
In high-dimensional statistics, the Lasso is a cornerstone method for simultaneous variable selection and parameter estimation. However, its reliance on the squared loss function renders it highly sensitive to outliers and heavy-tailed noise, potentially leading to unreliable model selection and bia…
- High Dimensional Distributed Gradient Descent with Arbitrary Number of Byzantine Attackers
Wenyu Liu, Tianqiang Huang, Pengfei Zhang, Zong Ke, Minghui Min, Puning Zhao · 19 November 2025
Adversarial attacks pose a major challenge to distributed learning systems, prompting the development of numerous robust learning methods. However, most existing approaches suffer from the curse of dimensionality, i.e. the error increases with the number of model parameters. In this paper, we make a…
- Learning and Testing Convex Functions
Renato Ferreira Pinto Jr., Cassandra Marcussen, Elchanan Mossel, Shivam Nadimpalli · 17 November 2025
We consider the problems of \emph{learning} and \emph{testing} real-valued convex functions over Gaussian space. Despite the extensive study of function convexity across mathematics, statistics, and computer science, its learnability and testability have largely been examined only in discrete or res…
- Spacing Test for Fused Lasso
Rieko Tasaka, Tatsuya Kimura, Joe Suzuki · 13 November 2025
Detecting changepoints in a one-dimensional signal is a classical yet fundamental problem. The fused lasso provides an elegant convex formulation that produces a stepwise estimate of the mean, but quantifying the uncertainty of the detected changepoints remains difficult. Post-selection inference (P…
- Variational Diffusion Unlearning: A Variational Inference Framework for Unlearning in Diffusion Models under Data Constraints
Subhodip Panda, MS Varun, Shreyans Jain, Sarthak Kumar Maharana, Prathosh A. P · 11 November 2025
For a responsible and safe deployment of diffusion models in various domains, regulating the generated outputs from these models is desirable because such models could generate undesired, violent, and obscene outputs. To tackle this problem, recent works use machine unlearning methodology to forget …
- MENSA: A Multi-Event Network for Survival Analysis with Trajectory-based Likelihood Estimation
Christian Marius Lillelund, Ali Hossein Gharari Foomani, Weijie Sun, Shi-ang Qi, Russell Greiner · 11 November 2025
Most existing time-to-event methods focus on either single-event or competing-risks settings, leaving multi-event scenarios relatively underexplored. In many healthcare applications, for example, a patient may experience multiple clinical events, that can be non-exclusive and semi-competing. A commo…
- Lassoed Forests: Random Forests with Adaptive Lasso Post-selection
Jing Shang, James Bannon, Benjamin Haibe-Kains, Robert Tibshirani · 11 November 2025
Random forests are a statistical learning technique that use bootstrap aggregation to average high-variance and low-bias trees. Improvements to random forests, such as applying Lasso regression to the tree predictions, have been proposed in order to reduce model bias. However, these changes can some…
- Structural Properties, Cycloid Trajectories and Non-Asymptotic Guarantees of EM Algorithm for Mixed Linear Regression
Zhankun Luo, Abolfazl Hashemi · 10 November 2025
This work investigates the structural properties, cycloid trajectories, and non-asymptotic convergence guarantees of the Expectation-Maximization (EM) algorithm for two-component Mixed Linear Regression (2MLR) with unknown mixing weights and regression parameters. Recent studies have established glo…
- Regularized least squares learning with heavy-tailed noise is minimax optimal
Mattes Mollenhauer, Nicole M\"ucke, Dimitri Meunier, Arthur Gretton · 7 November 2025
This paper examines the performance of ridge regression in reproducing kernel Hilbert spaces in the presence of noise that exhibits a finite number of higher moments. We establish excess risk bounds consisting of subgaussian and polynomial terms based on the well known integral operator framework. T…
- Variable Selection in Maximum Mean Discrepancy for Interpretable Distribution Comparison
Kensuke Mitsuzawa, Motonobu Kanagawa, Stefano Bortoli, Margherita Grossi, Paolo Papotti · 6 November 2025
We study two-sample variable selection: identifying variables that discriminate between the distributions of two sets of data vectors. Such variables help scientists understand the mechanisms behind dataset discrepancies. Although domain-specific methods exist (e.g., in medical imaging, genetics, an…
