Physical Sciences › Mathematics › Statistics and Probability
Statistical Methods and Inference
195 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual — últimos 12 meses
Últimos artículos
- Hybrid least squares for learning functions from highly noisy data
Ben Adcock, Bernhard Hientzsch, Akil Narayan, Yiming Xu · 26 de mayo de 2026
Motivated by the need for efficient estimation of conditional expectations, we consider a least-squares function approximation problem with heavily polluted data. Existing methods that are effective in the small-noise regime are suboptimal when large noise is present. To address this issue, we propo…
- Conformalised imprecise inference for robust extrapolation under limited data
Yu Chen, Scott Ferson · 26 de mayo de 2026
Recent advances in uncertainty quantification increasingly emphasise the distinction between aleatory and epistemic uncertainty in machine learning, motivating the need for more unified frameworks. However, despite much progress in producing reliable predictions, existing methods often lack rigorous…
- Variable Clustering via Distributionally Robust Nodewise Regression
Kaizheng Wang, Xiao Xu, Xun Yu Zhou · 26 de mayo de 2026
We study a multi-factor block model for variable clustering and connect it to regularized subspace clustering through a distributionally robust version of nodewise regression. To solve the latter problem, we derive a convex relaxation, provide a data-driven approach for selecting the size of the rob…
- Diffusion-based Denoising Beats Vanilla Score Matching in Parameter Estimation: A Theoretical Explanation
Benedikt L\"utke Schwienhorst, Nadja Klein, Johannes Lederer · 25 de mayo de 2026
Score matching is an alternative to maximum likelihood estimation when the normalizing constant is unknown or too costly to evaluate. However, vanilla score matching has shown to be inefficient relative to maximum likelihood estimation for multimodal distributions with well-separated modes, which ar…
- Computational-Statistical Trade-off in Kernel Two-Sample Testing with Random Fourier Features
Ikjun Choi, Ilmun Kim · 21 de mayo de 2026
Recent years have seen a surge in methods for two-sample testing, among which the Maximum Mean Discrepancy (MMD) test has emerged as an effective tool for handling complex and high-dimensional data. Despite its success and widespread adoption, the primary limitation of the MMD test has been its quad…
- Posterior Contraction of L\'evy Adaptive B-spline Regression in Besov Spaces
Jeunghun Oh, Sewon Park, Jaeyong Lee · 20 de mayo de 2026
We investigate the asymptotic properties of the L\'evy Adaptive B-spline (LABS) regression model, a Bayesian nonparametric method that incorporates B-spline kernels into the L\'evy Adaptive Regression Kernel (LARK) model. LABS applies splines of varying degrees with independently defined knots, yiel…
- A Scalable Nonparametric Continuous-Time Survival Model through Numerical Quadrature
Chaeyeon Lee, Sehwan Kim, Hyungrok Do · 18 de mayo de 2026
Flexible continuous-time survival modeling is critical for capturing complex time-varying hazard dynamics in high-dimensional data; however, training such models remains challenging due to the intractable integral required for likelihood estimation. We introduce QSurv, a scalable deep learning frame…
- Skew-adaptive conformal prediction
Paulo C. Marques F., Helton Graziadei · 18 de mayo de 2026
We develop a skew-adaptive extension of split conformal prediction for regression. The method starts from an asymmetric interval family centered at a point prediction and uses the gauge approach to deduce the conformity score induced by this family. The inverse hyperbolic sine transform of signed sc…
- Finite Sample Bounds for Learning with Score Matching
Devin Smedira, Abhijith Jayakumar, Sidhant Misra, Marc Vuffray, Andrey Y. Lokhov · 15 de mayo de 2026
Learning of continuous exponential family distributions with unbounded support remains an important area of research for both theory and applications in high-dimensional statistics. In recent years, score matching has become a widely used method for learning exponential families with continuous vari…
- Learning density ratios in causal inference using Bregman-Riesz regression
Oliver J. Hines, Caleb H. Miles · 13 de mayo de 2026
The ratio of two probability density functions is a fundamental quantity that appears in many areas of statistics and machine learning, including causal inference, reinforcement learning, covariate shift, outlier detection, independence testing, importance sampling, and diffusion modeling. Naively e…
- Linear Response Estimators for Singular Statistical Models
Chris Elliott, Daniel Murfet · 11 de mayo de 2026
We define susceptibilities as a measure of the response of an observable quantity of a parameterized statistical model to a perturbation of the data for a general class of observables. We define estimators for these susceptibilities as statistics in a sequence of n data-points and prove that these e…
- Kernel Selection is Model Selection: A Unified Complexity-Penalized Approach for MMD Two-Sample Tests
Yijin Ni, Xiaoming Huo · 11 de mayo de 2026
The Maximum Mean Discrepancy (MMD) is a cornerstone statistic for nonparametric two-sample testing, but its test power is dictated entirely by the chosen kernel. Because any fixed kernel inherently fails to distinguish certain distributions, the kernel must be dynamically optimized. However, data-dr…
- When Does Trimming Help Conformal Prediction? A Retained-Law Diagnostic under Calibration Contamination
Congye Wang · 8 de mayo de 2026
Trimming suspicious calibration points is a common response to contamination in conformal prediction. Its effect on clean-target coverage, however, is governed by the retained law induced by trimming, not by the contamination level alone. We analyse fixed-threshold trimming as conditioning rather th…
- Covariate Balancing and Riesz Regression Should Be Guided by the Neyman Orthogonal Score in Debiased Machine Learning
Masahiro Kato · 8 de mayo de 2026
This position paper argues that, in debiased machine learning, balancing functions should be derived from the Neyman orthogonal score, not chosen only as functions of covariates. Covariate balancing is effective when the regression error entering the score can be represented by functions of covariat…
- Proximal Projection for Doubly Sparse Regularized Models
Jia Wei He, R. Ayesha Ali, Gerarda Darlington · 7 de mayo de 2026
Regularization is often used in high-dimensional regression settings to generate a sparse model, which can save tremendous computing resources and identify predictors that are most strongly associated with the response. When the predictors can be represented by a Gaussian graphical model, the struct…
- Adaptive Estimation and Optimal Control in Offline Contextual MDPs without Stationarity
Riddhiman Bhattacharyya, Sayak Chakrabarty, Imon Banerjee · 6 de mayo de 2026
Contextual MDPs are powerful tools with wide applicability in areas from biostatistics to machine learning. However, specializing them to offline datasets has been challenging due to a lack of robust, theoretically backed methods. Our work tackles this problem by introducing a new approach towards a…
- A note on the unique properties of the Kullback--Leibler divergence for sampling via gradient flows
Francesca Romana Crucinio · 6 de mayo de 2026
We consider the problem of sampling from a probability distribution $\pi$ which admits a density w.r.t. a dominating measure. It is well known that this can be written as an optimisation problem over the space of probability distributions in which we aim to minimise a divergence from $\pi$. The opti…
- Denoising data using convex relaxations
Charles Fefferman, Aalok Gangopadhyay, Matti Lassas, Jonathan Marty, Hariharan Narayanan · 5 de mayo de 2026
We study the problem of denoising observations \(Y_i=X_i+Z_i\), where the latent variables \(X_i\) are sampled from a low-dimensional manifold in \(\mathbb{R}^n\) and the noise variables \(Z_i\) are isotropic Gaussian. We propose a convex-relaxation estimator that first reduces dimension by principa…
- Sparse Regression under Correlation and Weak Signals: A Reproducible Benchmark of Classical and Bayesian Methods
Hao Xiao · 5 de mayo de 2026
Choosing between classical and Bayesian sparse regression methods involves a real trade-off: penalized estimators like Lasso run in milliseconds but give no uncertainty estimates,while Horseshoe and Spike-and-Slab priors produce full posteriors but need MCMC chains that take minutes per fit.Surprisi…
- A Semi-Supervised Kernel Two-Sample Test
Gyumin Lee, Shubhanshu Shekhar, Ilmun Kim · 5 de mayo de 2026
We consider the problem of two-sample testing in a semi-supervised setting with abundant unlabeled covariate data. Standard two-sample tests neglect covariate information, which has the potential to significantly boost performance. However, incorporating covariates potentially breaks the exchangeabi…
- Linear Models, Variable Selection, Artificial Intelligence
By Riyadh Alrawkan, Edward Boone, Ryad Ghanam, Anton Westveld · 1 de mayo de 2026
Variable selection in linear regression models has been a problem since hypothesis testing began. Which variables to include or exclude from a model is not an easy task. Techniques such as Forward, Back ward, Stepwise Regression sequentially add or delete variables from a model. Penalized likelihood…
- Evaluating Computational Pathology Foundation Models for Prostate Cancer Grading under Distribution Shifts
Fredrik K. Gustafsson, Mattias Rantalainen · 29 de abril de 2026
Pathology foundation models (PFMs) have emerged as powerful pretrained encoders for computational pathology, but their robustness under clinically relevant distribution shifts remains insufficiently understood. We benchmark the robustness of recent PFMs in the setting of prostate cancer grading from…
- Statistical Test for Diffusion-Based Anomaly Localization via Selective Inference
Teruyuki Katsuoka, Tomohiro Shiraishi, Daiki Miwa, Vo Nguyen Le Duy, Ichiro Takeuchi · 28 de abril de 2026
Anomaly localization in images -- identifying regions that deviate from normal patterns -- is vital in applications such as medical diagnosis and industrial inspection. A recent trend is the use of image generation models in anomaly localization, where these models generate normal-looking counterpar…
- Flexible Deep Neural Networks for Partially Linear Survival Data: Estimation and Survival Inference
Asaf Ben Arie, Malka Gorfine · 28 de abril de 2026
We propose a flexible deep neural network (DNN) framework for modeling survival data within a partially linear regression structure. The approach preserves interpretability through a parametric linear component for covariates of primary interest, while a nonparametric DNN component captures complex …
- Calibrated Principal Component Regression
Yixuan Florence Wu, Yilun Zhu, Lei Cao, Naichen Shi · 27 de abril de 2026
We propose a new method for statistical inference in generalized linear models. In the overparameterized regime, Principal Component Regression (PCR) reduces variance by projecting high-dimensional data to a low-dimensional principal subspace before fitting. However, PCR incurs truncation bias whene…
