Physical Sciences › Mathematics › Statistics and Probability
Statistical Methods and Inference
195 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen — letzte 12 Monate
Neueste Paper
- Amortized Vine Copulas for High-Dimensional Density and Information Estimation
Houman Safaai · 23. April 2026
Modeling high-dimensional dependencies while keeping likelihoods tractable remains challenging. Classical vine-copula pipelines are interpretable but can be expensive, while many neural estimators are flexible but less structured. In this work, we propose Vine Denoising Copula (VDC), an amortized vi…
- A Ridge Too Far: Correcting Over-Shrinkage via Negative Regularization
Dongseok Kim, Gisung Oh · 21. April 2026
Conventional regularization is designed to control variance, but in small-data regression it can also aggravate underfitting when predictive signal is concentrated in weak directions of a restricted representation. We study a negative-capable ridge family that permits a feasible negative region when…
- Conformal Risk Control under Non-Monotone Losses: Theory and Finite-Sample Guarantees
Tareq Aldirawi, Yun Li, Wenge Guo · 21. April 2026
Conformal risk control (CRC) provides distribution-free guarantees for controlling the expected loss at a user-specified level. Existing theory typically assumes that the loss decreases monotonically with a tuning parameter that governs the size of the prediction set. However, this assumption is oft…
- MinShap: A Modified Shapley Value Approach for Feature Selection
Chenghui Zheng, Garvesh Raskutti · 17. April 2026
Feature selection is a classical problem in statistics and machine learning, and it continues to remain an extremely challenging problem especially in the context of unknown non-linear relationships with dependent features. On the other hand, Shapley values are a classic solution concept from cooper…
- Beyond Fixed False Discovery Rates: Post-Hoc Conformal Selection with E-Variables
Meiyi Zhu, Osvaldo Simeone · 14. April 2026
Conformal selection (CS) uses calibration data to identify test inputs whose unobserved outcomes are likely to satisfy a pre-specified minimal quality requirement, while controlling the false discovery rate (FDR). Existing methods fix the target FDR level before observing data, which prevents the us…
- Cost-optimal Sequential Testing via Doubly Robust Q-learning
Doudou Zhou, Yiran Zhang, Dian Jin, Yingye Zheng, Lu Tian, Tianxi Cai · 14. April 2026
Clinical decision-making often involves selecting tests that are costly, invasive, or time-consuming, motivating individualized, sequential strategies for what to measure and when to stop ascertaining. We study the problem of learning cost-optimal sequential decision policies from retrospective data…
- Stability of a Generalized Debiased Lasso with Applications to Resampling-Based Variable Selection
Jingbo Liu · 14. April 2026
We propose a generalized debiased Lasso estimator based on a stability principle. When a single column of the design matrix is perturbed, the estimator admits a simple update formula that can be computed from the original solution. Under sub-Gaussian designs with well-conditioned covariance, this ap…
- Choosing the Right Regularizer for Applied ML: Simulation Benchmarks of Popular Scikit-learn Regularization Frameworks
Benjamin S. Knight, Ahsaas Bajaj · 7. April 2026
This study surveys the historical development of regularization, tracing its evolution from stepwise regression in the 1960s to recent advancements in formal error control, structured penalties for non-independent features, Bayesian methods, and l0-based regularization (among other techniques). We e…
- Fused Multinomial Logistic Regression Utilizing Summary-Level External Machine-learning Information
Chi-Shian Dai, Jun Shao · 7. April 2026
In many modern applications, a carefully designed primary study provides individual-level data for interpretable modeling, while summary-level external information is available through black-box, efficient, and nonparametric machine-learning predictions. Although summary-level external information h…
- Sparse Max-Affine Regression
Haitham Kanj, Seonho Kim, Kiryung Lee · 7. April 2026
This paper presents Sparse Gradient Descent as a solution for variable selection in convex piecewise linear regression, where the model is given as the maximum of $k$-affine functions $ x \mapsto \max_{j \in [k]} \langle a_j^\star, x \rangle + b_j^\star$ for $j = 1,\dots,k$. Here, $\{ a_j^\star\}_{j…
- Distributional Statistics Restore Training Data Auditability in One-step Distilled Diffusion Models
Muxing Li, Zesheng Ye, Sharon Li, Andy Song, Guangquan Zhang, Feng Liu · 6. April 2026
The proliferation of diffusion models trained on web-scale, provenance-uncertain image collections has made it essential, yet technically unresolved, to determine whether a model has learned from specific copyrighted data without authorization. Current methods primarily rely on the memorization effe…
- Non-monotonicity in Conformal Risk Control
Tareq Aldirawi, Yun Li, Wenge Guo · 3. April 2026
Conformal risk control (CRC) provides distribution-free guarantees for controlling the expected loss at a user-specified level. Existing theory typically assumes that the loss decreases monotonically with a tuning parameter that governs the size of the prediction set. This assumption is often violat…
- On the Asymptotics of Self-Supervised Pre-training: Two-Stage M-Estimation and Representation Symmetry
Mohammad Tinati, Stephen Tu · 31. März 2026
Self-supervised pre-training, where large corpora of unlabeled data are used to learn representations for downstream fine-tuning, has become a cornerstone of modern machine learning. While a growing body of theoretical work has begun to analyze this paradigm, existing bounds leave open the question …
- On some practical challenges of conformal prediction
Liang Hong, Noura Raydan Nasreddine · 31. März 2026
Conformal prediction is a model-free machine learning method for constructing prediction regions at a guaranteed coverage probability level. However, a data scientist often faces three challenges in practice: (i) the determination of a conformal prediction region is only approximate, jeopardizing th…
- Uniform Laws of Large Numbers in Product Spaces
Ron Holzman, Shay Moran, Alexander Shlimovich · 26. März 2026
Uniform laws of large numbers form a cornerstone of Vapnik--Chervonenkis theory, where they are characterized by the finiteness of the VC dimension. In this work, we study uniform convergence phenomena in cartesian product spaces, under assumptions on the underlying distribution that are compatible …
- Beyond Consistency: Inference for the Relative risk functional in Deep Nonparametric Cox Models
Sattwik Ghosal, Xuran Meng, Yi Li · 26. März 2026
There remain theoretical gaps in deep neural network estimators for the nonparametric Cox proportional hazards model. In particular, it is unclear how gradient-based optimization error propagates to population risk under partial likelihood, how pointwise bias can be controlled to permit valid infere…
- Noise-contrastive Online Change Point Detection
Nikita Puchkin, Artur Goldman, Konstantin Yakovlev, Valeriia Dzis, Uliana Vinogradova · 24. März 2026
We suggest a novel procedure for online change point detection. Our approach expands an idea of maximizing a discrepancy measure between points from pre-change and post-change distributions. This leads to flexible algorithms suitable for both parametric and nonparametric scenarios. We prove non-asym…
- Unlearning in Diffusion models under Data Constraints: A Variational Inference Approach
Subhodip Panda, Varun M S, Shreyans Jain, Sarthak Kumar Maharana, Prathosh A. P · 24. März 2026
For a responsible and safe deployment of diffusion models in various domains, regulating the generated outputs from these models is desirable because such models could generate undesired, violent, and obscene outputs. To tackle this problem, recent works use machine unlearning methodology to forget …
- Exponential Family Discriminant Analysis: Generalizing LDA-Style Generative Classification to Non-Gaussian Models
Anish Lakkapragada · 24. März 2026
We introduce Exponential Family Discriminant Analysis (EFDA), a unified generative framework that extends classical Linear Discriminant Analysis (LDA) beyond the Gaussian setting to any member of the exponential family. Under the assumption that each class-conditional density belongs to a common exp…
- Starting Off on the Wrong Foot: Pitfalls in Data Preparation
Jiayi Guo, Panyi Dong, Zhiyu Quan · 20. März 2026
When working with real-world insurance data, practitioners often encounter challenges during the data preparation stage that can undermine the statistical validity and reliability of downstream modeling. This study illustrates that conventional data preparation procedures such as random train-test p…
- Precise Performance of Linear Denoisers in the Proportional Regime
Reza Ghane, Danil Akhtiamov, Babak Hassibi · 20. März 2026
In the present paper we study the performance of linear denoisers for noisy data of the form $\mathbf{x} + \mathbf{z}$, where $\mathbf{x} \in \mathbb{R}^d$ is the desired data with zero mean and unknown covariance $\mathbf{\Sigma}$, and $\mathbf{z} \sim \mathcal{N}(0, \mathbf{\Sigma}_{\mathbf{z}})$ …
- Consistency of the $k$-Nearest Neighbor Regressor under Complex Survey Designs
Caren Hasler · 19. März 2026
We study the consistency of the $k$-nearest neighbor regressor under complex survey designs. While consistency results for this algorithm are well established for independent and identically distributed data, corresponding results for complex survey data are lacking. We show that the $k$-nearest nei…
- Statistical Inference for Online Algorithms
Selina Carter, Arun K Kuchibhotla · 19. März 2026
The construction of confidence intervals and hypothesis tests for functionals is a cornerstone of statistical inference. Traditionally, the most efficient procedures - such as the Wald interval or the Likelihood Ratio Test - require both a point estimator and a consistent estimate of its asymptotic …
- Minimum Volume Conformal Sets for Multivariate Regression
Sacha Braun, Liviu Aolaritei, Michael I. Jordan, Francis Bach · 19. März 2026
Conformal prediction provides a principled framework for constructing predictive sets with finite-sample validity. While much of the focus has been on univariate response variables, existing multivariate methods either impose rigid geometric assumptions or rely on flexible but computationally expens…
- Transfer Learning with Distance Covariance for Random Forest: Error Bounds and an EHR Application
Chenze Li, Subhadeep Paul · 17. März 2026
We propose a method for transfer learning in nonparametric regression using a random forest (RF) with distance covariance-based feature weights, assuming the unknown source and target regression functions are sparsely different. Our method obtains residuals from a source domain-trained Centered RF (…
