Physical Sciences › Mathematics › Statistics and Probability
Statistical Methods and Inference
195 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen — letzte 12 Monate
Neueste Paper
- Beyond Marginal Validity: Finite-Sample Guarantees for Localized Conformal Prediction
Anton Conrad, Rustam Isaev, Denis Belomestny, Eric Moulines, Sergey Samsonov · 7. August 2026
Conformal prediction endows arbitrary black-box predictors with finite-sample, distribution-free marginal coverage, yet marginal validity can hide severe covariate-specific miscalibration, while exact distribution-free conditional coverage is finite-sample unattainable. Randomly localized conformal …
- Nonparametric Goodness-of-fit Testing under Covariate Shift
Zhen Hou, Dong Xia · 6. August 2026
This paper develops procedures for nonparametric goodness-of-fit testing under covariate shift, where labelled data are drawn from a source population but goodness-of-fit is evaluated for a target population. The distribution mismatch is quantified by either a bounded moment condition or a sub-expon…
- Double Descent in Gradient Boosting Decision Trees via Split-Candidate Scaling
Ryuichi Kanoh · 5. August 2026
Double descent is commonly studied by scaling an explicit capacity parameter, such as neural-network width. For gradient boosting decision trees (GBDTs), however, an analogous single-axis capacity parameter has not been established. We propose the number of split candidates as an operational capacit…
- A Simple Approximation to the Distribution of the Ridge Regression Estimator
Jos\'e Luis Montiel Olea, Ryan Strong, Amilcar Velez, Zhuoheng Xu, Haomin Yu · 4. August 2026
We present a simple Gaussian approximation to the finite-sample distribution of the classical ridge regression estimator. Our approximation captures the fact that, in finite samples, the ridge regression estimator trades off bias and variance to reduce estimation and prediction error. Our approximat…
- Active Regression for Single-Index Models with Unknown Link Functions
Chansophea Wathanak In, Yi Li, Wai Ming Tai, Xuan Wu · 4. August 2026
This paper studies active regression for single-index models under general $\ell_p$-loss with an unknown $1$-Lipschitz link function $f$, formulated as $\min_{f,x} \|f(Ax)-b\|_p^p$ with full access to $A$ but coordinate-query access to $b$. Prior work established upper bounds for known link function…
- Finite-Probe Total-Variation Certificates for Finite-Basis Drifting Models
Sam Andersson, Ricky Mol\'en · 4. August 2026
Drifting objectives compare a target and model distribution through a vector field observed noisily at finitely many locations. We ask what distributional conclusion such a frozen measurement system warrants. For integrable antisymmetric interactions and absolutely continuous laws in a declared fini…
- Beyond Modern Asymptotics for Log-Likelihood Ratios in Logistic Regression
Hugo Chardon, Reese Pathak, Nikita Zhivotovskiy · 4. August 2026
We characterize the finite sample behavior of the log-likelihood ratio statistic in binary logistic regression, uniformly over both the design and the target parameter. For $n\geq d\geq 3$, we determine, up to universal constants, its worst case $(1-\delta)$ quantile over all fixed collections of de…
- A reproducible and extensible framework for benchmarking competing risks survival models
Bego\~na B. Sierra, Colin McLean, Peter S. Hall, Sarah Friedrich-Welz, Catalina A. Vallejos · 4. August 2026
A wide range of statistical and machine learning methods have been proposed for survival analysis with competing risks, where the occurrence of one event (i.e., cancer death) precludes the occurrence of other events (i.e., cardiovascular disease death). Despite these methodological advances, their s…
- Who Wins Where? Conformal Model Comparison for Local Superiority
Yi Zhou, Baishi Li, Xuan Yao, Ke-Wei Huang · 3. August 2026
Standard model comparison is global, aggregating losses across the covariate space to declare a single winner. This can obscure heterogeneous performance, where different models are preferable in different regions. We introduce conformalized local model comparison, a split-sample framework for const…
- Retrospective Orthogonal Design: Response-Surface Reconstruction from Observational Data
Lawrence Fulton, Christopher Fulton, Arvind Sharma, Aleksandar Tomic · 30. Juli 2026
Regression estimates from observational data can depend on specification under multicollinearity, while sequential sums of squares (SS) depend on term order. We introduce Retrospective Orthogonal Design (ROD), which reconstructs conditional mean surfaces on a probability-balanced lattice. ROD preser…
- When Kernel Ridge Regression Meets the H\"older-Zygmund Class: Minimax Optimality and Failure of Properness
Yuxuan Hou · 30. Juli 2026
We study kernel ridge regression for nonparametric regression over the H\"older-Zygmund class. Using an RKHS equivalent to a Sobolev space of smoothness s+d/2, we prove that misspecified KRR attains the minimax L2 rate n^{-2s/(2s+d)}. We also show that properness fails in the H\"older-Zygmund norm: …
- Breaking the Curse with BAND: Nonparametric Distribution Estimation in High Dimensions
Shuo-Chieh Huang, Chien-Ming Chi, Jau-er Chen · 30. Juli 2026
Minimax-optimal rates for multivariate distribution estimation are known to suffer from the curse of dimensionality. We propose a sparse Bayesian network approach in which each conditional probability is estimated using sparsity-aware conditional mean methods. The resulting estimator, \textit{BAyesi…
- Distributional Split Criteria for Random Forests: Extensions, Shrinkage, and the Robustness of Mean Splitting
Silas Koemen · 28. Juli 2026
Distributional random forests replace mean-based CART splitting with criteria that compare the full conditional response distribution in candidate children. We implement and systematically study a family of such criteria inside a single honest-forest implementation: isotropic random-Fourier-feature …
- Automatic knot selection in smooth additive models
Nicol\'as Carrizosa, Vanesa Guerrero, Mar\'ia Durb\'an · 24. Juli 2026
B-spline regression constitutes a widely used framework for nonparametric modeling. The performance of this methodology depends on specifying the number and placement of changepoints, known as knots, prior to the estimation process. Such knot sequence determines the dimension of the B-spline basis u…
- Adaptive deep nonparametric regression from dependent data under covariate shift
William Kengne, Ehud Mossa Ockegna · 23. Juli 2026
Covariate shift often occurs because, in many real applications, the source and the target observations may be generated from different distributions. In this case, the standard metric under the source distribution is not appropriate. This paper considers deep neural network estimators for nonparame…
- Isotonic Conformal Prediction
Daniel Bensimon, Sean Xiang Yu, Eric D. Kolaczyk, Archer Y. Yang · 21. Juli 2026
A point prediction that is well calibrated on average can still be systematically biased conditional on its own value, undermining its use in downstream decision-making. We consider two objectives for reliable uncertainty quantification: self-calibration, requiring a point prediction to be unbiased …
- Pitfalls of Administrative Censoring in Survival Models with Time-Indexed Inputs
Yanqi Xu, Hui Dai, Carlos Fernandez-Granda, Krzysztof J. Geras, Yiqiu Shen · 14. Juli 2026
Survival models can model time-to-event outcomes using partially observed data. They are widely used in clinical prediction, including cancer risk, disease progression, treatment response, and mortality. Recent models often rely on rich inputs collected at a specific clinical encounter, such as medi…
- Sharp Concentration Bounds for Bundle-Valued Statistics on Manifolds
Swagatam Das, Vaclav Snasel · 14. Juli 2026
Many geometric statistics and manifold learning pipelines routinely produce observations -- such as tangent vectors or local frames -- whose natural home is a varying family of fibers attached to different points of a base manifold, rather than a single shared vector space. Forming empirical average…
- The Regularization Parameter: Sparse Precision Matrix Estimation
Aryan Eftekhari, Daniel Sergio Vega, Ernst-Jan Camiel Wit, Olaf Schenk · 10. Juli 2026
Sparse precision matrix estimation provides an interpretable and computationally efficient framework for modeling conditional dependencies in high-dimensional, low-sample-size data. A recurring challenge is appropriately selecting the regularization parameter that controls estimator sparsity and str…
- Approximate full conformal prediction in an RKHS
Davidson Lova Razafindrakoto, Alain Celisse, J\'er\^ome Lacaille · 9. Juli 2026
Full conformal prediction is a framework that implicitly formulates distribution-free confidence prediction regions for a wide range of estimators. However, a classical limitation of the full conformal framework is the computation of the confidence prediction regions, which is usually impossible sin…
- K-ABENA: K-Adaptive Backpropagation with Error-based N-exclusion Algorithm : (Compensated Loss-Based Sample Exclusion with Unbiased Gradient Estimation)
Jean-Francois Bonbhel · 8. Juli 2026
We present K-ABENA (K-Adaptive Backpropagation with Error-based N-exclusion Algorithm), a selective gradient computation framework that reduces per-iteration training cost by excluding a fraction of low-loss ("minor") observations from the backward pass. Its canonical form (v3) combines a defensive-…
- Distribution-free Deviation Bounds and The Role of Domain Knowledge in Learning via Model Selection with Cross-validation Risk Estimation
Diego Marcondes, Cl\'audia Peixoto · 7. Juli 2026
Cross-validation is one of the most widely used tools for risk estimation and model selection in statistics and machine learning, yet its theoretical properties when embedded in a learning procedure remain insufficiently understood. This paper develops a general, distribution-free framework for lear…
- Nonparametric Control Koopman Operators
Petar Bevanda, Bas Driessen, Lucian Cristian Iacob, Stefan Sosnowski, Roland T\'oth, Sandra Hirche · 7. Juli 2026
This paper presents a novel Koopman composition operator representation framework for control systems in reproducing kernel Hilbert spaces (RKHSs) that is free of explicit dictionary or input parametrizations. By establishing fundamental equivalences between different model representations, we are a…
- On Pairwise Quantile Regression -- Statistical Guarantees and Applications
Romain Th\'er\'ezien, Stephan Cl\'emen\c{c}on, Fantin Girard, Hamza El-Abdouni · 7. Juli 2026
Quantile regression provides a powerful tool for summarizing the conditional distribution of a real valued random variable (r.v.) of interest $Y$ as a function of covariates $Z$ in cases where it shows a large dispersion with high probability, going beyond the situation where standard least square r…
- Aggregation with Exponential Weights is Optimal in Expectation
Mikael M{\o}ller H{\o}gsgaard, Patrick Rebeschini, Tobias Wegel · 3. Juli 2026
The aggregation with exponential weights (AEW) estimator is not fully understood in the basic setting of model selection aggregation with squared loss. In particular, whether it is minimax-rate optimal in expectation for large enough fixed temperatures and under random design has been an open proble…
