Physical Sciences › Mathematics › Statistics and Probability
Statistical Methods and Inference
195 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual — últimos 12 meses
Últimos artículos
- Learning False Discovery Rate Control via Model-Based Neural Networks
Arnau Vilella, Jasin Machkour, Michael Muma, Daniel P. Palomar · 6 de febrero de 2026
Controlling the false discovery rate (FDR) in high-dimensional variable selection requires balancing rigorous error control with statistical power. Existing methods with provable guarantees are often overly conservative, creating a persistent gap between the realized false discovery proportion (FDP)…
- Unified Unbiased Variance Estimation for Maximum Mean Discrepancy: Robust Finite-Sample Performance with Imbalanced Data and Exact Acceleration under Null and Alternative Hypotheses
Shijie Zhong, Yikun Yang, Da Gong, Jiangfeng Fu · 5 de febrero de 2026
The maximum mean discrepancy (MMD) is a kernel-based nonparametric statistic for two-sample testing, whose inferential accuracy depends critically on variance characterization. Existing work provides various finite-sample estimators of the MMD variance, often differing under the null and alternative…
- PCA of probability measures: Sparse and Dense sampling regimes
Gachon Erell, J\'er\'emie Bigot, Elsa Cazelles · 3 de febrero de 2026
A common approach to perform PCA on probability measures is to embed them into a Hilbert space where standard functional PCA techniques apply. While convergence rates for estimating the embedding of a single measure from $m$ samples are well understood, the literature has not addressed the setting i…
- On the calibration of survival models with competing risks
Julie Alberge (DREES), Tristan Haugomat (DREES), Ga\"el Varoquaux (SODA, IP Paris), Judith Ab\'ecassis (SODA, IP Paris) · 3 de febrero de 2026
Survival analysis deals with modeling the time until an event occurs, and accurate probability estimates are crucial for decision-making, particularly in the competing-risks setting where multiple events are possible. While recent work has addressed calibration in standard survival analysis, the com…
- Multivariate Standardized Residuals for Conformal Prediction
Sacha Braun, Eug\`ene Berta, Michael I. Jordan, Francis Bach · 3 de febrero de 2026
While split conformal prediction guarantees marginal coverage, approaching the stronger property of conditional coverage is essential for reliable uncertainty quantification. Naive conformal scores, however, suffer from poor conditional coverage in heteroskedastic settings. In univariate regression,…
- Tabular Foundation Models Can Do Survival Analysis
Da In Kim, Wei Siang Lai, Kelly W. Zhang · 2 de febrero de 2026
While tabular foundation models have achieved remarkable success in classification and regression, adapting them to model time-to-event outcomes for survival analysis is non-trivial due to right-censoring, where data observations may end before the event occurs. We develop a classification-based fra…
- Efficient Group Lasso Regularized Rank Regression with Data-Driven Parameter Determination
Meixia Lin, Meijiao Shi, Yunhai Xiao, Qian Zhang · 29 de enero de 2026
High-dimensional regression often suffers from heavy-tailed noise and outliers, which can severely undermine the reliability of least-squares based methods. To improve robustness, we adopt a non-smooth Wilcoxon score based rank objective and incorporate structured group sparsity regularization, a na…
- Kernel smoothing on manifolds
Eunseong Bae, Wolfgang Polonik · 26 de enero de 2026
Under the assumption that data lie on a compact (unknown) manifold without boundary, we derive finite sample bounds for kernel smoothing and its (first and second) derivatives, and we establish asymptotic normality through Berry-Esseen type bounds. Special cases include kernel density estimation, ke…
- Estimation of discrete distributions in relative entropy, and the deviations of the missing mass
Jaouad Mourtada · 26 de enero de 2026
We study the problem of estimating a distribution over a finite alphabet from an i.i.d. sample, with accuracy measured in relative entropy (Kullback-Leibler divergence). While optimal bounds on the expected risk are known, high-probability guarantees remain less well-understood. First, we analyze th…
- Finite-Sample Inference for Sparsely Permuted Linear Regression
Hirofumi Ota, Masaaki Imaizumi · 22 de enero de 2026
We study a noisy linear observation model with an unknown permutation called permuted/shuffled linear regression, where responses and covariates are mismatched and the permutation forms a discrete, factorial-size parameter. This unknown permutation is a key component of the data-generating process, …
- Statistical Learning Theory for Distributional Classification
Christian Fiedler · 22 de enero de 2026
In supervised learning with distributional inputs in the two-stage sampling setup, relevant to applications like learning-based medical screening or causal learning, the inputs (which are probability distributions) are not accessible in the learning phase, but only samples thereof. This problem is p…
- Large Data Limits of Laplace Learning for Gaussian Measure Data in Infinite Dimensions
Zhengang Zhong, Yury Korolev, Matthew Thorpe · 22 de enero de 2026
Laplace learning is a semi-supervised method, a solution for finding missing labels from a partially labeled dataset utilizing the geometry given by the unlabeled data points. The method minimizes a Dirichlet energy defined on a (discrete) graph constructed from the full dataset. In finite dimension…
- Approximate full conformal prediction in RKHS
Davidson Lova Razafindrakoto, Alain Celisse, J\'er\^ome Lacaille · 21 de enero de 2026
Full conformal prediction is a framework that implicitly formulates distribution-free confidence prediction regions for a wide range of estimators. However, a classical limitation of the full conformal framework is the computation of the confidence prediction regions, which is usually impossible sin…
- Unified Unbiased Variance Estimation for MMD: Robust Finite-Sample Performance with Imbalanced Data and Exact Acceleration under Null and Alternative Hypotheses
Shijie Zhong, Jiangfeng Fu, Yikun Yang · 21 de enero de 2026
The maximum mean discrepancy (MMD) is a kernel-based nonparametric statistic for two-sample testing, whose inferential accuracy depends critically on variance characterization. Existing work provides various finite-sample estimators of the MMD variance, often differing under the null and alternative…
- Variable transformations in consistent loss functions
Hristos Tyralis, Georgia Papacharalampous · 21 de enero de 2026
The empirical use of variable transformations within (strictly) consistent loss functions is widespread, yet a theoretical understanding is lacking. To address this gap, we develop a theoretical framework that establishes formal characterizations of (strict) consistency for such transformed loss fun…
- The Interpolating Information Criterion for Overparameterized Models
Liam Hodgkinson, Chris van der Heide, Robert Salomone, Fred Roosta, Michael W. Mahoney · 13 de enero de 2026
The problem of model selection is considered for the setting of interpolating estimators, where the number of model parameters exceeds the size of the dataset. Classical information criteria typically consider the large-data limit, penalizing model size. However, these criteria are not appropriate i…
- Fast Conformal Prediction using Conditional Interquantile Intervals
Naixin Guo, Rui Luo, Zhixin Zhou · 7 de enero de 2026
We introduce Conformal Interquantile Regression (CIR), a conformal regression method that efficiently constructs near-minimal prediction intervals with guaranteed coverage. CIR leverages black-box machine learning models to estimate outcome distributions through interquantile ranges, transforming th…
- Conformal Prediction for Dose-Response Models with Continuous Treatments
Jarne Verhaeghe, Jef Jonkers, Sofie Van Hoecke · 7 de enero de 2026
Understanding the dose-response relation between a continuous treatment and the outcome for an individual can greatly drive decision-making, particularly in areas like personalized drug dosing and personalized healthcare interventions. Point estimates are often insufficient in these high-risk enviro…
- Conformal Blindness: A Note on $A$-Cryptic change-points
Johan Hallberg Szabadv\'ary · 6 de enero de 2026
Conformal Test Martingales (CTMs) are a standard method within the Conformal Prediction framework for testing the crucial assumption of data exchangeability by monitoring deviations from uniformity in the p-value sequence. Although exchangeability implies uniform p-values, the converse does not hold…
- Copula Discrepancy: Benchmarking Dependence Structure
Agnideep Aich, Ashit Baran Aich · 30 de diciembre de 2025
We study a simple statistic for benchmarking how well a sample preserves a known bivariate dependence structure. Given a target copula family (Clayton or Gumbel) and parameter $\theta_P$, the Copula Discrepancy (CD) compares the target Kendall's tau $\tau(\theta_P)$ with the Kendall's tau implied by…
- Multivariate Conformal Prediction via Conformalized Gaussian Scoring
Sacha Braun, Eug\`ene Berta, Michael I. Jordan, Francis Bach · 30 de diciembre de 2025
While achieving exact conditional coverage in conformal prediction is unattainable without making strong, untestable regularity assumptions, the promise of conformal prediction hinges on finding approximations to conditional guarantees that are realizable in practice. A promising direction for obtai…
- A general framework for deep learning
William Kengne, Modou Wade · 30 de diciembre de 2025
This paper develops a general approach for deep learning for a setting that includes nonparametric regression and classification. We perform a framework from data that fulfills a generalized Bernstein-type inequality, including independent, $\phi$-mixing, strongly mixing and $\mathcal{C}$-mixing obs…
- An Efficient Minimax Optimal Estimator For Multivariate Convex Regression
Gil Kur, Eli Putterman · 30 de diciembre de 2025
This work studies the computational aspects of multivariate convex regression in dimensions $d \ge 5$. Our results include the \emph{first} estimators that are minimax optimal (up to logarithmic factors) with polynomial runtime in the sample size for both $L$-Lipschitz convex regression, and $\Gamma…
- Subgroup Discovery with the Cox Model
Zachary Izzo, Iain Melvin · 25 de diciembre de 2025
We study the problem of subgroup discovery for survival analysis, where the goal is to find an interpretable subset of the data on which a Cox model is highly accurate. Our work is the first to study this particular subgroup problem, for which we make several contributions. Subgroup discovery meth…
- Optimal Anytime-Valid Tests for Composite Nulls
Shubhanshu Shekhar · 24 de diciembre de 2025
We consider the problem of designing optimal level-$\alpha$ power-one tests for composite nulls. Given a parameter $\alpha \in (0,1)$ and a stream of $\mathcal{X}$-valued observations $\{X_n: n \geq 1\} \overset{i.i.d.}{\sim} P$, the goal is to design a level-$\alpha$ power-one test $\tau_\alpha$ fo…
