In this paper, we consider the feature screening problem of ultra-high dimensional longitudinal heterogeneous data, which significantly extends the existing frameworks on ultra-high dimensional heterogeneous data and ultra-high dimensional longitudinal data focusing only on mean regression. A quantile adaptive feature screening approach is proposed by integrating independence screening with quadratic inference functions (QIF). This framework offers two distinctive features: (1) it takes into account the within-subject dependency and is more efficient than that of ignoring the correlation and assuming independence for each subject; (2) it allows the set of active variables to vary with different quantiles, thereby providing a more comprehensive description for the real data and flexibility to accommodate heterogeneity. The sure screening property is shown under some regularity conditions. Some simulation studies and a real data analysis are conducted to assess the effectiveness of the proposed screening method. The numerical results indicate that the proposed method is an effective tool to handle with the ultra-high dimensional longitudinal data.
In this paper, we consider the feature screening problem of ultra-high dimensional longitudinal heterogeneous data, which significantly extends the existing frameworks on ultra-high dimensional heterogeneous data and ultra-high dimensional longitudinal data focusing only on mean regression. A quantile adaptive feature screening approach is proposed by integrating independence screening with quadratic inference functions (QIF). This framework offers two distinctive features: (1) it takes into account the within-subject dependency and is more efficient than that of ignoring the correlation and assuming independence for each subject; (2) it allows the set of active variables to vary with different quantiles, thereby providing a more comprehensive description for the real data and flexibility to accommodate heterogeneity. The sure screening property is shown under some regularity conditions. Some simulation studies and a real data analysis are conducted to assess the effectiveness of the proposed screening method. The numerical results indicate that the proposed method is an effective tool to handle with the ultra-high dimensional longitudinal data.
Price includes VAT (Russian Federation)
Instant access to the full article PDF.

The datasets analyzed during the current study are publicly available. All data and code supporting the results are available from the corresponding author upon reasonable request.
Brown B, Wang Y (2005) Standard errors and covariance matrices for smoothed rank estimators. Biometrika 92:149–158
Candès EJ, Tao T (2007) The dantzig selector: Statistical estimation when \(p\) is much larger than \(n\). Ann Stat 35(6):2313–2351
Cheng MY, Honda T, Li JL et al (2014) Nonparametric independence screening and structure identification for ultra-high dimensional longitudinal data. Ann Stat 42:1819–1849
Cheng XW, Li G, Wang H (2024) The concordance filter: an adaptive model-free feature screening procedure. Comput Stat 39:2413–2436
Chu WH, Li RZ, Reimherr M (2016) Feature screening for time-varying coefficient models with ultrahigh-dimensional longitudinal data. Ann Appl Stat 10:596–617
Diao TB, Li B (2025) Communication-efficient feature screening for ultrahigh-dimensional data under quantile regression. Stat 14:e70055
Fan JQ, Li RZ (2001) Variable selection via nonconcave penalized likelihood and its oracle properties. J Am Stat Assoc 96:1348–1360
Fan JQ, Lv JC (2008) Sure independence screening for ultrahigh dimensional feature space. J Roy Stat Soc B 70(5):849–911
Fan JQ, Song R (2010) Sure independence screening in generalized linear models with np-dimensionality. Ann Stat 38:3567–3604
Fan JQ, Samworth R, Wu YC (2009) Ultrahigh dimensional feature selection: beyond the linear model. J Mach Learn Res 10:1829–1853
Fan JQ, Feng Y, Song R (2011) Nonparametric independence screening in sparse ultra-high dimensional additive models. J Am Stat Assoc 106:544–557
Fu L, Wang YG (2012) Quantile regression for longitudinal data with a working correlation model. Comput Stat Data Anal 56:2526–2538
Hansen PL (1982) Large sample properties of generalized method of moments estimators. Econometrica 50:1029–1054
He XM, Wang L, Hong HG (2013) Quantile-adaptive model-free variable screening for high-dimensional heterogeneous data. Ann Stat 41:342–369
Kaslow R, Ostrow D, Detels R et al (1987) The multicenter aids cohort study: rationale, organization and selected characteristics of the participants. Am J Epidemiol 126:310–318
Lai P, Liang WJ, Wang FJ et al (2020) Feature screening of quadratic inference functions for ultrahigh dimensional longitudinal data. J Stat Comput Simul 90:2614–2630
Leng CL, Zhang WP (2014) Smoothing combined estimating equations in quantile regression for longitudinal data. Stat Comput 24:123–136
Li GR, Yang YP (2015) Semiparametric Models with Longitudinal Data (in Chinese). Science Press, Beijing
Li GR, Peng H, Zhang J et al (2012) Robust rank correlation based screening. Ann Stat 40:1846–1877
Li RZ, Zhong W, Zhu LP (2012) Feature screening via distance correlation learning. J Am Stat Assoc 107:1129–1139
Li YJ, Li GR, Lian H et al (2017) Profile forward regression screening for ultra-high dimensional semiparametric varying coefficient partially linear models. J Multivar Anal 155:133–150
Liang K, Zeger S (1986) Longitudinal data analysis using generalized linear models. Biometrika 73:13–22
Lin L, Sun J, Zhu LX (2013) Nonparametric feature screening. Comput Stat Data Anal 67:162–174
Liu JY (2016) Feature screening and variable selection for partially linear models with ultrahigh-dimensional longitudinal data. Neurocomputing 195:202–210
Ma SJ, Li RZ, Tsai CL (2017) Variable screening via quantile partial correlation. J Am Stat Assoc 112:650–663
Niu Y, Zhang RQ, Liu JC et al (2018) Nonparametric independence screening for ultra-high-dimensional longitudinal data under additive models. J Nonparametric Stat 30:884–905
Pang NW, Xia XC (2024) Distributed conditional feature screening via Pearson partial correlation with FDR control pp 1–43. arXiv:2403.05792v1
Qu A, Li RZ (2006) Quadratic inference functions for varying-coefficient models with longitudinal data. Biometrics 62:379–391
Qu A, Lindsay BG (2003) Building adaptive estimating equations when inverse of covariance estimation is difficult. J R Stat Soc B 65:127–142
Qu A, Song PXK (2004) Assessing robustness of generalised estimating equations and quadratic inference functions. Biometrika 91:447–459
Qu A, Lindsay BG, Li B (2000) Improving generalised estimating equations using quadratic inference functions. Biometrika 87:823–836
Song PXK, Jiang ZC, Park E et al (2009) Quadratic inference functions in marginal models for longitudinal data. Stat Med 26:3683–3696
Song R, Yi F, Zou H (2014) On varying-coefficient independence screening for high-dimensional varying-coefficient models. Stat Sin 24:1735–1752
Tibshirani R (1996) Regression shrinkage and selection via the lasso. J R Stat Soc B 58(1):267–288
Tong ZX, Cai ZR, Yang SS et al (2022) Model-free conditional feature screening with fdr control. J Am Stat Assoc 118(544):2575–2587
Wang LM, Li XX, Wang XQ et al (2022) Unified mean-variance feature screening for ultrahigh-dimensional regression. Comput Stat 37:1887–1918
Xia LL, Tang NS (2023) Feature screening via distance correlation for ultrahigh dimensional data with responses missing at random. Stat Sin 33:1169–1191
Xu PR, Zhu LX, Li Y (2014) Ultrahigh dimensional time course feature selection. Biometrics 70:356–365
Xue L, Qu A, Zhou JH (2010) Consistent model selection for marginal generalized additive model for correlated data. J Am Stat Assoc 105:1518–1530
Zhang JY, Zhang RQ, Lu ZP (2016) Quantile-adaptive variable screening in ultra-high dimensional varying coefficient models. J Appl Stat 43:643–654
Zhang S, Zhao PX, Li GR et al (2019) Nonparametric independence screening for ultra-high dimensional generalized varying coefficient models with longitudinal data. J Multivar Anal 171:37–52
Zhao WH, Lian H, Song XY (2017) Composite quantile regression for correlated data. Comput Stat Data Anal 109:15–33
Zhao WH, Zhang WP, Lian H (2020) Marginal quantile regression for varying coefficient models with longitudinal data. Ann Inst Stat Math 72:213–234
Zhou JH, Qu A (2012) Informative estimation and selection of correlation structure for longitudinal data. J Am Stat Assoc 107:701–710
Zhu LP, Li LX, Li RZ et al (2011) Model-free feature screening for ultrahigh-dimensional data. J Am Stat Assoc 106:1464–1475
Zhu ZY, Liu JC, Zhang RQ (2026) A novel martingale difference correlation via data splitting with applications in feature screening. J Multivar Anal 211:105508
The authors want to sincerely thank the Editor, an Associate Editor, and two referees for their thoughtful suggestions and detailed feedback, which led to substantial improvements in the presentation and overall quality of the manuscript.
The research was supported by the National Natural Science Foundation of China (Nos. 12271046 and 12201306)
the National Statistical Science Research Project of China (No. 2024LY088)
the Priority Academic Program Development of Jiangsu Higher Education Institutions (Statistics)
the Open Project of Joint Lab for Statistics and Finance of NAU (Nos. 2025JLSF306 and 2025JLSF320).
School of Statistics and Data Science, Nanjing Audit University, Nanjing, 211815, China
Lili Yue & Xiong Cai
Joint Lab for Statistics and Finance, Nanjing Audit University, Nanjing, 211815, China
Lili Yue & Xiong Cai
College of Business and Economics, California State University, Fullerton, 90802, USA
Daoji Li
School of Statistics, Beijing Normal University, Beijing, 100875, China
Gaorong Li
Authors
Correspondence to Gaorong Li.
The authors declare that they have no conflict of interest.
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
In this Appendix, we shall give the proof of the Theorem 1. Review the definitions of \(\hat{f}_k\) and \(f_k\), we have
$$\begin{aligned} \Vert \hat{f}_k\Vert ^2=&\mathbb {P}_nX_{(k)}^2\hat{\beta }_k^2+\{F^{-1}_Y(\tau )\}^2-2F^{-1}_Y(\tau )\cdot \mathbb {P}_nX_{(k)}\hat{\beta }_k,\\ \Vert f_k\Vert ^2=&EX_{(k)}^2\beta _{0k}^2+\{Q_\tau (Y)\}^2-2Q_\tau (Y)\cdot EX_{(k)}\beta _{0k}, \end{aligned}$$
where \(\mathbb {P}_nX_{(k)}^2=n^{-1}\sum _{i=1}^nm_i^{-1}\sum _{j=1}^{m_i}X_{ijk}^2\) and \(\mathbb {P}_nX_{(k)}=n^{-1}\sum _{i=1}^nm_i^{-1}\sum _{j=1}^{m_i}X_{ijk}\). Let \(D_n=\mathbb {P}_nX_{(k)}^2-EX_{(k)}^2\). Then,
$$\begin{aligned} \Vert \hat{f}_k\Vert ^2-\Vert f_k\Vert ^2=&(\hat{\beta }_k-\beta _{0k})^2\mathbb {P}_nX_{(k)}^2+2(\hat{\beta }_k-\beta _{0k})\mathbb {P}_nX_{(k)}^2\beta _{0k}+\beta _{0k}^2D_n+[\{F^{-1}_Y(\tau )\}^2\\ &-\{Q_\tau (Y)\}^2]\\&-2F^{-1}_Y(\tau )\{\mathbb {P}_nX_{(k)}\hat{\beta }_k-EX_{(k)}\beta _{0k}\}-2\{F^{-1}_Y(\tau )-Q_\tau (Y)\}EX_{(k)}\beta _{0k}\\ =&(\hat{\beta }_k-\beta _{0k})^2D_n+(\hat{\beta }_k-\beta _{0k})^2EX_{(k)}^2 \\ &+2(\hat{\beta }_k-\beta _{0k})D_n\beta _{0k}+2(\hat{\beta }_k-\beta _{0k})EX_{(k)}^2\beta _{0k}\\&+\beta _{0k}^2D_n+[\{F^{-1}_Y(\tau )\}^2-\{Q_\tau (Y)\}^2]-2F^{-1}_Y(\tau )\mathbb {P}_nX_{(k)}(\hat{\beta }_k-\beta _{0k})\\&-2F^{-1}_Y(\tau )\{\mathbb {P}_nX_{(k)}-EX_{(k)}\}\beta _{0k}-2\{F^{-1}_Y(\tau )-Q_\tau (Y)\}EX_{(k)}\beta _{0k}\\ \triangleq&S_{k1}+S_{k2}+S_{k3}+S_{k4}+S_{k5}+S_{k6}+S_{k7}+S_{k8}+S_{k9}. \end{aligned}$$
Similar with the discuss of He et al. (2013), \(|S_{k6}|=O(n^{-1/2}(\log n)^{1/2})=o(n^{-\xi })\), \(|S_{k9}|=O(n^{-1/2}(\log n)^{1/2})=o(n^{-\xi })\) using conditions (C2), (C4), (C5) and \(\textrm{max}_{i}\{m_i\}=m_a<\infty \). Then, for sufficiently large n, we have
$$\begin{aligned}&P\big (\big |\Vert \hat{f}_k\Vert ^2-\Vert f_k\Vert ^2\big |\ge \epsilon \big ) \nonumber \\&\le P(|S_{k1}|+|S_{k2}|+|S_{k3}|+|S_{k4}|+|S_{k5}|+|S_{k7}|+|S_{k8}|\ge \epsilon )\nonumber \\&=P(|S_{k1}|+|S_{k2}|+|S_{k3}|+|S_{k4}|+|S_{k5}|+|S_{k7}|+|S_{k8}|\ge \epsilon ,|D_n|\ge \epsilon _1)\nonumber \\&+P(|S_{k1}|+|S_{k2}|+|S_{k3}|+|S_{k4}|+|S_{k5}|+|S_{k7}|+|S_{k8}|\ge \epsilon ,|D_n|<\epsilon _1)\nonumber \\&\le P(|S_{k1}|+|S_{k2}|+|S_{k3}|+|S_{k4}|+|S_{k5}|+|S_{k7}|+|S_{k8}|\ge \epsilon ,|D_n|<\epsilon _1)\nonumber \\&+P(|D_n|\ge \epsilon _1). \end{aligned}$$
(A.1)
For the second part of the inequality (Appendix A.1), we have
$$\begin{aligned} & P(|D_n|\ge \epsilon _1)\le P\left( \sum _{j=1}^{m_a}\Big |\dfrac{1}{n}\sum _{i=1}^nX_{ijk}^2-EX_{ijk}^2\Big |\ge \epsilon _1\right) \\ & \le m_aP\left( \Big |\dfrac{1}{n}\sum _{i=1}^nX_{ijk}^2-EX_{ijk}^2\Big |\ge \epsilon _1/m_a\right) . \end{aligned}$$
Let \(\epsilon _1=m_a\delta _1/n\). Combine Bernstein’s inequality, we have
$$\begin{aligned} P(|D_n|\ge m_a\delta _1/n)\le 2m_a\exp \left( -\dfrac{\delta _1^2}{4\sum _{i=1}^nDX_{ijk}^2+2c_3\delta _1}\right) , \end{aligned}$$
(A.2)
where \(c_3\) is a positive constant.
By the conditions (C2)–(C5), the first part of the inequality (A.1) can be expressed as
$$\begin{aligned}&P(|S_{k1}|+|S_{k2}|+|S_{k3}|+|S_{k4}|+|S_{k5}|+|S_{k7}|+|S_{k8}|\ge \epsilon ,|D_n|<\epsilon _1)\\&\hspace{20pt}\le P\big ((\hat{\beta }_k-\beta _{0k})^2\epsilon _1+(\hat{\beta }_k-\beta _{0k})^2C_1+2\epsilon _1|\hat{\beta }_k-\beta _{0k}||\beta _{0k}|+2C_1|\hat{\beta }_k-\beta _{0k}||\beta _{0k}|\\&\hspace{30pt}+\beta _{0k}^2|D_n|+2C_2(C_1+\epsilon _1)^{1/2}|\hat{\beta }_k-\beta _{0k}| +2|F^{-1}_Y(\tau )||\{\mathbb {P}_nX_{(k)}-EX_{(k)}\}||\beta _{0k}|\ge \epsilon \big )\\&\hspace{20pt}\le P\big (\beta _{0k}^2|D_n|+2|F^{-1}_Y(\tau )||\mathbb {P}_nX_{(k)}-EX_{(k)}||\beta _{0k}|\ge \epsilon /2\big ), \end{aligned}$$
the last inequality is derived using the \(\sqrt{n}\)-consistency of QIF estimator \(\hat{\beta }_k\), \(C_1\) and \(C_2\) are some positive constants. Note that,
$$\begin{aligned} P(\beta _{0k}^2|D_n|\ge \epsilon /4)\le&P(|D_n|\ge C_3\epsilon /4)\\ \le&P\left( \sum _{j=1}^{m_a}\Big |\dfrac{1}{n}\sum _{i=1}^nX_{ijk}^2-EX_{ijk}^2\Big |\ge C_3\epsilon /4\right) \\ \le&m_aP\left( \Big |\dfrac{1}{n}\sum _{i=1}^nX_{ijk}^2-EX_{ijk}^2\Big |\ge C_3\epsilon /4m_a\right) ,\end{aligned}$$
where \(C_3\) is a positive constant. Let \(\epsilon =4m_a\delta _2/n\), and combine Bernstein’s inequality, we have
$$\begin{aligned} P(\beta _{0k}^2|D_n|\ge m_a\delta _2/n)\le 2m_a\exp \left( -\dfrac{C_3^2\delta _2^2}{4\sum _{i=1}^nDX_{ijk}^2+2c_4C_3\delta _2}\right) , \end{aligned}$$
(A.3)
where \(c_4\) is a positive constant. Similarly, we have
$$\begin{aligned} P\big (2|F^{-1}_Y(\tau )||\mathbb {P}_nX_{(k)}-EX_{(k)}||\beta _{0k}|\ge \epsilon /4\big ) \le&m_aP\left( \Big |\dfrac{1}{n}\sum _{i=1}^nX_{ijk}-EX_{ijk}\Big |\ge C_4\epsilon /4m_a\right) \nonumber \\ \le&2m_a\exp \left( -\dfrac{C_4^2\delta _2^2}{4\sum _{i=1}^nDX_{ijk}+2c_5C_4\delta _2}\right) , \end{aligned}$$
(A.4)
where \(c_5\) and \(C_4\) are some positive constants.
$$\begin{aligned} P\big (\big |\Vert \hat{f}_k\Vert ^2-\Vert f_k\Vert ^2\big |\ge 4m_a\delta _2/n\big )\le&P(|D_n|\ge m_a\delta _1/n)+P(\beta _{0k}^2|D_n|\ge m_a\delta _2/n)\\&+P\big (2|F^{-1}_Y(\tau )||\mathbb {P}_nX_{(k)}-EX_{(k)}||\beta _{0k}|\ge m_a\delta _2/n\big )\\ \le&2m_a\exp \left( -\dfrac{\delta _1^2}{4\sum _{i=1}^nDX_{ijk}^2+2c_3\delta _1}\right) \\&+2m_a\exp \left( -\dfrac{C_3^2\delta _2^2}{4\sum _{i=1}^nDX_{ijk}^2+2c_4C_3\delta _2}\right) \\&+2m_a\exp \left( -\dfrac{C_4^2\delta _2^2}{4\sum _{i=1}^nDX_{ijk}+2c_5C_4\delta _2}\right) . \end{aligned}$$
Let \(4m_a\delta _2/n=m_a\delta _1/n=Cn^{-\xi }\) for any given constant \(C>0\), that is, \(\delta _2=Cn^{1-\xi }/(4m_a)\) and \(\delta _1=Cn^{1-\xi }/m_a\). Then,
$$\begin{aligned} P\Big (\mathop \textrm{max}\limits _{1\le k\le p}|\Vert \hat{f}_k\Vert ^2-\Vert f_k\Vert ^2|\ge Cn^{-\xi } \Big )\le 6m_ap\exp (-c_2n^{1-2\xi }). \end{aligned}$$
Similar with the Lemma 3.1 in He et al. (2013), there is a given constant \(c_1>0\) such that \(\textrm{min}_{k\in \mathcal{M}_\tau }\Vert f_k\Vert ^2\ge c_1n^{-\xi }/8\). Take the threshold value \(\nu _n=\delta n^{-\xi }\) with \(\delta <c_1/16\), we have
$$\begin{aligned} P\Big (\mathcal{M}_\tau \subset \widehat{\mathcal{M}}_\tau \Big )\ge&P\Big (\mathop \textrm{min}\limits _{k\in \mathcal{M}_\tau }\Vert \hat{f}_k\Vert ^2\ge \nu _n\Big )\\ \ge&P\Big (\mathop \textrm{min}\limits _{k\in \mathcal{M}_\tau }\Vert f_k\Vert ^2- \mathop \textrm{max}\limits _{k\in \mathcal{M}_\tau }\big |\Vert \hat{f}_k\Vert ^2-\Vert f_k\Vert ^2\big | \ge \nu _n\Big )\\ =&1-P\Big (\mathop \textrm{max}\limits _{k\in \mathcal{M}_\tau }\big |\Vert \hat{f}_k\Vert ^2-\Vert f_k\Vert ^2\big | \ge \mathop \textrm{min}\limits _{k\in \mathcal{M}_\tau }\Vert f_k\Vert ^2-\nu _n\Big )\\ \ge&1-P\Big (\mathop \textrm{max}\limits _{k\in \mathcal{M}_\tau }\big |\Vert \hat{f}_k\Vert ^2-\Vert f_k\Vert ^2\big | \ge c_1n^{-\xi }/16\Big ). \end{aligned}$$
Then, the proof of Theorem 1 is finished. \(\square \)
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
Yue, L., Li, D., Cai, X. et al. Quantile adaptive feature screening for ultra-high dimensional longitudinal heterogeneous data. Comput Stat 41, 109 (2026). https://doi.org/10.1007/s00180-026-01789-5
Received: 18 December 2025
Accepted: 15 July 2026
Published: 24 July 2026
Version of record: 24 July 2026
DOI: https://doi.org/10.1007/s00180-026-01789-5