Compiled Data for "Parsimonious Subset Selection for Generalized Linear Models with Biomedical Applications"
This deposit contains the compiled, model-ready datasets used in the two real-data applications presented in the Biometrics manuscript “Parsimonious Subset Selection for Generalized Linear Models with Biomedical Applications” (manuscript BIOM2026208M). The rice dataset contains the processed genome-wide association study design matrix for 1,155 rice accessions, comprising 158,210 SNP predictors, together with the binary grain-length response used in the logistic regression analysis. The underlying rice data are publicly available through NCBI GEO (GSE71553), BioProject (PRJNA291537), and dbSNP (batch ID 1062024). The Khan small round blue cell tumor dataset contains the established training and independent test partitions used in the multinomial classification analysis. The training data contain 63 observations and the test data contain 20 observations, with 2,308 gene-expression predictors and one response variable. These publicly available partitions are also distributed through the ISLR2 R package. README files describe the contents, processing, provenance, checksums, and instructions for reading the deposited files. Analysis code and additional reproduction instructions accompany the associated manuscript.
资源说明
This deposit contains the compiled, model-ready datasets used in the two real-data applications presented in the Biometrics manuscript “Parsimonious Subset Selection for Generalized Linear Models with Biomedical Applications” (manuscript BIOM2026208M). The rice dataset contains the processed genome-wide association study design matrix for 1,155 rice accessions, comprising 158,210 SNP predictors, together with the binary grain-length response used in the logistic regression analysis. The underlying rice data are publicly available through NCBI GEO (GSE71553), BioProject (PRJNA291537), and dbSNP (batch ID 1062024). The Khan small round blue cell tumor dataset contains the established training and independent test partitions used in the multinomial classification analysis. The training data contain 63 observations and the test data contain 20 observations, with 2,308 gene-expression predictors and one response variable. These publicly available partitions are also distributed through the ISLR2 R package. README files describe the contents, processing, provenance, checksums, and instructions for reading the deposited files. Analysis code and additional reproduction instructions accompany the associated manuscript.
研究使用指引
资源信息需要更新?
可报告链接失效、文件异常、信息错误或版本变化,我们会核对并更新资源记录。