开放资源库 · 研究数据集

Compiled Data for "Parsimonious Subset Selection for Generalized Linear Models with Biomedical Applications"

This deposit contains the compiled, model-ready datasets used in the two real-data applications presented in the Biometrics manuscript “Parsimonious Subset Selection for Generalized Linear Models with Biomedical Applications” (manuscript BIOM2026208M). The rice dataset contains the processed genome-wide association study design matrix for 1,155 rice accessions, comprising 158,210 SNP predictors, together with the binary grain-length response used in the logistic regression analysis. The underlying rice data are publicly available through NCBI GEO (GSE71553), BioProject (PRJNA291537), and dbSNP (batch ID 1062024). The Khan small round blue cell tumor dataset contains the established training and independent test partitions used in the multinomial classification analysis. The training data contain 63 observations and the test data contain 20 observations, with 2,308 gene-expression predictors and one response variable. These publicly available partitions are also distributed through the ISLR2 R package. README files describe the contents, processing, provenance, checksums, and instructions for reading the deposited files. Analysis code and additional reproduction instructions accompany the associated manuscript.

← 返回资源筛选结果
RESOURCE OVERVIEW

资源说明

This deposit contains the compiled, model-ready datasets used in the two real-data applications presented in the Biometrics manuscript “Parsimonious Subset Selection for Generalized Linear Models with Biomedical Applications” (manuscript BIOM2026208M). The rice dataset contains the processed genome-wide association study design matrix for 1,155 rice accessions, comprising 158,210 SNP predictors, together with the binary grain-length response used in the logistic regression analysis. The underlying rice data are publicly available through NCBI GEO (GSE71553), BioProject (PRJNA291537), and dbSNP (batch ID 1062024). The Khan small round blue cell tumor dataset contains the established training and independent test partitions used in the multinomial classification analysis. The training data contain 63 observations and the test data contain 20 observations, with 2,308 gene-expression predictors and one response variable. These publicly available partitions are also distributed through the ISLR2 R package. README files describe the contents, processing, provenance, checksums, and instructions for reading the deposited files. Analysis code and additional reproduction instructions accompany the associated manuscript.

聚变工程工程验证
RESEARCH USE PROFILE

研究使用指引

适用任务诊断分析、模型校准、代理训练、跨装置比较与基准测试
使用准备核对字段、单位、缺失值、采样条件、训练测试划分和许可
核验重点需要检查数据泄漏、分布偏移以及装置和工况适用范围
RESOURCE FEEDBACK

资源信息需要更新?

可报告链接失效、文件异常、信息错误或版本变化,我们会核对并更新资源记录。