遇见数据集

NHANES Gaussian phenotype atlas

收藏
Zenodo2026-07-09 更新2026-08-01 收录
官方服务:

资源简介:

All biological measurements analyzed in this study were obtained from the National Health and Nutrition Examination Survey (NHANES) 1999–2023, which follows a complex multistage probability sampling design described previously. NHANES is a nationally representative health survey of the United States population that includes anthropometric measurements, laboratory tests, physiological examinations, imaging data, environmental exposure assessments, and questionnaire-based variables. Publicly available datasets and data dictionaries were downloaded from the Centers for Disease Control and Prevention (CDC) NHANES website. Only numerical variables representing biological measurements were included in the present study. Administrative variables, participant identifiers, dates, sampling weights, and other non-biological variables were excluded. Each numerical variable was evaluated independently. Variables were excluded if they contained fewer than three finite observations, had zero variance, or appeared to represent categorical variables. Categorical-like variables were operationally defined as variables having fewer unique values than the square root of the number of finite observations. Variables were analyzed separately for each NHANES cycle. Identical biological measurements obtained in different survey cycles were treated as independent variables when constructing the reference distribution. Because the objective of this study was to characterize the Gaussian phenotype of biological measurements rather than administrative metadata, variables describing examination procedures, time intervals, sampling weights, identifiers, or other operational information were excluded from biological interpretation. The Gaussian phenotype of each biological variable was quantified using the root mean square error of a normal quantile–quantile plot (QQ-RMSE). For each variable, non-finite values, including missing values, NaN, and infinite values, were removed. Variables were excluded if they were non-numeric, contained fewer than three finite observations, had zero variance, or appeared to represent categorical variables. Categorical-like variables were operationally defined as variables for which the number of unique values was smaller than the square root of the number of finite observations. For each remaining variable, values were standardized to z-scores by subtracting the sample mean and dividing by the sample standard deviation. The standardized values were then sorted in ascending order. The corresponding theoretical quantiles were obtained from the standard normal distribution using plotting positions generated by the ppoints function in R.

提供机构:
Zenodo
创建时间:
2026-07-09
二维码
社区交流群
二维码
科研交流群
商业服务