遇见数据集

Data sample and example model for random forest photolysis emulation.

收藏
Zenodo2026-05-21 更新2026-05-26 收录
官方服务:

资源简介:

This record contains dataset samples to use to train machine learning emulators of UKCA photolysis, and a trained random forest photolysis emulator along with its own test and prediction data samples. Recommended software for working with these data: Python, numpy, scikit-learn, joblib. Files provided: 1982_45m.npy - 8.3 GB numpy training dataset of shape, 49 features x ~45 million samples, of 32-bit floats. The random forest was trained on a dataset of 182 million samples but it is too large to upload and store here. This is a random sample of it. 1982_metadata.txt - Specifications of the sample training datasets including an ordered list of their features. rf.pkl - 1.1 GB trained scikit-learn random forest regression model (64-bit) which emulates UKCA photolysis, in pickle format. Read in to Python using joblib and scikit. rf_structure.nc - The same random forest, decomposed into its constituent arrays (32-bit) and saved as a NetCDF file for use in Fortran or other languages. rf_metadata.txt - Specifications of the random forest. rf_inputs.npy - 70 MB numpy sample of 32-bit input data, of shape, 10 features x ~1 million samples, of test data for the random forest. rf_targets.npy - 181 MB numpy sample of 32-bit target output data, of shape, 26 features x ~1 million samples, of test data for the random forest, corresponding to the provided input data. rf_preds.npy - 361 MB numpy sample of 64-bit output predictions from the random forest, of shape, 26 features x ~1 million samples, corresponding to the provided test data. All of the numpy datasets are 2D, of shape features x samples, and were randomly sampled from a year (1982) of hourly AMIP-nudged UM vn13.9 output data at a spatial resolution of 144 latitudes x 192 longitudes x 85 vertical levels up to 85 km. The features of the full data are (in order): day hour relative level hybrid height latitude longitude specific humidity cloud fraction sza upward shortwave flux downward shortwave flux pressure temperature O3 col JO3 JNO2 Jnitrates JHCHOr JHCHOm JMeCOCHO JprodCO JprodOH JO2 JCl2O2 JNO3strat JO1D JOCS JSO3 JMeONO2 JNALD JISON JMeCHO->MeOO Jpropanal JNO3 JH2O JHOBr JHOCl JHNO3 JHNO4 JH2O2 JMeOOH JO2->O3P JO3 JN2O JMACR JMACROOH JMeCHO->CH4 JNO JNO2 (from Strat-trop) The features of the input test data sample are (in order): day hour relative level hybrid height latitude longitude specific humidity cloud fraction sza upward shortwave flux downward shortwave flux pressure temperature The features of the target and predicted output J-value test data sample are (in order): HCHOr HCHOm MeCOCHO prodCO prodOH O2 Cl2O2 NO3strat O1D OCS SO3 MeONO2 NALD ISON MeCHO->MeOO propanal NO3 H2O HOBr HOCl HNO3 HNO4 H2O2 MeOOH O2->O3P O3 N2O MACR MACROOH MeCHO->CH4 NO NO2 (from Strat-trop)

提供机构:
Zenodo
创建时间:
2026-05-21
二维码
社区交流群
二维码
科研交流群
商业服务