遇见数据集

Models and Predictions for "The Proper Care and Feeding of CAMELS: How Limited Training Data Affects Streamflow Prediction"

收藏
Zenodo2020-07-30 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

<strong>Models and Predictions</strong> This dataset contains the trained XGBoost and EA-LSTM models and the models' predictions for the paper <em>The Proper Care and Feeding of CAMELS: How Limited Training Data Affects Streamflow Prediction</em>. For each input sequence length (10, 30, 100, 270*, 365*) and each combination of model (XGBoost, EA-LSTM), training years (3, 6, 9), number of basins (13, 26, 53, 265, 531), and seed (111-888), there are five folders. Each corresponds to a random basin sample (for 531 basins there's only one folder, since it's all basins).<br> In each folder, there are three files: \(\texttt{model.pkl}\) (XGBoost) or <em>\(\texttt{model_epoch30.pt}\)</em> (EA-LSTM), which stores the pickled trained model <em>\(\texttt{xgboost_seedNNN.p}\)</em> or <em>\(\texttt{ealstm_seedNNN.p}\)</em>, which stores a pickled dictionary that maps each basin to the DataFrame of predicted and actual daily streamflow. \(\texttt{attributes.db}\), which stores static catchment attributes needed for inference. In addition to each folder, there is a SLURM submission script called <em>\(\texttt{&lt;foldername&gt;.sbatch}\)</em> that was used to create and evaluate the model in the folder. * sequence lengths 270 and 365 only contain data for EA-LSTM.

「模型与预测结果」 本数据集包含为论文《CAMELS的合理养护与投喂:有限训练数据如何影响径流预测》(The Proper Care and Feeding of CAMELS: How Limited Training Data Affects Streamflow Prediction)训练得到的极端梯度提升树(XGBoost)与进化长短期记忆网络(EA-LSTM)模型,以及这些模型的预测结果。针对每一种输入序列长度(10、30、100、270*、365*)、每一组模型(XGBoost、EA-LSTM)、训练年限(3、6、9)、流域数量(13、26、53、265、531)以及随机种子(111-888)的组合,均对应5个文件夹。每个文件夹对应一组随机抽取的流域样本(当流域数量为531时仅存在1个文件夹,因为此时包含全部流域)。 每个文件夹内包含三个文件:`model.pkl`(适用于XGBoost)或`model_epoch30.pt`(适用于EA-LSTM),用于存储序列化后的训练模型;`xgboost_seedNNN.p`(XGBoost场景下)或`ealstm_seedNNN.p`(EA-LSTM场景下),用于存储序列化字典,该字典将每个流域映射至包含预测与实测日径流的数据框(DataFrame);以及`attributes.db`,用于存储推理所需的静态集水区属性。除上述文件夹外,还附带一个名为`<foldername>.sbatch`的SLURM提交脚本,用于创建并评估对应文件夹内的模型。 *注:序列长度270与365仅包含EA-LSTM的相关数据。

提供机构:
Zenodo
创建时间:
2019-11-17
二维码
社区交流群
二维码
科研交流群
商业服务