CAMELSH: A Large-Sample Hourly Hydrometeorological Dataset and Attributes at Watershed-Scale for Contiguous United States
收藏资源简介:
We present CAMELSH (Catchment Attributes and Hourly HydroMeteorology for Large-Sample Studies), the first hourly large-sample hydrometeorological dataset for the contiguous United States. CAMELSH provides hourly streamflow and meteorological time series, along with catchment attributes and boundaries from GAGES-II and HydroATLAS, covering 7,211 catchments across diverse climatic, hydrological, and anthropogenic conditions. Among these, 2,569 catchments include observed streamflow, while the remaining contain only meteorological forcings. The dataset spans 45 years (1980–2024) with 11 meteorological variables from the NLDAS-2 forcing dataset, from which we compute nine climate indices related to precipitation, evapotranspiration, seasonality, and snow fraction. Additionally, CAMELSH includes two sets of catchment attributes: 439 from GAGES-II and 195 derived from HydroATLAS. Developed in accordance with FAIR (Findability, Accessibility, Interoperability, and Reusability) principles, CAMELSH is the first large-sample dataset at an hourly timescale, supporting machine learning applications for short-term streamflow (flood) prediction and advancing data-driven hydrological research across multiple timescales. The dataset is organized into the following subfolders: • The attributes folder comprises three subfolders, each corresponding to a source dataset, and contains CSV (comma-separated values) files that store basin attributes. Specifically, the ‘Attributes-Climate’ subfolder includes the ‘attribute_climate.csv’ file, which contains nine climate attributes (Table 2) derived from NLDAS-2 data. The ‘Attributes-GAGESII’ subfolder contains a total of 26 CSV files and one Excel (*.xlsx) file. Among these, the 26 CSV files include information related to basins/streamflow gauging stations and 439 basin attributes. The name of each file represents a distinct group of attributes, as described in Table S.1. The remaining file, named ‘Var_description.xlsx’, provides explanatory details regarding the variable names included in the 26 CSV files, with information similar to that presented in Table S.1. The final subfolder, ‘Attributes-HydroAtlas’, contains a single file, ‘attribute_HydroATLAS.csv’, which includes 195 basin attributes derived from the HydroATLAS dataset. The first column in all CSV files, labeled ‘STAID,’ contains the identification (ID) names of the streamgages. These IDs are assigned by the USGS and are sourced from the original GAGES-II dataset. • The shapefiles folder contains two sets of shapefiles for the catchment boundary. The first set, CAMELSH_shapefile.shp, is derived from the original GAGES-II dataset and is used to obtain the corresponding climate forcing data for each catchment. The second set, CAMELSH_shapefile_hydroATLAS.shp, includes catchment boundaries derived from the HydroATLAS dataset. Each polygon in both shapefiles contains a field named GAGE_ID, which represents the ID of the stream gauges. • The timeseries (7zip) file contains a compressed archive (7zip) that includes time series data for 7,211 basins. Within this 7zip file, there are a total of 7,211 NetCDF files, each corresponding to a specific basin. The name of each NetCDF file matches the stream gauge ID. Each file contains an hourly time series from 1980-01-01 00:00:00 to 2024-12-31 23:00:00 for streamflow (denoted as “Streamflow” in the NetCDF file) and 11 climate variables (see Table 1). The streamflow data series includes missing values, which are represented as “NaN”. Among these, a total of 2,569 stations contain observed streamflow data, while 4,642 stations do not have observed data. Note that streamflow is measured in cubic meters per second. All meteorological forcing data and streamflow records have been standardized to the +0 UTC time zone. • The info.csv file, located in the main directory of dataset, contains basic information for 7,211 stream stations. This includes the stream gauge ID, the total number of observed hourly data points over 45 years (from 1980 to 2024), and the number of observed hourly data points for each individual year from 1980 to 2024. Stations with and without observed data are distinguished by the value in the second column, where stations without observed streamflow data have a corresponding value of 0.
本研究介绍了CAMELSH(Catchment Attributes and Hourly HydroMeteorology for Large-Sample Studies,面向大样本研究的流域属性与小时尺度水文气象数据集),这是首个针对美国本土的小时尺度大样本水文气象数据集。CAMELSH提供小时尺度径流与气象时间序列,同时包含源自GAGES-II与HydroATLAS的流域属性数据与流域边界,覆盖全美7211个具有多样气候、水文及人为扰动条件的流域。其中2569个流域带有实测径流数据,其余仅包含气象强迫数据。 该数据集时间跨度为45年(1980–2024年),包含11个源自NLDAS-2强迫数据集的气象变量,研究团队基于这些变量计算得到9项与降水、蒸散发、季节性及积雪占比相关的气候指数。此外,CAMELSH包含两套流域属性数据集:439项源自GAGES-II,195项源自HydroATLAS。 本数据集遵循FAIR(可发现性、可访问性、互操作性、可复用性)原则构建,是首个小时尺度的大样本水文数据集,可支持短期径流(洪水)预测的机器学习应用,并推动多时间尺度数据驱动水文研究的发展。 数据集按以下子文件夹组织: • 属性文件夹包含三个子文件夹,分别对应不同的源数据集,内部存储以逗号分隔值(CSV)格式保存的流域属性文件。具体而言,“Attributes-Climate”子文件夹包含`attribute_climate.csv`文件,其中存储了9项由NLDAS-2数据衍生得到的气候属性(见表2)。“Attributes-GAGESII”子文件夹共包含26个CSV文件与1个Excel(*.xlsx)文件。其中26个CSV文件涵盖流域/径流监测站相关信息与439项流域属性,每个文件的名称对应一类独立的属性组,详见表S.1。剩余的`Var_description.xlsx`文件则对26个CSV文件中的变量名称进行了解释说明,信息与表S.1一致。最后一个子文件夹“Attributes-HydroAtlas”仅包含一个文件`attribute_HydroATLAS.csv`,其中存储了195项源自HydroATLAS数据集的流域属性。所有CSV文件的第一列均标注为“STAID”,存储由美国地质调查局(USGS)分配的径流监测站标识ID,该ID源自原始GAGES-II数据集。 • Shapefile文件夹包含两套流域边界矢量图层。第一套为`CAMELSH_shapefile.shp`,源自原始GAGES-II数据集,用于获取每个流域对应的气象强迫数据。第二套为`CAMELSH_shapefile_hydroATLAS.shp`,包含源自HydroATLAS数据集的流域边界。两套矢量图层中的每个多边形均包含名为`GAGE_ID`的字段,用于标识对应的径流监测站。 • 时间序列压缩包子文件夹包含一个7zip格式的压缩归档,内含7211个流域的时间序列数据。该压缩包中共包含7211个网络通用数据格式(NetCDF)文件,每个文件对应一个特定流域,文件名与径流监测站ID一致。每个文件包含1980年1月1日00:00:00至2024年12月31日23:00:00的小时尺度时间序列,涵盖径流(NetCDF文件中记为"Streamflow")与11项气候变量(详见表1)。径流数据序列包含缺失值,以“NaN”表示。其中共计2569个监测站带有实测径流数据,剩余4642个监测站无实测数据。请注意,径流的单位为立方米每秒。所有气象强迫数据与径流记录均已标准化至协调世界时(UTC)+0时区。 • 数据集主目录下的`info.csv`文件包含7211个径流监测站的基础信息,涵盖监测站ID、45年(1980–2024年)间的总实测小时数据点数,以及1980至2024年各年度的实测小时数据点数。有无实测径流数据的监测站可通过第二列的数值区分,无实测径流数据的监测站对应值为0。



