dazhiyang/bsrn-merra2
收藏资源简介:
--- language: en license: mit configs: - config_name: abs data_files: - split: abs path: abs/*.parquet - config_name: ale data_files: - split: ale path: ale/*.parquet - config_name: asp data_files: - split: asp path: asp/*.parquet - config_name: bar data_files: - split: bar path: bar/*.parquet - config_name: ber data_files: - split: ber path: ber/*.parquet - config_name: bil data_files: - split: bil path: bil/*.parquet - config_name: bon data_files: - split: bon path: bon/*.parquet - config_name: bos data_files: - split: bos path: bos/*.parquet - config_name: bou data_files: - split: bou path: bou/*.parquet - config_name: brb data_files: - split: brb path: brb/*.parquet - config_name: bud data_files: - split: bud path: bud/*.parquet - config_name: cab data_files: - split: cab path: cab/*.parquet - config_name: cam data_files: - split: cam path: cam/*.parquet - config_name: cap data_files: - split: cap path: cap/*.parquet - config_name: car data_files: - split: car path: car/*.parquet - config_name: clh data_files: - split: clh path: clh/*.parquet - config_name: cnr data_files: - split: cnr path: cnr/*.parquet - config_name: coc data_files: - split: coc path: coc/*.parquet - config_name: daa data_files: - split: daa path: daa/*.parquet - config_name: dar data_files: - split: dar path: dar/*.parquet - config_name: dom data_files: - split: dom path: dom/*.parquet - config_name: dra data_files: - split: dra path: dra/*.parquet - config_name: dwn data_files: - split: dwn path: dwn/*.parquet - config_name: e13 data_files: - split: e13 path: e13/*.parquet - config_name: ena data_files: - split: ena path: ena/*.parquet - config_name: eur data_files: - split: eur path: eur/*.parquet - config_name: flo data_files: - split: flo path: flo/*.parquet - config_name: fpe data_files: - split: fpe path: fpe/*.parquet - config_name: fua data_files: - split: fua path: fua/*.parquet - config_name: gan data_files: - split: gan path: gan/*.parquet - config_name: gcr data_files: - split: gcr path: gcr/*.parquet - config_name: gim data_files: - split: gim path: gim/*.parquet - config_name: gob data_files: - split: gob path: gob/*.parquet - config_name: gur data_files: - split: gur path: gur/*.parquet - config_name: gvn data_files: - split: gvn path: gvn/*.parquet - config_name: how data_files: - split: how path: how/*.parquet - config_name: ilo data_files: - split: ilo path: ilo/*.parquet - config_name: ino data_files: - split: ino path: ino/*.parquet - config_name: ish data_files: - split: ish path: ish/*.parquet - config_name: iza data_files: - split: iza path: iza/*.parquet - config_name: kwa data_files: - split: kwa path: kwa/*.parquet - config_name: lau data_files: - split: lau path: lau/*.parquet - config_name: ler data_files: - split: ler path: ler/*.parquet - config_name: lin data_files: - split: lin path: lin/*.parquet - config_name: lmp data_files: - split: lmp path: lmp/*.parquet - config_name: lrc data_files: - split: lrc path: lrc/*.parquet - config_name: lyu data_files: - split: lyu path: lyu/*.parquet - config_name: man data_files: - split: man path: man/*.parquet - config_name: mnm data_files: - split: mnm path: mnm/*.parquet - config_name: nau data_files: - split: nau path: nau/*.parquet - config_name: new data_files: - split: new path: new/*.parquet - config_name: nya data_files: - split: nya path: nya/*.parquet - config_name: ohy data_files: - split: ohy path: ohy/*.parquet - config_name: pal data_files: - split: pal path: pal/*.parquet - config_name: par data_files: - split: par path: par/*.parquet - config_name: pay data_files: - split: pay path: pay/*.parquet - config_name: psu data_files: - split: psu path: psu/*.parquet - config_name: ptr data_files: - split: ptr path: ptr/*.parquet - config_name: qiq data_files: - split: qiq path: qiq/*.parquet - config_name: reg data_files: - split: reg path: reg/*.parquet - config_name: rlm data_files: - split: rlm path: rlm/*.parquet - config_name: run data_files: - split: run path: run/*.parquet - config_name: sap data_files: - split: sap path: sap/*.parquet - config_name: sbo data_files: - split: sbo path: sbo/*.parquet - config_name: sel data_files: - split: sel path: sel/*.parquet - config_name: sms data_files: - split: sms path: sms/*.parquet - config_name: son data_files: - split: son path: son/*.parquet - config_name: sov data_files: - split: sov path: sov/*.parquet - config_name: spo data_files: - split: spo path: spo/*.parquet - config_name: sxf data_files: - split: sxf path: sxf/*.parquet - config_name: syo data_files: - split: syo path: syo/*.parquet - config_name: tam data_files: - split: tam path: tam/*.parquet - config_name: tat data_files: - split: tat path: tat/*.parquet - config_name: tik data_files: - split: tik path: tik/*.parquet - config_name: tir data_files: - split: tir path: tir/*.parquet - config_name: tor data_files: - split: tor path: tor/*.parquet - config_name: xia data_files: - split: xia path: xia/*.parquet - config_name: yus data_files: - split: yus path: yus/*.parquet tags: - solar - radiation - bsrn pretty_name: BSRN MERRA-2 Atmospheric Inputs --- <!-- This file is uploaded as the Hugging Face dataset README; keep content user-facing (no internal maintainer or upload-only notes). --> # BSRN MERRA-2 Atmospheric Inputs Point-extracted MERRA-2 reanalysis data for [Baseline Surface Radiation Network (BSRN)](https://bsrn.awi.de/) stations. These parquet files provide atmospheric and aerosol inputs for the **REST2** clear-sky radiation model [2]. ## Dataset Description Each file contains hourly MERRA-2 variables [1] at a single BSRN station location. **Extraction is performed via Google Earth Engine (GEE)** from NASA's 0.5° × 0.625° global grid. Data are aligned to the MERRA-2 grid cell nearest the station coordinates. ### File Structure Station folders use the **lowercase** BSRN three-letter code (e.g. `ber`, `qiq`, or `spo`), with the exception of `e13`. ``` {station}/ {station}{MM}{YY}_merra2.parquet # One file per month ``` Examples: - `qiq/qiq0124_merra2.parquet` — QIQ, January 2024 - `ber/ber0325_merra2.parquet` — BER, March 2025 ### Variables | Column | Description | MERRA-2 Source | Units (raw) | |--------|--------------------------------------------------|----------------|---------------| | AOD55 | Aerosol optical depth at 550 nm | TOTEXTTAU | dimensionless | | ALPHA | Ångström exponent | TOTANGSTR | dimensionless | | ALBEDO | Surface albedo | ALBEDO | [0–1] | | TQV | Total column precipitable water vapor | TQV | kg/m² | | TO3 | Total column ozone | TO3 | Dobson | | PS | Surface pressure | PS | Pa | - **Index**: UTC `DatetimeIndex` (hourly, MERRA-2 native resolution). - **Time coverage**: MERRA-2 spans 1980–present; files are generated only for months with BSRN station-to-archive data on the FTP. ### Use with REST2 These parquet files are designed for the REST2 clear-sky model [2]. The `bsrn` Python package fetches MERRA-2 from this dataset **into RAM** (no disk cache) and provides: - `fetch_rest2(index, station_code)` — fetch from HF into RAM, reindex to 1-min target, interpolate, derive BETA, and convert units for REST2 - Raw parquet: use `pandas.read_parquet` on a path from `huggingface_hub` (see below) REST2 expects: **PS** (hPa), **ALBEDO**, **ALPHA**, **BETA** (derived from AOD55 and ALPHA), **TO3** (atm·cm), **TQV** (atm·cm). **Conversion tips** (raw → REST2): | Variable | Raw unit | REST2 unit | Conversion | |----------|---------------|---------------|-------------------------------------------| | PS | Pa | hPa | ÷ 100 | | ALBEDO | [0–1] | [0–1] | no conversion | | ALPHA | dimensionless | dimensionless | no conversion | | BETA | — | — | AOD55 × 0.55^ALPHA (use 0.001 if AOD55=0) | | TO3 | Dobson | atm·cm | ÷ 1000 | | TQV | kg/m² | atm·cm | ÷ 10 | ## Usage ### Load from Hugging Face ```python from huggingface_hub import hf_hub_download import pandas as pd # Download a single file; path is {station}/{station}{MM}{YY}_merra2.parquet path = hf_hub_download( repo_id="dazhiyang/bsrn-merra2", filename="qiq/qiq0124_merra2.parquet", repo_type="dataset", ) df = pd.read_parquet(path) # df has DatetimeIndex (UTC) and columns: AOD55, ALPHA, ALBEDO, TQV, TO3, PS ``` ### Use with bsrn package The `bsrn` package fetches MERRA-2 from Hugging Face **into RAM** (no disk cache). You will see `Fetching MERRA-2 from Hugging Face: {filename}` when it runs. ```python from bsrn.modeling.clear_sky import add_clearsky_columns # Option 1: Load raw parquet (from path, e.g. after hf_hub_download) df = load_merra2_parquet(path) # Option 2: REST2-ready inputs (fetches from HF into RAM, interpolated to 1-min, units converted) # target_index = your BSRN 1-min DatetimeIndex rest2_inputs = fetch_rest2(target_index, station_code="QIQ") # Option 3: Add clear-sky columns to BSRN data (fetches MERRA-2 from HF into RAM automatically) df = add_clearsky_columns(df, station_code="QIQ", model="rest2") ``` ## Data Sources - **MERRA-2**: [NASA GMAO](https://gmao.gsfc.nasa.gov/gmao-products/merra-2/), [GES DISC](https://disc.gsfc.nasa.gov/) - **Extraction**: Google Earth Engine (GEE) — this dataset uses GEE; NCSS is an alternative source. - **Station inventory**: BSRN FTP (months with `.dat.gz` files only) > **GEE extraction and validation** > > The data in this dataset is **extracted via Google Earth Engine (GEE)**. GEE's MERRA-2 pixel boundaries are offset by half a cell in latitude relative to raw NetCDF. To correct this, a **−0.25° latitude shift** is applied when querying GEE so that the returned pixel aligns with the MERRA-2 grid cell used by raw MERRA-2 and NASA GESDISC NCSS. > > **The dataset creator has confirmed the correctness** of the GEE extraction by comparing it against raw MERRA-2 data from NASA and point extractions from NASA GESDISC THREDDS NCSS. ## References 1. Gelaro, R., McCarty, W., Suárez, M. J., Todling, R., Molod, A., Takacs, L., ... & Zhao, B. (2017). The modern-era retrospective analysis for research and applications, version 2 (MERRA-2). *Journal of Climate*, 30(14), 5419–5454. 2. Gueymard, C. A. (2008). REST2: High-performance solar radiation model for cloudless-sky irradiance, illuminance, and photosynthetically active radiation—Validation with a benchmark dataset. *Solar Energy*, 82(3), 272–285. 3. Sun, X., Bright, J. M., Gueymard, C. A., Acord, B., Wang, P., & Engerer, N. A. (2019). Worldwide performance assessment of 75 global clear-sky irradiance models using principal component analysis. *Renewable and Sustainable Energy Reviews*, 111, 550–570. ## Citation If you use this dataset, please cite the references above and the **bsrn** package: [dazhiyang/bsrn](https://github.com/dazhiyang/bsrn). ## License This dataset mirrors publicly available MERRA-2 reanalysis data. MERRA-2 is produced by NASA and is freely available. See [NASA's data use policy](https://www.earthdata.nasa.gov/what-is-nasa-earth-data) for terms of use.
### 数据集元数据 - 语言:英语 - 许可证:MIT协议 - 配置信息:本数据集包含多组配置,每组配置以BSRN站点的三字母小写代码作为配置名称(`e13`除外),每个配置对应一个同名数据分片,数据文件路径格式为`{配置名称}/*.parquet`。 - 展示名称:BSRN MERRA-2大气输入数据集 --- # BSRN MERRA-2大气输入数据集 针对**基线地表辐射网络(Baseline Surface Radiation Network, BSRN)**站点提取的MERRA-2再分析单点数据。本数据集提供的Parquet文件可为**REST2**晴空辐射模型[2]提供大气与气溶胶输入参数。 ## 数据集说明 每个文件包含单个BSRN站点位置处的逐小时MERRA-2变量[1]。数据提取工作通过**谷歌地球引擎(Google Earth Engine, GEE)**从NASA的0.5°×0.625°全球网格完成,数据将对齐至最接近站点坐标的MERRA-2网格单元。 ### 文件结构 站点文件夹使用小写的BSRN三字母代码命名(例如`ber`、`qiq`或`spo`),`e13`为例外。文件命名格式如下: {station}/ {station}{MM}{YY}_merra2.parquet # 每月一个文件 示例: - `qiq/qiq0124_merra2.parquet` —— QIQ站点2024年1月数据 - `ber/ber0325_merra2.parquet` —— BER站点2025年3月数据 ### 变量说明 | 列名 | 说明 | MERRA-2 数据源 | 原始单位 | |--------|-------------------------------------------|----------------|----------------| | AOD55 | 550nm处气溶胶光学厚度 | TOTEXTTAU | 无量纲 | | ALPHA | 安斯特朗指数(Ångström exponent) | TOTANGSTR | 无量纲 | | ALBEDO | 地表反照率 | ALBEDO | [0–1] | | TQV | 整柱可降水量 | TQV | kg/m² | | TO3 | 整柱臭氧总量 | TO3 | 多布森单位(Dobson) | | PS | 地表气压 | PS | 帕斯卡(Pa) | - **索引**:UTC时间索引(逐小时,与MERRA-2原生分辨率一致)。 - **时间覆盖范围**:MERRA-2数据集覆盖1980年至今,本数据集仅生成包含BSRN站点存档FTP数据的对应月份文件。 ### 适配REST2模型 本数据集的Parquet文件专为REST2晴空辐射模型[2]设计。`bsrn` Python包可从本数据集将MERRA-2数据加载至内存(无磁盘缓存),并提供以下功能: 1. `fetch_rest2(index, station_code)`:从Hugging Face加载数据至内存,将时间序列重采样至1分钟目标步长、进行插值、推导BETA参数,并为REST2模型转换单位格式; 2. 原始Parquet文件读取:可通过`huggingface_hub`获取文件路径,使用`pandas.read_parquet`直接读取。 REST2模型所需的输入参数包括:**PS**(单位:百帕)、**ALBEDO**、**ALPHA**、**BETA**(由AOD55与ALPHA推导得到)、**TO3**(单位:大气厘米)、**TQV**(单位:大气厘米)。 #### 单位转换指南(原始数据 → REST2格式) | 变量名 | 原始单位 | REST2要求单位 | 转换方法 | |--------|---------------|---------------|-------------------------------------------| | PS | Pa | hPa | 除以100 | | ALBEDO | [0–1] | [0–1] | 无需转换 | | ALPHA | 无量纲 | 无量纲 | 无需转换 | | BETA | — | — | AOD55 × 0.55^ALPHA(若AOD55=0则使用0.001) | | TO3 | 多布森单位 | atm·cm | 除以1000 | | TQV | kg/m² | atm·cm | 除以10 | ## 使用方法 ### 从Hugging Face加载数据 python from huggingface_hub import hf_hub_download import pandas as pd # 下载单个文件;文件路径格式为 {station}/{station}{MM}{YY}_merra2.parquet path = hf_hub_download( repo_id="dazhiyang/bsrn-merra2", filename="qiq/qiq0124_merra2.parquet", repo_type="dataset", ) df = pd.read_parquet(path) # df包含UTC时间索引与列:AOD55、ALPHA、ALBEDO、TQV、TO3、PS ### 使用bsrn包 `bsrn`包可从Hugging Face将MERRA-2数据加载至内存(无磁盘缓存),运行时会显示`Fetching MERRA-2 from Hugging Face: {filename}`提示信息。 python from bsrn.modeling.clear_sky import add_clearsky_columns # 选项1:加载原始Parquet文件(需先通过hf_hub_download获取本地路径) df = load_merra2_parquet(path) # 选项2:获取REST2格式的输入数据(从Hugging Face加载至内存,重采样至1分钟并转换单位) # target_index = 你的BSRN站点1分钟时间索引 rest2_inputs = fetch_rest2(target_index, station_code="QIQ") # 选项3:向BSRN原始数据中添加晴空辐射列(自动从Hugging Face获取MERRA-2数据) df = add_clearsky_columns(df, station_code="QIQ", model="rest2") ## 数据来源 1. **MERRA-2**:[NASA GMAO](https://gmao.gsfc.nasa.gov/gmao-products/merra-2/)、[GES DISC](https://disc.gsfc.nasa.gov/) 2. **数据提取**:谷歌地球引擎(GEE),本数据集使用GEE完成提取,NCSS为替代提取数据源。 3. **站点清单**:BSRN FTP服务器(仅包含`.dat.gz`文件的月份) > #### GEE提取与验证说明 > 本数据集的数据通过GEE提取得到。GEE的MERRA-2像素边界在纬度方向较原始NetCDF数据偏移半个单元格,为修正此问题,查询GEE时应用了−0.25°的纬度偏移,使返回的像素与原始MERRA-2及NASA GESDISC NCSS使用的MERRA-2网格单元对齐。 > > 数据集创建者已通过将本数据集结果与NASA原始MERRA-2数据、NASA GESDISC THREDDS NCSS的单点提取结果对比,验证了GEE提取流程的正确性。 ## 参考文献 1. Gelaro, R., McCarty, W., Suárez, M. J., Todling, R., Molod, A., Takacs, L., ... & Zhao, B. (2017). The modern-era retrospective analysis for research and applications, version 2 (MERRA-2). *Journal of Climate*, 30(14), 5419–5454. 2. Gueymard, C. A. (2008). REST2: High-performance solar radiation model for cloudless-sky irradiance, illuminance, and photosynthetically active radiation—Validation with a benchmark dataset. *Solar Energy*, 82(3), 272–285. 3. Sun, X., Bright, J. M., Gueymard, C. A., Acord, B., Wang, P., & Engerer, N. A. (2019). Worldwide performance assessment of 75 global clear-sky irradiance models using principal component analysis. *Renewable and Sustainable Energy Reviews*, 111, 550–570. ## 引用说明 若使用本数据集,请引用上述参考文献以及**bsrn**包:[dazhiyang/bsrn](https://github.com/dazhiyang/bsrn)。 ## 许可证 本数据集镜像公开可用的MERRA-2再分析数据。MERRA-2由NASA制作并免费公开,使用条款请参见[NASA数据使用政策](https://www.earthdata.nasa.gov/what-is-nasa-earth-data)。



