Data in Sgr B2(N)
收藏资源简介:
This repository contains the data used for column density prediction in the Sagittarius B2(N) region. This dataset is compiled from molecular column density tables reported in the peer-reviewed literature, specifically from: "The physical and chemical structure of Sagittarius B2. VIII. Full molecular line survey of hot cores." All data included here are derived from previously published results. No new observational data are introduced. The purpose of this repository is to provide a curated,machine-readable version of the published tables to facilitate reuse, reproducibility, and data-driven modeling. The dataset is provided as a single ZIP archive. After decompression, it contains 20 folders labeled A01–A20, each corresponding to one sub-region / sample in Sgr B2(N). Each folder includes 7 data files representing the associated molecular or physical features used in the prediction tasks. The dataset is organized as follows: - A01–A20: Each folder represents one sample or sub-region in Sgr B2(N), containing 7 data files. - training_data: Training data without explicit sample indices. - training_data_total: Training data with explicit sample indices added for unified indexing. - testing_data: Testing data without explicit sample indices. - testing_data_total: Testing data with explicit sample indices added for unified indexing. - training_label: Column density labels corresponding to the training data. - testing_label: Column density labels corresponding to the testing data. - total_data: The complete dataset combining both training and testing data. The indexed (“_total”) versions are provided to facilitate reproducible data splitting, unified indexing across samples, and direct use in machine learning workflows. The non-indexed versions preserve the original data organization.
本仓库包含用于人马座B2(N)(Sagittarius B2(N))区域柱密度(column density)预测的相关数据。 本数据集源自经同行评议的学术文献中报道的分子柱密度表,具体来自论文:《人马座B2的物理与化学结构 Ⅷ. 热核团的全分子线巡天》(The physical and chemical structure of Sagittarius B2. VIII. Full molecular line survey of hot cores.)。 本数据集收录的所有数据均来自已发表的研究成果,未引入任何新的观测数据。本仓库的目的是提供经精选整理且可被机器读取(machine-readable)的已发表表格版本,以促进数据复用、研究可重复性以及数据驱动建模。 本数据集以单个ZIP归档(ZIP archive)形式提供。解压后包含20个命名为A01至A20的文件夹,每个文件夹对应人马座B2(N)中的一个子区域或样本。每个文件夹内包含7个数据文件,代表预测任务中使用的相关分子或物理特征。 本数据集的组织结构如下: - A01–A20:每个文件夹对应人马座B2(N)中的一个样本或子区域,内含7个数据文件。 - training_data:未添加显式样本索引的训练数据。 - training_data_total:添加了显式样本索引以实现统一索引的训练数据。 - testing_data:未添加显式样本索引的测试数据。 - testing_data_total:添加了显式样本索引以实现统一索引的测试数据。 - training_label:与训练数据对应的柱密度标签。 - testing_label:与测试数据对应的柱密度标签。 - total_data:整合训练与测试数据的完整数据集。 带索引(后缀为_total)的版本旨在支持可复现的数据划分、跨样本的统一索引,以及直接应用于机器学习工作流(machine learning workflows)。无索引版本则保留了原始数据的组织形式。



