HuangBoYuan/SkyGraph-1.5B
收藏资源简介:
SkyGraph-1.5B是一个大规模中文文本语料库,以HDF5格式打包,包含图结构语言标注。每个样本对应一个句子,包括分词器输入ID、词级token ID、依存句法分析目标、AMR风格语义图目标以及外部token嵌入分片的引用。该数据集包含36,721,440个句子,分为74个分片,磁盘大小约57 GB,提供依存图和AMR图标注,并附带词汇表文件。数据集适用于图感知语言建模、语法语义表示学习、图解码和中文文本建模研究。
SkyGraph-1.5B is a large-scale Chinese text corpus packaged in HDF5 format with graph-structured linguistic annotations. Each example corresponds to one sentence and includes tokenizer input ids, word-level token ids, dependency parsing targets, AMR-style semantic graph targets, and references to external token embedding shards. The dataset contains 36,721,440 sentences, divided into 74 shards, with a disk size of approximately 57 GB, providing dependency graph and AMR graph annotations, along with vocabulary files. It is intended for research on graph-aware language modeling, syntax-semantics representation learning, graph decoding, and Chinese text modeling.




