KRadim/czech-punctuation-pos-syntax
收藏资源简介:
这是一个捷克语的结构化、语言学标注语料库,专门设计用于标点恢复任务、词性标注和词级句法嵌入(如nanoGPT自定义元数据训练)。与纯原始文本语料库不同,该数据集通过斯坦福的先进Stanza管道(基于布拉格依存树库标准训练)直接提供形态特征(词类、语法格)和依存句法角色(主语、谓语、宾语)的确定性1:1词级映射。数据集包含三个官方子集,采用严格可复现的80%/10%/10%分布进行分割,并经过种子随机打乱。数据以优化的Apache Parquet格式提供,包含原始句子、标点类型和词法注释等列。
This dataset is a structured, linguistically annotated corpus of the Czech language, specifically designed for Punctuation Restoration tasks, Part-of-Speech (POS) tagging, and token-level syntax embedding (such as nanoGPT custom metadata training). Unlike pure raw text corpora, this dataset provides a deterministic 1:1 token-level mapping of morphological features (word class, grammatical case) and dependency syntax roles (subject, predicate, object) directly derived via Stanfords state-of-the-art Stanza pipeline (trained on the Prague Dependency Treebank standard).



