遇见数据集

Preprocessed Penn Tree Bank

收藏
Zenodo2020-06-27 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

Penn Tree Bank is originally a corpus of English sentences with linguistic structure annotations. This is a variant originally distributed by Mikolov here and Zaremba here which omits the annotation and splits the dataset into raining, validation, and test. The dataset consists of <strong>929k</strong> <strong>training</strong> words, <strong>73k</strong> <strong>validation</strong> words, and <strong>82k</strong> test words. As part of the preprocessing, words were <strong>lower-cased</strong>, <strong>numbers</strong> were replaced with <strong>N</strong>, newlines were replaced with <strong>&lt;eos&gt;</strong>, and all other punctuation was removed. The vocabulary is the most frequent <strong>10k</strong> words with the rest of the tokens replaced by an <strong>&lt;unk&gt;</strong> token.

提供机构:
Zenodo
创建时间:
2020-06-26
二维码
社区交流群
二维码
科研交流群
商业服务