Preprocessed Penn Tree Bank
收藏资源简介:
Penn Tree Bank is originally a corpus of English sentences with linguistic structure annotations. This is a variant originally distributed by Mikolov here and Zaremba here which omits the annotation and splits the dataset into raining, validation, and test. The dataset consists of <strong>929k</strong> <strong>training</strong> words, <strong>73k</strong> <strong>validation</strong> words, and <strong>82k</strong> test words. As part of the preprocessing, words were <strong>lower-cased</strong>, <strong>numbers</strong> were replaced with <strong>N</strong>, newlines were replaced with <strong><eos></strong>, and all other punctuation was removed. The vocabulary is the most frequent <strong>10k</strong> words with the rest of the tokens replaced by an <strong><unk></strong> token.



