Princeton Prosody Archive Found Poems Dataset
收藏资源简介:
This is a dataset of poetry excerpts found in the pages of the Princeton Prosody Archive (PPA). It consists of 1,263,587 excerpts identified when searching for quotes from 246,839 reference poems within the 1,982,317 pages of PPA texts. The dataset includes details for the excerpts, along with poem and PPA work-level metadata for analyzing the excerpts in context. The poem and PPA metadata files include summary information about the number of excerpts found. Refer to the dataset readme or datapackage for a complete list of files and fields. Poetry excerpts were identified by three different methods, but the overwhelming majority (99%) were found by running passim against reference poetry corpora. This is a preliminary version of the dataset; it is known to include temporally implausible excerpts (misidentified quotations before the poem was first published) as well as duplicate and overlapping excerpts. We plan to address these in future versions.



