philosophy-corpus
收藏资源简介:
Philosophy Corpus 是一个包含54部经典哲学文本的数据集,专为训练字符级语言模型而设计。数据集涵盖多个哲学流派和时期,包括柏拉图(12部作品)、亚里士多德(12部作品)、斯多葛学派(4部作品)、罗马时期(4部作品)、早期现代(6部作品)以及启蒙/19世纪(10部作品)等哲学家的著作。数据以文本字符串形式存储,包含训练集(284,485个样本,20.5MB)和验证集(31,610个样本,2.3MB)两个分割。数据集提供三种文件格式:单独清洗的文本文件(54个文件)、合并后的段落文本(全部小写,85K段落,约25MB)以及完整保留原始大小写的合并语料(约26MB)。数据来源为古腾堡计划和MIT互联网经典档案馆的公共领域翻译文本。该数据集已成功应用于JuliaGPT和JuliaFluxGPT等字符级Transformer模型的训练。
Philosophy Corpus is a dataset containing 54 classic philosophical texts, specifically designed for training character-level language models. The dataset covers multiple philosophical schools and periods, including works by Plato (12 works), Aristotle (12 works), Stoicism (4 works), Roman period (4 works), Early Modern period (6 works), Enlightenment/19th century (10 works) and other philosophers' writings. The data is stored as text strings, with two splits: training set (284,485 samples, 20.5 MB) and validation set (31,610 samples, 2.3 MB). Three file formats are provided: separately cleaned text files (54 files), merged paragraph text (all lowercase, 85K paragraphs, ~25 MB), and merged corpus that fully retains the original capitalization (approximately 26 MB). The data sources are public domain translations from Project Gutenberg and the MIT Internet Classics Archive. This dataset has been successfully applied to the training of character-level Transformer models such as JuliaGPT and JuliaFluxGPT.



