遇见数据集

Tokenized Forms of Jane Austen Novels with Positional Information

收藏
DataONE2024-05-03 更新2024-10-19 收录
官方服务:

资源简介:

This dataset contains tokenized forms of four Jane Austen novels sourced from Project Gutenberg--Emma, Persuasion, Pride and Prejudice, and Sense and Sensibility--that are broken down by chapter (and volume where appropriate). Each file also includes positional data for each row which will be used for further analysis. This was created to hold the data for the final project for COSC426: Introduction to Data Mining, a class at the University of Tennessee.

本数据集收录源自古腾堡计划(Project Gutenberg)的四部简·奥斯汀小说的Token化文本,分别为《爱玛》(Emma)、《劝导》(Persuasion)、《傲慢与偏见》(Pride and Prejudice)与《理智与情感》(Sense and Sensibility),已按章节(必要时按卷)拆分。每个文件还包含各行的位置数据,可用于后续研究分析。本数据集专为田纳西大学(University of Tennessee)开设的COSC426《数据挖掘导论(Introduction to Data Mining)》课程的期末项目而构建。

创建时间:
2024-09-24
二维码
社区交流群
二维码
科研交流群
商业服务