遇见数据集

PPORTAL: Public domain Portuguese-language literature Dataset

收藏
Zenodo2024-07-03 更新2026-05-26 收录
官方服务:

资源简介:

Combining human expertise with information from book-consumer digital data may generate what it takes to face the following changes in such a critical market. Along with the publishing industry, researchers rely on book-related data to develop tools and applications, drawing constructive conclusions to make better informed and faster decisions. Such solutions range from best-selling prediction models to natural language processing to classify raw text. Besides require complex Artificial Intelligence (AI) methods, all of them are essentially data-dependent, mainly book-related data-dependent. Data, and more specifically data growth, is essential for developing and performing such AI-powered applications. None of these efforts can be achieved without a preliminary collection of data on literary works, readers, and their reading habits. Therefore, it is critically important to build and make available datasets that fully comprise the essential elements of the book industry ecosystem. Although some efforts have been made for English language books, little has been done regarding other lesser-spoken languages, such as Portuguese. The evaluation of specific data is of fundamental importance for literature analysis, as Portuguese has its own literary peculiarities. Hence, we present PPORTAL, a Public domain PORTuguese-lAnguage Literature dataset. PPORTAL's contributions are summarized as follows: Data integration of numerous public domain works from three digital libraries; Enriched metadata for works, authors and online reviews extracted from Goodreads; Feature engineering on the metadata to create meaningful additional features; and Unrestricted access in two formats (SQL database and compressed .csv files

将人类专业知识与图书消费者数字数据信息相结合,可为应对这一关键市场中的后续变革提供必要支撑。与出版行业同理,研究人员依托图书相关数据开发工具与应用,通过得出建设性结论以实现更科学高效的决策。此类解决方案涵盖畅销预测模型、用于原始文本分类的自然语言处理技术等范畴。除需依托复杂人工智能(Artificial Intelligence, AI)方法外,所有此类方案本质上均依赖数据,尤以图书相关数据为核心依托。 数据,尤其是数据规模的增长,对于开发并运行此类人工智能赋能的应用至关重要。若未预先收集文学作品、读者及其阅读习惯相关数据,则无法推进上述任一工作。因此,构建并开放涵盖图书产业生态系统核心要素的数据集,具有至关重要的意义。尽管针对英语图书已开展了相关建设工作,但针对葡萄牙语等小众语言的相关探索仍十分匮乏。由于葡萄牙语拥有独特的文学特质,针对性数据的评估对于文学分析具有基础性意义。据此,我们推出了PPORTAL:公共领域葡萄牙语文学数据集(Public domain PORTuguese-lAnguage Literature dataset)。PPORTAL的核心贡献总结如下: - 整合来自三家数字图书馆的海量公共领域作品; - 丰富从Goodreads平台提取的作品、作者及在线评论的元数据; - 针对元数据开展特征工程,构建具备实际应用价值的新增特征; - 提供两种格式的无限制获取途径:SQL数据库与压缩.csv文件

提供机构:
Zenodo
创建时间:
2024-07-03
二维码
社区交流群
二维码
科研交流群
商业服务