PPORTAL: Public domain Portuguese-language literature Dataset
收藏资源简介:
Combining human expertise with information from book-consumer digital data may generate what it takes to face the following changes in such a critical market. Along with the publishing industry, researchers rely on book-related data to develop tools and applications, drawing constructive conclusions to make better informed and faster decisions. Such solutions range from best-selling prediction models to natural language processing to classify raw text. Besides require complex Artificial Intelligence (AI) methods, all of them are essentially data-dependent, mainly book-related data-dependent. Data, and more specifically data growth, is essential for developing and performing such AI-powered applications. None of these efforts can be achieved without a preliminary collection of data on literary works, readers, and their reading habits. Therefore, it is critically important to build and make available datasets that fully comprise the essential elements of the book industry ecosystem. Although some efforts have been made for English language books, little has been done regarding other lesser-spoken languages, such as Portuguese. The evaluation of specific data is of fundamental importance for literature analysis, as Portuguese has its own literary peculiarities. Hence, we present <strong><em>PPORTAL</em></strong>, a <strong>P</strong>ublic domain <strong>PORT</strong>uguese-l<strong>A</strong>nguage <strong>L</strong>iterature dataset. <strong><em>PPORTAL</em></strong>'s contributions are summarized as follows: Data integration of numerous public domain works from three digital libraries; Enriched metadata for works, authors and online reviews extracted from Goodreads; Feature engineering on the metadata to create meaningful additional features; and Unrestricted access in two formats (SQL database and compressed <strong>.csv</strong> files
将人类专业知识与图书消费者数字数据信息相结合,可提供应对这一关键市场后续变革所需的核心支撑。与出版行业同理,研究人员依托图书相关数据开发工具与应用,通过推导建设性结论以制定更具前瞻性且更高效的决策。此类解决方案涵盖畅销书预测模型、用于原始文本分类的自然语言处理技术等多个方向。尽管此类方案均需运用复杂的人工智能(Artificial Intelligence, AI)技术,但本质上均高度依赖数据,尤其是图书相关数据。数据,尤其是数据体量的增长,对于开发并运行此类人工智能驱动的应用而言至关重要。若未预先收集文学作品、读者及其阅读习惯相关数据,则无法开展上述所有研究工作。因此,构建并开放包含图书产业生态系统核心要素的数据集,具备极高的战略重要性。尽管目前已针对英语图书开展了相关数据集构建工作,但针对葡萄牙语等小众语言的相关研究仍寥寥无几。由于葡萄牙语拥有独特的文学特性,针对性数据的评估对于文学分析而言具有根本性的重要意义。为此,我们推出<strong><em>PPORTAL</em></strong>,一款公有领域葡萄牙语文学数据集。<strong><em>PPORTAL</em></strong> 的核心贡献总结如下:整合来自三个数字图书馆的海量公有领域作品;丰富从Goodreads平台提取的作品、作者与在线评论的元数据内容;对元数据开展特征工程处理,生成具备实际价值的新增特征;提供两种格式的无限制访问权限:SQL数据库与压缩后的<strong>.csv</strong>文件



