遇见数据集

"Basisklassifikation" (BK) Training Dataset for Automatic Subject Indexing

收藏
Zenodo2025-09-01 更新2026-05-26 收录
官方服务:

资源简介:

This is a training dataset for automatic subject indexing containing more than 6 Mio. titles and their corresponding subjects (classes) from the "Basisklassifikation" (BK). Initially introduced in the 1980s, today the Basisklassifikation constitutes one of the most widely used classification system for subject indexing within the Berlin State Library. As of August 2024, around 11% of all of the works (~8.7 Mio. titles) included in the K10plus catalogue – the union catalogue of the German library networks GBV and SWB – have already received BK notations. The dataset consists of files in .tsv format, intended to be used together with the Annif tool for automatic subject indexing and a vocabulary file of the BK (see BK Download). In this way, Annif models for the prediction of BK classes have been trained which can be accessed via the Staatsbibliothek zu Berlin – Preußischer Kulturbesitz community at HuggingFace. The dataset was created by the team of the research project "Mensch.Maschine.Kultur – Künstliche Intelligenz für das Digitale Kulturelle Erbe" at Berlin State Library (SBB) which was funded by the Federal Government Commissioner for Culture and the Media (BKM), project grant no. 2522DIG002. More specifically, it was created in the context of sub-project 3 "AI-supported content analysis and subject indexing".

创建时间:
2025-09-01
二维码
社区交流群
二维码
科研交流群
商业服务