遇见数据集

CzechTopic: A Benchmark for Zero-Shot Topic Localization in Historical Czech Documents

收藏
Zenodo2026-03-09 更新2026-05-26 收录
官方服务:

资源简介:

CzechTopic is a benchmark dataset of historical Czech documents designed for topic localization and document classification in a zero-shot setting. Each document contains 768–1024 characters and is written in Czech. The dataset consists of two parts: a development set and a test set. The development set contains 15,245 documents and 19,107 topics. Each topic is annotated in 10 documents. The annotations for the development set were generated using the GPT-5-2 model. The test set contains 525 documents and 364 human-created topics, with each topic annotated in five documents. All annotations are provided as character spans, indicating the exact locations in the text where a topic appears. We evaluate models at two levels: text level and word level. At the text level, the task is to determine whether a given topic is present in a document. At the word level, the task is to identify which words correspond to a given topic.

提供机构:
Zenodo
创建时间:
2026-03-09
二维码
社区交流群
二维码
科研交流群
商业服务