遇见数据集

Automatic translation and multilingual cultural heritage retrieval: a case study with transcriptions in Europeana (dataset)

收藏
Zenodo2021-09-10 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

The dataset contains all the data required to reproduce the experiments done in the paper "Automatic translation and multilingual cultural heritage retrieval: a case study with transcriptions in Europeana", published in the 25th International Conference on Theory and Practice of Digital Libraries (TPDL'21). In that work we run an experiment using the Europeana CH digital library as a use case, and we evaluated the effectiveness of a multilingual information retrieval strategy using machine translations to English as pivot language. We used the CEF translation service (eTranslation) for the translation of queries and content to English (https://ec.europa.eu/cefdigital/wiki/display/CEFDIGITAL/eTranslation). The dataset is also available at https://rnd-2.eanadev.org/share/crosslingual-search/, and it is organized in four main folders: <strong>queries</strong>: sample of 68 queries and their translations to English. The queries were issued in languages other than English from the Europeana Portal, using the Europeana’s 1914-1918 thematic collection, between January and August 2019. <strong>transcriptions</strong>: sample of 18,257 handwriting transcriptions and its translations to English. The transcriptions are taken from the Europeana 1914-1918 thematic collection, and obtained from the Transcribathon crowdsourcing platform (https://europeana.transcribathon.eu/). <strong>solr_configuration</strong>: Apache Solr search engine configuration used in the experiments (which replicates the one used in Europeana). <strong>results</strong>: manual evaluation of the query translations, and automatic evaluation of the multilingual retrieval.

本数据集包含复现发表于第25届国际数字图书馆理论与实践会议(TPDL'21)的论文《自动翻译与多语言文化遗产检索:以Europeana中的转录文本为例》(Automatic translation and multilingual cultural heritage retrieval: a case study with transcriptions in Europeana)中所开展实验的全部必要数据。 该研究以Europeana文化遗产数字图书馆作为应用案例开展实验,并评估了以机器翻译至英语作为枢纽语言的多语言信息检索策略的有效性。 本次实验使用CEF翻译服务(eTranslation)完成查询语句与文本内容到英语的翻译,相关服务详情可访问:https://ec.europa.eu/cefdigital/wiki/display/CEFDIGITAL/eTranslation。 本数据集同时可于https://rnd-2.eanadev.org/share/crosslingual-search/获取,其分为四个核心文件夹: **queries(查询集)**:包含68条查询语句及其英语译稿。这些查询语句于2019年1月至8月间,通过Europeana门户针对Europeana 1914-1918专题馆藏发起,且均非英语撰写。 **transcriptions(转录文本集)**:包含18257条手写转录文本及其英语译稿。该转录文本取自Europeana 1914-1918专题馆藏,来源于Transcribathon众包平台(https://europeana.transcribathon.eu/)。 **solr_configuration(Solr配置集)**:实验中使用的Apache Solr搜索引擎配置文件(该配置复刻了Europeana所使用的配置)。 **results(结果集)**:包含查询语句翻译的人工评估结果,以及多语言检索任务的自动评估结果。

提供机构:
Zenodo
创建时间:
2021-06-30
二维码
社区交流群
二维码
科研交流群
商业服务