遇见数据集

4lang: Open Access Dataset for Cross-Lingual Plagiarism Detection

收藏
NIAID Data Ecosystem2026-05-01 收录
官方服务:

资源简介:

A dataset for cross-lingual plagiarism evaluation. 4collection.zip: a subset of Wikipedia articles on 4 languages (ru, hy, es, en). 4query.zip: wikipedia documents in each of the four languages with translated sentences with Google Translate API from collection. The archieve contains text documents and XML-markup for them. For the markup description and evaluation see http://pan.webis.de/clef13/pan13-web/plagiarism-detection.html

创建时间:
2023-04-05
二维码
社区交流群
二维码
科研交流群
商业服务