遇见数据集

Catalan-Chinese Wikimedia dataset

收藏
Zenodo2025-06-11 更新2026-05-26 收录
官方服务:

资源简介:

Dataset Summary The Catalan-Chinese Wikimedia dataset includes sentences from three sources: Catalan Wikinews, Catalan Wikipedia, and Spanish Wikivoyage, with an average sentence length of 26 words. Sentences from Catalan Wikinews and Wikipedia were translated into Spanish and Mandarin Chinese using OpenAI’s GPT-4. Likewise, sentences from Spanish Wikivoyage were translated into Catalan and Mandarin Chinese with the same model. All Mandarin Chinese translations were later reviewed and post-edited by a professional human translator. In total, the dataset contains 1,022 parallel sentences in Catalan, Spanish, and Mandarin Chinese. Supported Tasks and Leaderboards The dataset can serve both as a human preference dataset for Reinforcement Learning and as a benchmark for evaluate Bilingual Machine Translation models between Catalan and Chinese in any direction, as well as Multilingual Machine Translation models. Languages The sentences included in the dataset are in Catalan (CA), Spanish (ES), and simplified Chinese (ZH). Dataset Structure Data Instances Each data instance contains parallel sentences in three languages along with metadata. Below is a sample instance: { "id": 1, "article_id": 1, "sentence_id": 1, "URL": "https://ca.wikinews.org/wiki/La_NASA_suspèn_el_llançament_de_la_nau_Orion", "domain": "Wikinews (Catalan)", "topic": "Ciència i tecnologia", "sentence_ca": "La nau havia d'orbitar el planeta Terra dues vegades a una altitud aproximada de 5800 quilòmetres, abans de caure a l'oceà Pacífic, a prop de la costa de Baixa Califòrnia.", "sentence_es": "La nave debía orbitar el planeta Tierra dos veces a una altitud aproximada de 5800 kilómetros, antes de caer en el océano Pacífico, cerca de la costa de Baja California. ", "sentence_zh_gpt": "这艘航天器需要在大约5800公里的高度绕地球轨道飞行两次,然后坠入太平洋,靠近下加利福尼亚州的海岸。", "sentence_zh_human": "这艘飞船需要在大约5800公里的高度绕地球轨道飞行两次,之后坠入靠近下加利福尼亚州海岸的太平洋。" } Data Fields id: Unique row number for each data entry, ranging from 1 to 1022. article_id: Identifier for the source article, ranging from 1 to 296. sentence_id: Identifier for the sentence within the article, ranging from 1 to 5. URL: Link to the original article in Catalan or Spanish on Wikimedia platforms. domain: Source of the sentence, including Wikinews (Catalan, https://ca.wikinews.org/wiki/Portada), Wikipedia (Catalan, https://ca.wikipedia.org/wiki/Portada), or Wikivoyage (Spanish, https://es.wikivoyage.org/wiki/P%C3%A1gina_principal). topic: Topic category of the sentence. For Catalan Wikinews, there are 10 categories: Ciència i tecnologia, Cultura i esplai, Dret, Economia, Esports, Medi ambient, Política, Necrologia, Salut, and Successos. sentence_ca: Sentence in Catalan. For Spanish Wikivoyage entries, this was translated from Spanish by GPT-4. As such, some sentences are original Catalan Wikimedia content, while others are GPT-4 translations. sentence_es: Sentence in Spanish. For entries from Catalan Wikinews and Catalan Wikipedia, this was translated from Catalan by GPT-4. As such, it includes both original Spanish Wikimedia content and GPT-4 translations. sentence_zh_gpt: Simplified Chinese translation generated by GPT-4, translated from the original sentence in either Catalan or Spanish. sentence_zh_human: Human post-edited version of the Simplified Chinese translation. Data Splits The dataset contains a single split: test. Curation Rationale This dataset is aimed at promoting the development of Machine Translation between Catalan and other languages, specifically Mandarin Chinese. Source Data Sentences are extracted from Catalan Wikinews, Catalan Wikipedia, and Spanish Wikivoyage. Data Collection and Translation The dataset includes sentences collected from three primary sources: Catalan Wikinews, Catalan Wikipedia, and Spanish Wikivoyage. Approximately one-third of the sentences come from each source. Source CA Wikinews CA Wikipedia ES Wikivoyage Number of sentences 328 341 353 Article Selection:Approximately 100 articles were randomly selected from each source using the Wikimedia API. Sentence Selection: For each article, 3–5 contiguous sentences were selected. For Catalan Wikinews and Catalan Wikipedia articles, sentences were sampled approximately equally from the beginning, middle, and end of the article to ensure diversity. For Spanish Wikivoyage articles, efforts were made to ensure sentences represent different topics, such as "beber y salir," "clima," "comprar," "flora y fauna", etc. Sentence position Count Middle 231 End 222 Begining 216 llegar 30 comprender 30 ver 29 desplazarse 26 comer 25 comprar 25 hacer 25 mantente seguro 20 dormir 19 beber y salir 17 siguiente destino 15 introduction 14 itinerario 12 salud 10 pasear 6 clima 6 empleos disponibles 5 Organizaciones de albergues 5 Cascadas más famosos 4 consejos 4 Flora y fauna 4 tornado 4 destinos 3 madera 3 respeto 3 Características propias 3 costas y playas 3 equipamiento 3 Who are the source language producers? Catalan Wikinews Catalan Wikipedia Spanish Wikivoyage. Annotations Annotation process Translations were initially performed using OpenAI GPT-4, followed by manual post-editing by a native Mandarin Chinese speaker who is also a professional translator between Spanish and Chinese. Who are the annotators? Xixian Liao (xixian.liao@bsc.es) Personal and Sensitive Information The dataset is derived from publicly available content on Wikipedia, Wikinews, and Wikivoyage, which are community-edited platforms. No specific anonymization process has been applied to the dataset. While these sources generally avoid publishing personal or sensitive information, it is possible that such information may be present. This needs to be considered when using the data for training models. Considerations for Using the Data Social Impact of Dataset By providing this resource, we intend to promote the use of Catalan across NLP tasks, thereby improving the accessibility and visibility of the Catalan language. Discussion of Biases No specific bias mitigation strategies were applied to this dataset. Inherent biases may exist within the data. Other Known Limitations The dataset contains data of a general domain. Applications of this dataset in more specific domains such as biomedical, legal etc. would be of limited use. Additional Information Dataset Curators Language Technologies Unit at the Barcelona Supercomputing Center (langtech@bsc.es). This work has been promoted and financed by the Generalitat de Catalunya through the Aina project. Licensing Information This work is licensed under Creative Commons Attribution-ShareAlike.

提供机构:
Zenodo
创建时间:
2025-06-11
二维码
社区交流群
二维码
科研交流群
商业服务