遇见数据集

CEREAL II, el Corpus del Español REAL

收藏
Zenodo2025-02-07 更新2026-05-26 收录
官方服务:

资源简介:

Content: CEREAL (visit the project website) is a document-level corpus of documents in Spanish extracted from OSCAR. The documents are classified according to their country of origin. CEREALsentence is the sentence-level corpus extracted from CEREAL. It covers 24 countries where Spanish is spoken. Following OSCAR, we provide our annotations with CCO license, but we do not hold the copyright of the content text which comes from OSCAR and therefore from Common Crawl. The process to build the corpus and its characteristics can be found in: Cristina España-Bonet and Alberto Barrón-Cedeño. "Elote, Choclo and Mazorca: on the Varieties of Spanish." In proceedings of the 2024 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL 2024), Ciudad de México, Mexico, June 2024. This repository contains the corpora (train/validation/test) used to train the classifier, the sentence-level version of CEREAL and the word embeddings calculated from it. The complete document-level corpus is available inhttps://zenodo.org/records/11387864 Files Description: See the README.txt file

提供机构:
Zenodo
创建时间:
2024-06-15
二维码
社区交流群
二维码
科研交流群
商业服务