CUBANSPVARIETY
收藏资源简介:
CUBANSPVARIETY数据集是首个专注于古巴或加勒比西班牙语变体识别的数据集,由1762条手动标注的推文组成,标注由三名古巴母语者完成。数据集内容涵盖古巴西班牙语变体、非古巴变体以及常见示例,旨在解决西班牙语变体识别中的常见示例分类问题。数据集的创建过程包括从Twitter上收集数据并进行手动标注,标注过程中考虑了推文的语言变体信息。该数据集主要应用于自然语言处理中的语言变体识别任务,特别是用于提高模型在处理常见示例时的准确性和鲁棒性。
The CUBANSPVARIETY Dataset is the first dataset dedicated to the identification of Cuban or Caribbean Spanish varieties. It comprises 1,762 manually annotated tweets, with annotation completed by three native Cuban speakers. The dataset covers Cuban Spanish varieties, non-Cuban Spanish varieties, and common samples, aiming to address the common sample classification issues in Spanish variety identification. The dataset creation process includes collecting data from Twitter and performing manual annotation, where the annotation procedure takes into account the language variety information of the tweets. This dataset is primarily applied to language variety identification tasks in Natural Language Processing (NLP), and is specifically used to enhance the accuracy and robustness of models when handling common samples.




