遇见数据集

Dataset for: Using Transformer-based models for Vietnamese language detection

收藏
Zenodo2026-02-04 更新2026-05-26 收录
官方服务:

资源简介:

Abstract: This dataset contains the training and testing data used in the research manuscript "Using Transformer-based models for Vietnamese language detection" (PLOS ONE). The data consists of labeled text sequences distinguishing Vietnamese language content from other languages, utilized to fine-tune Transformer-based models. Data Content: The dataset aggregates data from multiple sources to ensure robustness against orthographic noise and mixed-language contexts. It includes: Self-collected data: News articles crawled from VNExpress (Vietnamese and English editions), adhering to the source's robots.txt and terms of service. Open-source subsets: Selected samples from Binhvq News Corpus, vi-error-correction-2.0, and OPUS Tatoeba. Hugging Face Repository: For ease of access, code integration, and the full version of the dataset including any future updates, please visit the authors' Hugging Face repository: URL: https://huggingface.co/datasets/Cerberose/vietnamese-classification-dataset Usage: This dataset can be loaded directly in Python using the datasets library. Files Included: train.csv: The training set containing labeled text samples. test.csv: The hold-out test set used for evaluating model performance. Citation: If you use this dataset, please cite the associated PLOS ONE article.

提供机构:
Zenodo
创建时间:
2026-02-04
二维码
社区交流群
二维码
科研交流群
商业服务