遇见数据集

VTSB: A benchmark-ready Vietnamese toxic speech dataset with standardized labels, leakage auditing, and independent out-of-domain evaluation corpus

收藏
Zenodo2026-08-13 更新2026-08-20 收录
官方服务:

资源简介:

Version 1.0: Initial public release, including the raw annotation sheet for the independent evaluation corpus. This is the recommended version for use. The dataset is organized into a structured directory format to facilitate access and reuse. It consists of text-free record-level manifests, a de-identified independent evaluation corpus together with the raw annotation sheet it is built from, aggregate statistical and audit reports, documentation and integrity metadata, and the Python code that reconstructs and validates the resource. The primary purpose of this dataset is to provide a benchmark-ready Vietnamese binary toxicity resource with standardized labels, explicit split and source provenance, and a documented normalized-text leakage audit. By fixing the label mapping, preserving the official partitions of the source releases, and publishing record-level audit keys, the dataset supports consistent and reproducible evaluation across different algorithmic approaches. Comment text from ViHSD and UIT-ViCTSD is not redistributed. For the primary corpus the release provides record-level metadata and SHA-256 audit keys only; users obtain the two source datasets from their authoritative repositories and run the included reconstruction code to materialize the standardized corpus on their own machine. The independent evaluation corpus was collected by the authors and is therefore released in full as text. Beyond benchmarking, the dataset can also be used in related research contexts. The manifests alone support recomputing class distributions, duplicate statistics and cross-split overlap without access to any comment text, which makes them suitable forwork on class imbalance, data contamination and dataset auditing. The independent corpus provides a standardized out-of domain resource for cross-corpus generalization studies, and its paired raw sheet and de-identified output allow the labelling and de-identification procedures themselves to be inspected and re-executed. The primary corpus comprises 43,400 records, of which 36,523 are non-toxic and 6,877 are toxic, giving a toxic rate of 15.85% and an imbalance ratio of 5.31:1. The official partitions are preserved without re-splitting, at 31,048 training, 4,672 development and 7,680 test records. The independent evaluation corpus adds 1,031 human-verified comments, of which 921 are non-toxic and 110 are toxic, giving a toxic rate of 10.67% and an imbalance ratio of 8.37:1, with zero normalized-string overlap with the primary corpus. The manifests directory contains text-free record-level metadata, with one CSV file per official split. Each row carries a deterministic record identifier, the standardized binary label and its human-readable form, the split, the originating dataset, the original source label, the SHA-256 of the normalized text, the size of the duplicate group the record belongs to within its split, and a flag indicating whether its key also occurs in training. No comment text and no platform-specific identifier appears in any column. A fourth file, leakage_free_test_ids.csv, lists the 6,867 record identifiers that form the deterministic leakage-free test view, obtained by removing from the 7,680-record test partition the 813 records whose normalized key also occurs in training. The published SHA-256 values are exact-match audit keys rather than an anonymization mechanism. The independent_corpus directory contains the out-of-domain evaluation corpus. The file vtsb_independent_corpus.csv holds 1,031 de-identified Vietnamese comments with record identifiers and binary toxicity labels. The file data_vihate_indipendent.xlsx is the raw annotation sheet from which that corpus is built, released so that the labelling rule and the de-identification procedure can be re-executed end to end; it is the input to de-identification rather than its output. A README documents the corpus statistics, the annotation protocol, the de-identification patterns and the limits of the sampling frame. The reports directory contains aggregate results in machine-readable and tabular form. The file statistics.json records label, split, source and whitespace-token-length distributions. The file leakage_audit.json records internal duplicate counts, directional cross-split overlap, cross-source overlap and machine-readable definitions of every audit term. Three Markdown files reproduce Tables 1 to 3 of the data article directly from the data. The metadata directory contains documentation and integrity files. The file datasheet.md follows the datasheet format of Gebru et al. (2021). The file schema.json is a per-file column dictionary carrying the redistribution notice. The file provenance.json records the source files, the label mapping, the random seed, the dropped source columns and the per-pattern de-identification counts. The file checksums.json holds the SHA-256 digest of the other 30 files, allowing a download to be verified. The src/vtsb directory contains the nine Python modules that produced everything above: text normalization and audit-key generation, local reconstruction of the primary corpus, construction of the independent corpus, the leakage audit, the statistics and table generators, the release exporter and its pre-upload gate, and a validator that asserts a local reconstruction against every published figure. The tests directory contains 29 unit tests; those requiring the original source releases skip automatically when the files are absent. The notebook run_benchmark.ipynb documents the modelling setup behind the reference results reported in the data article, covering four transformer configurations over three random seeds. These results are provided solely as reference operating points that establish the feasibility and reproducibility of evaluation on this dataset. They are not intended as a comparative study of models or of training strategies, and no configuration is presented as preferable to another. Overall, the dataset comprises 31 files totalling approximately 5.6 MB: 4 manifest files (CSV), 3 independent-corpus files (CSV, XLSX, Markdown), 5 aggregate report files (JSON and Markdown), 4 metadata files (Markdown and JSON), 9 Python modules, 1 test module, 1 Jupyter notebook, and 4 top-level files comprising the README, the licence, the citation file and the package configuration. All released files can be used directly in computational experiments without additional preprocessing. Attaching comment text to the primary-corpus manifests additionally requires the local reconstruction step described above, for which the necessary code and instructions are included. This dataset is associated with the data article: "VTSB: A benchmark-ready Vietnamese toxic speech dataset with standardized labels, leakage auditing, and independent out-of-domain evaluation corpus" by Pham Duc Cua and Nguyen Kieu Linh.

提供机构:
Zenodo
创建时间:
2026-08-13
二维码
社区交流群
二维码
科研交流群
商业服务