typhoon-ai/ThaiSafetyBench
收藏资源简介:
--- license: apache-2.0 viewer: false task_categories: - question-answering language: - th size_categories: - 1K<n<10K dataset_info: features: - name: prompt dtype: string - name: risk_area dtype: string - name: types_of_harm dtype: string - name: subtypes_of_harm dtype: string - name: thai_related dtype: bool - name: ai_gen dtype: bool - name: source dtype: string - name: id dtype: int64 splits: - name: test num_bytes: 549743 num_examples: 1889 download_size: 137554 dataset_size: 565298 configs: - config_name: default data_files: - split: test path: data/test-* --- # ThaiSafetyBench <span style="color: red;">⚠️ **Warning:** This dataset contains harmful and toxic language. It is intended for academic purposes only.</span> [[ArXiv Paper](https://arxiv.org/abs/2603.04992v2)] [[Github](https://github.com/trapoom555/ThaiSafetyBench)] [[Hugging Face Leaderboard 🤗](https://huggingface.co/spaces/typhoon-ai/ThaiSafetyBench-Leaderboard)] The ThaiSafetyBench dataset comprises 1,889 malicious Thai-language prompts across various categories. In addition to translated malicious prompts, it includes prompts tailored to Thai culture, offering deeper insights into culturally specific attacks. > **Note:** The *Monarchy* type of harm has been removed from the dataset to comply with Thai regulations. This filtering reduced the dataset from 1,954 to 1,889 samples. ## Dataset Distribution The dataset is categorized into hierarchical categories, as illustrated in the figure below.  ## Dataset Fields | **Field** | **Description** | |--------------------|---------------------------------------------------------------------------------| | `id` | A unique identifier for each sample in the dataset. | | `prompt` | The malicious prompt written in Thai. | | `risk_area` | The primary hierarchical category of risk associated with the prompt. | | `types_of_harm` | A subcategory of the `risk_area`, specifying the type of harm. | | `subtypes_of_harm` | A further subcategory under `types_of_harm`, providing more granular detail. | | `thai_related` | A boolean (`true`/`false`) indicating whether the prompt is related to Thai culture. | | `ai_gen` | A boolean (`true`/`false`) indicating whether the prompt was generated by AI. | | `source` | The source from which the data was collected. | **Note:** The hierarchical structure (`risk_area` → `types_of_harm` → `subtypes_of_harm`) allows for detailed classification of the potential harm caused by each prompt. ## More Dataset Statistics  ## Citation ``` @misc{ukarapol2026thaisafetybenchassessinglanguagemodel, title={ThaiSafetyBench: Assessing Language Model Safety in Thai Cultural Contexts}, author={Trapoom Ukarapol and Nut Chukamphaeng and Kunat Pipatanakul and Pakhapoom Sarapat}, year={2026}, eprint={2603.04992}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2603.04992}, } ```



