Small-SEA-Instruct-2602
收藏官方服务:
资源简介:
This is a training dataset for paper "[DuDi: Dual-Signal Distillation with Cross-Lingual Verbalizer ](https://arxiv.org/pdf/2606.04694)". We use [SEA-Instruct](https://huggingface.co/datasets/aisingapore/SEA-Instruct-2602), which covers seven SEA languages: Indonesian, Vietnamese, Thai, Tamil, Tagalog, Malay, and Burmese. The dataset contains open-source prompts, each paired with a synthetic response and quality estimate. We sample 4,000 high-quality examples per language, as labeled by the original dataset, resulting 28,000 samples. Random sampling constraints preserve the distribution of domains, task types, and prompt complexity.
提供机构:
maas创建时间:
2026-06-01



