aisingapore/Small-SEA-Instruct-2602
收藏资源简介:
这是一个用于论文“DuDi: Dual-Signal Distillation with Cross-Lingual Verbalizer”的训练数据集。它基于SEA-Instruct数据集,覆盖了七种东南亚语言:印尼语、越南语、泰语、泰米尔语、他加禄语、马来语和缅甸语。数据集包含开源提示,每个提示都配有一个合成响应和质量估计。从原始数据集中,每种语言采样了4000个高质量示例(根据原始标注),总计28000个样本。通过随机采样约束,保持了领域、任务类型和提示复杂度的分布。
This is a training dataset for paper "DuDi: Dual-Signal Distillation with Cross-Lingual Verbalizer". We use SEA-Instruct, which covers seven SEA languages: Indonesian, Vietnamese, Thai, Tamil, Tagalog, Malay, and Burmese. The dataset contains open-source prompts, each paired with a synthetic response and quality estimate. We sample 4,000 high-quality examples per language, as labeled by the original dataset, resulting 28,000 samples. Random sampling constraints preserve the distribution of domains, task types, and prompt complexity.




