遇见数据集

Small-SEA-Instruct-2602

收藏
魔搭社区2026-06-14 更新2026-08-16 收录
官方服务:

资源简介:

This is a training dataset for paper "[DuDi: Dual-Signal Distillation with Cross-Lingual Verbalizer ](https://arxiv.org/pdf/2606.04694)". We use [SEA-Instruct](https://huggingface.co/datasets/aisingapore/SEA-Instruct-2602), which covers seven SEA languages: Indonesian, Vietnamese, Thai, Tamil, Tagalog, Malay, and Burmese. The dataset contains open-source prompts, each paired with a synthetic response and quality estimate. We sample 4,000 high-quality examples per language, as labeled by the original dataset, resulting 28,000 samples. Random sampling constraints preserve the distribution of domains, task types, and prompt complexity.

提供机构:
maas
创建时间:
2026-06-01
二维码
社区交流群
二维码
科研交流群
商业服务