遇见数据集

tiny-think-dpo

收藏
魔搭社区2026-04-30 更新2026-08-09 收录
官方服务:

资源简介:

# Shekswess/tiny-think-dpo ## Overview Merged Direct Preference Optimization (DPO) dataset combining chat + instruct, math + STEM, and code subsets. Built from `allenai/Dolci-Think-DPO-7B` using the `facebook/MobileLLM-R1-140M-base` tokenizer and chat template. ## Dataset Details - **Build date:** 2026-01-10 - **Rows:** 10,483 - **Tokens:** 29,997,608 - **Max sequence length:** 4096 tokens per example (both `chosen` and `rejected`) - **Columns:** `prompt`,`chosen`, `rejected`, `dataset_source`, `token_count` ## Data Schema - `prompt`: user input prompt - `chosen`: preferred chat message list (role/content) - `rejected`: non-preferred chat message list (role/content) - `dataset_source`: string identifier for the originating source dataset. - `token_count`: total tokens for the pair under the tokenizer. ## Intended Use DPO for small models on single-GPU setups, providing a broad general-purpose preference mix.

提供机构:
maas
创建时间:
2026-01-08
二维码
社区交流群
二维码
科研交流群
商业服务