遇见数据集

tiny-think-dpo-code

收藏
魔搭社区2026-07-18 更新2026-08-02 收录
官方服务:

资源简介:

# Shekswess/tiny-think-dpo-code ## Overview Direct Preference Optimization (DPO) dataset built from `allenai/Dolci-Think-DPO-7B` using the `facebook/MobileLLM-R1-140M-base` tokenizer and chat template. This dataset targets code-related preferences. ## Dataset Details - **Build date:** 2026-01-10 - **Sources:** 3 - **Rows:** 2,708 - **Tokens:** 9,999,838 - **Max sequence length:** 4096 tokens per example (both `chosen` and `rejected`) - **Token budget:** 10,000,000 tokens (equal strategy) - **Columns:** `prompt`,`chosen`, `rejected`, `dataset_source`, `token_count` ## Data Schema - `prompt`: user input prompt - `chosen`: preferred chat message list (role/content) - `rejected`: non-preferred chat message list (role/content) - `dataset_source`: string identifier for the originating source dataset. - `token_count`: total tokens for the pair under the tokenizer. ## Intended Use DPO for small models on single-GPU setups, emphasizing python code preferences.

提供机构:
maas
创建时间:
2026-01-08
二维码
社区交流群
二维码
科研交流群
商业服务