Mini-LLM数据集是用于Mini-LLM项目的训练数据,主要包括OpenCSG Fineweb-Edu-Chinese-V2.1数据集(约70GB)和Jiangshu Tech Large Model Dataset(约16GB),用于分词器训练、预训练和SFT数据。数据集提供下载脚本,支持模型训练和优化。
This dataset was created using [LeRobot](https://github.com/huggingface/lerobot). <a class="flex" href="https://huggingface.co/spaces/lerobot/visualize_dataset?path=lerobot/utokyo_xarm_pick_and_pla
Dataset of [MegaStyle](https://jeoyal.github.io/MegaStyle/). MegaStyle-1.4M is a large-scale style dataset built through a scalable pipeline that leverages consistent text-to-image style mapping of Q
# Dataset Information A Chain of Thought (CoT) version of the TAT-QA arithmetic dataset (hosted at https://huggingface.co/datasets/nvidia/ChatQA-Training-Data). The dataset was synthetically generat