Ling-flash-2.0-open-perfectblend-regenerate
收藏资源简介:
## Dataset Overview **Ling-Flash-2.0-open-perfectblend-regenerate** is a high-quality synthetic instruction dataset derived from the **Open-PerfectBlend** dataset. It serves as a foundational component for the **Ling-Flash-2.0** ecosystem. This dataset consists of **1.4 million samples** that have been regenerated using the **Ling-Flash-2.0** model. By replacing original human annotations with model-generated responses, this dataset ensures higher consistency and better alignment with the model's internal distribution. This leads to significantly improved instruction understanding, response quality, and acceptance rates. --- ## Dataset Source - **Base Dataset:** Open-PerfectBlend (1.4M samples). - **Generator Model:** Ling-Flash-2.0. - **Method:** Automated inference and regeneration to refine response consistency. --- ## Dataset Usage The primary purpose of this dataset is to train the **EAGLE3 architecture** within the **Ling-Flash-2.0** model. It provides high-quality, multi-task instruction samples (covering Chat, Math, Code, and Instruction Following) that enhance the performance and stability of the speculative decoding process. --- ## Citation If you utilize this dataset in your research or application, please cite the following: ```bibtex @misc{lingflash2025regenerate, title={Ling-Flash-2.0-open-perfectblend-regenerate: High-Quality Synthetic Instructions for EAGLE3}, author={Ant AQ Team}, year={2025}, }



