llm-classification-distilled-v2-sharded
收藏资源简介:
该数据集名为LLM Classification Distilled v2 Sharded,是一个用于文本分类任务的英文数据集。它通过教师-评判者蒸馏流程生成,并以分片CSV文件的形式存储,共包含4个分片。该数据集是中间分片版本,最终处理后会合并并上传至多个不同的版本仓库,包括完整版、过滤版、安全过滤版以及教师困难版。数据集的具体内容、分类类别、样本数量及字段结构在README中未详细说明,其主要用途是作为大语言模型蒸馏流程中的中间产物,用于生成最终的训练数据。
The dataset named LLM Classification Distilled v2 Sharded is an English-language dataset for text classification tasks. It is generated via a teacher-judge distillation pipeline and stored as sharded CSV files, with a total of 4 shards. This is an intermediate sharded version. After final processing, it will be merged and uploaded to multiple version repositories including the full version, filtered version, safety-filtered version, and teacher-hard example version. The specific content, classification categories, sample count, and field structure of the dataset are not detailed in the README. Its primary purpose is to serve as an intermediate product in the LLM distillation pipeline for generating final training datasets.




