遇见数据集

aimlresearch2023/ClimbMix100

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

ClimbMix100是ClimbMix1K数据集的100个样本子集,保持了原始数据集中各簇的比例。数据集格式为单个Parquet文件,包含text(文本)、cluster_id(簇ID)和topics(主题)三列。总样本数为100,使用种子42确保采样过程的可重复性。数据集中包含20个簇,每个簇至少有一个样本,样本分布覆盖了数学、历史、教育、健康、技术、环境、体育等多个主题领域,具体比例通过最大余数法计算得出,以确保比例保持。

ClimbMix100 is a 100-sample subsample of the ClimbMix1K dataset, preserving the original ratio of clusters. The dataset is in a single Parquet file with columns [text, cluster_id, topics]. Total samples: 100, with seed 42 for deterministic sampling. It includes 20 clusters, each with at least one sample, covering topics such as Mathematics, History, Education, Health, Technology, Environment, Sports, and more. The sampling method uses the largest-remainder method with minimum-1 enforcement to maintain proportional representation.

提供机构:
aimlresearch2023
二维码
社区交流群
二维码
科研交流群
商业服务