遇见数据集

aixk/vlite3.6-dataset

收藏
Hugging Face2026-05-18 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个用于自然语言处理任务的数据集,包含训练集,共有787,890个样本,总大小约为863.5 MB。数据集特征包括input_ids(uint16列表类型,表示输入序列的编码)和loss_mask(bool列表类型,用于指示损失计算中的掩码部分),适用于序列建模或掩码语言建模等任务。数据以默认配置提供,文件路径为data/train-*。

This dataset is designed for natural language processing tasks, containing a training set with 787,890 examples and a total size of approximately 863.5 MB. The features include input_ids (a list of uint16, representing encoded input sequences) and loss_mask (a list of bool, indicating mask portions for loss calculation), suitable for tasks such as sequence modeling or masked language modeling. The data is provided in a default configuration with file paths at data/train-*.

提供机构:
aixk
二维码
社区交流群
二维码
科研交流群
商业服务