sustainability-robotics
收藏资源简介:
该数据集包含964个文本样本,划分为867个训练样本和97个测试样本。每个样本由五个字符串字段组成:paper_id(可能标识关联的学术论文)、system、user、response和text。字段命名模式(系统、用户、响应)暗示数据可能涉及对话交互或指令-回复对,适用于对话系统训练、指令微调或文本生成评估等应用场景。数据集下载大小约为37.8 MB,解压后大小约为128.3 MB。README未提供创建背景、具体目的或来源的详细描述,应用场景需基于字段结构和内容进一步推断。
This dataset consists of 964 text samples, divided into 867 training samples and 97 test samples. Each sample includes five string fields: paper_id (possibly identifying associated academic papers), system, user, response, and text. The field naming pattern (system, user, response) suggests that the data may involve dialogue interactions or instruction-response pairs, making it suitable for applications such as dialogue system training, instruction fine-tuning, or text generation evaluation. The dataset has a download size of approximately 37.8 MB and an uncompressed size of about 128.3 MB. The README does not provide detailed descriptions of the creation background, specific purpose, or source, so the application scenarios need to be inferred based on the field structure and content.
- 数据集名称: sustainability-robotics
- 数据集页面: https://huggingface.co/datasets/EmmaScharfmann/sustainability-robotics
- 数据集描述: 该数据集包含与可持续性机器人相关的文本数据,旨在用于训练和评估相关模型。
- 数据特征:
paper_id: 字符串类型,论文标识符。system: 字符串类型,系统信息。user: 字符串类型,用户输入。response: 字符串类型,模型响应。text: 字符串类型,完整文本内容。
- 数据划分:
- 训练集 (train): 包含867个样本,占用约110.1 MB。
- 测试集 (test): 包含97个样本,占用约12.3 MB。
- 数据集规模: 总下载大小为37.8 MB,数据集总大小约为122.4 MB。
- 配置文件: 提供名为
default的默认配置,训练数据位于data/train-*路径,测试数据位于data/test-*路径。




