axentx/surrogate-2-biz-mastery
收藏资源简介:
该数据集是一个对话格式的数据集,包含162,903个训练样本,总大小为约624.6 MB。每个样本由两个主要部分组成:messages和metadata。messages字段是一个列表,每个元素包含角色(role,字符串类型)和内容(content,字符串类型),用于表示多轮对话中的交互信息。metadata字段为字符串类型,可能包含样本的附加元数据。数据集仅提供train分割,适用于自然语言处理任务,如对话生成、模型训练等。
This dataset is a dialogue-formatted dataset containing 162,903 training examples with a total size of approximately 624.6 MB. Each example consists of two main components: messages and metadata. The messages field is a list where each element includes a role (string type) and content (string type), representing interactions in multi-turn dialogues. The metadata field is of string type and may contain additional metadata for each example. The dataset only provides a train split and is suitable for natural language processing tasks such as dialogue generation and model training.




