遇见数据集

singhalrk/combined-wildchat-llama-3.1-8b-responses

收藏
Hugging Face2026-05-02 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个多语言对话数据集,包含超过101万个训练样本,总大小约2.93 GB。数据特征包括对话哈希、模型标识、时间戳、对话轮次、语言类型、毒性标记、编辑状态、国家信息、哈希IP地址、用户查询、描述标签模型、描述文本、意图标签模型、原始输出、类别列表、完成原因、样本索引、模型响应、思考过程和响应完成原因等字段。数据集可能用于AI对话模型的训练或评估,涉及内容安全分类(如毒性检测)和意图分析。

This dataset is a multilingual dialogue dataset containing over 1.01 million training samples, with a total size of approximately 2.93 GB. Features include conversation hash, model identifier, timestamp, turn count, language type, toxicity flag, redaction status, country information, hashed IP address, user query, description label model, description text, intention label model, raw output, category list, finish reason, sample index, model response, thinking process, and response finish reason. The dataset is likely intended for training or evaluating AI dialogue models, involving content safety classification (e.g., toxicity detection) and intention analysis.

提供机构:
singhalrk
二维码
社区交流群
二维码
科研交流群
商业服务