bus_speeds_inner_hcm
收藏资源简介:
Chinese-LLM-Alignment是一个专为中文大语言模型(LLM)对齐研究而构建的数据集。该数据集旨在通过人类反馈来提升模型在指令遵循、无害性、有用性等方面的表现。数据来源于众包平台(如Amazon Mechanical Turk),由人类参与者针对多样化指令提供自然语言响应,并对多个模型生成的响应进行偏好判断。数据集规模约为10,000个指令-响应对,以及约100,000条人类偏好标签。数据以JSONL格式组织,每个指令条目包含原始指令、多个候选模型响应以及对应的人类偏好排序或评分。该数据集适用于训练和评估中文LLM的对齐能力,可用于监督微调、基于人类反馈的强化学习(RLHF)等任务。数据收集过程遵循伦理规范,已获得参与者知情同意,并进行了匿名化处理。
Chinese-LLM-Alignment is a dataset specifically constructed for alignment research of Chinese large language models (LLMs). It aims to enhance model performance in areas such as instruction following, harmlessness, and helpfulness through human feedback. The data is sourced from crowdsourcing platforms (e.g., Amazon Mechanical Turk), where human participants provide natural language responses to diverse instructions and make preference judgments on multiple model-generated responses. The dataset comprises approximately 10,000 instruction-response pairs and about 100,000 human preference labels. It is organized in JSONL format, with each instruction entry containing the original instruction, multiple candidate model responses, and corresponding human preference rankings or scores. This dataset is suitable for training and evaluating the alignment capabilities of Chinese LLMs and can be used for tasks such as supervised fine-tuning and reinforcement learning from human feedback (RLHF). The data collection process adheres to ethical standards, with informed consent obtained from participants and anonymization applied.
数据集名称
bus_speeds_inner_hcm
许可证
MIT(麻省理工学院开源许可证)
数据集描述
该数据集位于Hugging Face平台,具体内容涉及胡志明市(HCM)内部公交车的速度数据。数据集由用户 lnanhduc12 上传,未提供其他详细的描述或使用说明。
文件与结构
该数据集页面未提供详细的文件列表、数据格式、字段说明或样本数据。
用途与许可
数据集采用 MIT 许可证,允许自由使用、修改和分发,但需保留原版权声明。具体应用场景(如交通分析、城市研究等)需用户根据数据内容自行探索。




