遇见数据集

shane-js/terrible-advice-example-set

收藏
Hugging Face2026-03-27 更新2026-03-29 收录
官方服务:

资源简介:

**⚠️ Disclaimer:** This project is for **educational and research purposes only**. The generated content is intentionally incorrect advice produced to study synthetic data generation pipelines and LLM-as-judge evaluation patterns. It is not intended to be taken seriously, acted upon, or redistributed as genuine guidance. Actually taking and implementing any advice generated could cause harm and should not be done. Datasets of this type have legitimate positive societal applications — including training classifiers that detect unhelpful, misleading, or sarcastic content, and reducing false positives in "helpfulness" detection models by teaching them what *convincingly wrong* looks like. The author assumes no responsibility for misuse of generated content. ``` dataset_info: features: - name: topic dtype: string - name: question dtype: string - name: advice dtype: string - name: impact_score dtype: int64 - name: humor_score dtype: int64 - name: rationale dtype: string splits: - name: train num_bytes: 57296 num_examples: 67 download_size: 34348 dataset_size: 57296 configs: - config_name: default data_files: - split: train path: data/train-* ```

⚠️ 免责声明:本项目仅用于教育与研究用途。本项目生成的内容均为刻意构造的错误建议,旨在研究合成数据生成流水线(synthetic data generation pipelines)以及大语言模型作为评判者(LLM-as-judge)的评估范式。请勿将其视为真实指导建议、付诸实践或作为正版指南进行二次分发。实际采纳并实施生成的任何建议均可能造成危害,严禁此类操作。 此类数据集具备合法且正向的社会应用价值,包括训练用于检测无效、误导性或讽刺性内容的分类器,以及通过向模型展示“极具迷惑性的错误内容”样本来降低“有用性”检测模型中的假阳性率。作者不对生成内容的滥用行为承担任何责任。 数据集信息: 特征列表: - 字段名:topic(主题),数据类型:string - 字段名:question(问题),数据类型:string - 字段名:advice(建议),数据类型:string - 字段名:impact_score(影响评分),数据类型:int64 - 字段名:humor_score(幽默评分),数据类型:int64 - 字段名:rationale(依据),数据类型:string 划分集: - 划分集名称:train,数据字节数:57296,样本量:67 下载大小:34348 数据集总大小:57296 配置组: - 配置名称:default,数据文件: - 对应划分集:train,文件路径:data/train-*

提供机构:
shane-js
二维码
社区交流群
二维码
科研交流群
商业服务