遇见数据集

jupyter-agent/kaggle-notebooks-edu-v0

收藏
Hugging Face2024-10-11 更新2025-09-13 收录
官方服务:

资源简介:

--- dataset_info: features: - name: text dtype: string - name: id dtype: string - name: file_path dtype: string - name: response dtype: string - name: label dtype: int64 - name: contains_outputs dtype: bool splits: - name: error num_bytes: 31664883.4359375 num_examples: 112 - name: '0' num_bytes: 15266997.370898437 num_examples: 54 - name: '1' num_bytes: 128073144.61142579 num_examples: 453 - name: '2' num_bytes: 277915896.58505857 num_examples: 983 - name: '3' num_bytes: 1227296955.3161132 num_examples: 4341 - name: '4' num_bytes: 700020101.6730468 num_examples: 2476 - name: '5' num_bytes: 514837078.00751954 num_examples: 1821 download_size: 2107964624 dataset_size: 2895075057.0 configs: - config_name: default data_files: - split: error path: data/error-* - split: '0' path: data/0-* - split: '1' path: data/1-* - split: '2' path: data/2-* - split: '3' path: data/3-* - split: '4' path: data/4-* - split: '5' path: data/5-* --- # Kaggle Notebooks LLM Filtered - Model: `meta-llama/Meta-Llama-3.1-70B-Instruct` - Sample: `12,400` - Source dataset: `data-agents/kaggle-notebooks` - Prompt: ``` Below is an extract from a Jupyter notebook. Evaluate whether it has a high analysis value and could help a data scientist. The notebooks are formatted with the following tokens: START <text block> # Here comes markdown content <input block> <lang python> # Here comes python code <output block> # Here comes code output # More blocks END Use the additive 5-point scoring system described below. Points are accumulated based on the satisfaction of each criterion, so stop counting if any of the criteria is not fulfilled: - Add 1 point if the notebook contains valid code, even if it's not educational, like boilerplate code, configs, and niche concepts. - Add another point if the notebook successfully loads a dataset e.g. a CSV or JSON file, even if it lacks further analysis and contains the code outputs. - Award a third point if the notebook runs some analysis on the dataset by running statistics or plotting useful properties, even if they are mostly uncommented. - Give a fourth point if the majority of the notebook contains text between the code cells explaining insights and performing reasoning. - Give a fifth point if the notebook is clean and outstanding in it's analysis, creates insightful, explained plots and contains consistent, multi-step reasoning connected across the whole notebook and gains useful insights from the data. The extract: START {} END After examining the extract: - Briefly justify your total score, up to 100 words. - Conclude with the score using the format: "Educational score: <total points>" where <total points> is just a one digit number. ```

数据集信息: 特征字段: - 名称:text,数据类型:字符串型 - 名称:id,数据类型:字符串型 - 名称:file_path,数据类型:字符串型 - 名称:response,数据类型:字符串型 - 名称:label,数据类型:64位整型 - 名称:contains_outputs,数据类型:布尔型 数据拆分: - 拆分名称:error,字节大小:31664883.4359375,样本数量:112 - 拆分名称:'0',字节大小:15266997.370898437,样本数量:54 - 拆分名称:'1',字节大小:128073144.61142579,样本数量:453 - 拆分名称:'2',字节大小:277915896.58505857,样本数量:983 - 拆分名称:'3',字节大小:1227296955.3161132,样本数量:4341 - 拆分名称:'4',字节大小:700020101.6730468,样本数量:2476 - 拆分名称:'5',字节大小:514837078.00751954,样本数量:1821 下载大小:2107964624 数据集总大小:2895075057.0 配置项: - 配置名称:default,数据文件: - 拆分:error,路径:data/error-* - 拆分:'0',路径:data/0-* - 拆分:'1',路径:data/1-* - 拆分:'2',路径:data/2-* - 拆分:'3',路径:data/3-* - 拆分:'4',路径:data/4-* - 拆分:'5',路径:data/5-* # 经大语言模型(Large Language Model)筛选的Kaggle笔记本数据集 - 所用模型:`meta-llama/Meta-Llama-3.1-70B-Instruct` - 样本量:12,400 - 源数据集:`data-agents/kaggle-notebooks` - 提示词: 以下为一段Jupyter笔记本的节选内容,请评估其是否具备较高的分析价值,能否为数据科学家提供助力。 笔记本采用如下标记格式: START <文本块> # 此处为Markdown内容 <输入块> <lang python> # 此处为Python代码 <输出块> # 此处为代码输出 # 更多代码块 END 请使用下述加分制5分评分体系,每满足一项准则即可获得对应分值,若任意准则未满足则停止计分: - 若笔记本包含有效代码(即便不具备教学意义,如样板代码、配置项及小众概念代码),加1分。 - 若笔记本成功加载数据集(如CSV或JSON文件),即便缺乏后续分析且包含代码输出,再加1分。 - 若笔记本对数据集执行了分析操作(如运行统计或绘制有效属性图表),即便多数代码未添加注释,授予第3分。 - 若笔记本的大部分内容为代码单元格间的文本,用于阐释见解与开展推理,授予第4分。 - 若笔记本整洁且分析质量出众,生成了具备洞察力且附带说明的图表,且全笔记本贯穿连贯的多步推理,从数据中获取了有效见解,授予第5分。 节选内容: START {} END 审阅节选内容后: - 简要阐述总得分的理由,字数不超过100词。 - 以格式“Educational score: <总分>”收尾,其中<总分>为一位数字。

提供机构:
jupyter-agent
二维码
社区交流群
二维码
科研交流群
商业服务