遇见数据集

illume/Opus-4.6-Reasoning

收藏
Hugging Face2026-03-23 更新2026-03-29 收录
官方服务:

资源简介:

--- license: apache-2.0 --- [<img src="https://huggingface.co/crownelius/Crow-9B-Opus-4.6-Distill-Heretic_Qwen3.5/resolve/main/banner.png" width="350"/>](https://ko-fi.com/abcuo) # Opus-4.6-Reasoning-3000x (Cleaned) This dataset has been automatically cleaned to remove: - Empty or missing responses - Responses shorter than 10 characters - Refusal responses ("problem is incomplete", "cannot solve", etc.) - Responses with no substantive content - Responses that just echo the problem ## Cleaning Report - **Original rows:** 3,305 - **Clean rows:** 2,160 - **Removed:** 1,145 (34.6%) - **Columns:** ['id', 'problem', 'thinking', 'solution', 'difficulty', 'category', 'timestamp', 'hash'] ### Rejection Breakdown | Reason | Count | % | |--------|------:|---:| | refusal/incomplete_problem | 1,090 | 33.0% | | empty_response | 18 | 0.5% | | response_no_substance | 17 | 0.5% | | problem_too_short | 14 | 0.4% | | response_too_short | 6 | 0.2% | ## Usage ```python from datasets import load_dataset ds = load_dataset("crownelius/Opus-4.6-Reasoning-3000x", split="train") ``` *Cleaned on 2026-02-12 21:49 UTC* ### This dataset was extremely expensive, around $500 CAD I would really appreciate a follow or a small contribution to make more of these as models are released. -https://ko-fi.com/abcuo --- ## Stats | Metric | Value | |--------|-------| | Total prompt tokens | 0 | | Total completion tokens | 1,636,368 | | Total tokens | 1,636,368 | | Total cost | $40.91 (USD) | | Average turns | 1.00 | | Average tool calls | 0.00 | | Average tokens per row | 757.58 | *Cost estimated using Claude Opus 4.6 pricing on [OpenRouter](https://openrouter.ai) ($5.0/M input, $25.0/M output)*

--- 许可证:Apache 2.0 --- [<img src="https://huggingface.co/crownelius/Crow-9B-Opus-4.6-Distill-Heretic_Qwen3.5/resolve/main/banner.png" width="350"/>](https://ko-fi.com/abcuo) # Opus-4.6-Reasoning-3000x(清洗版) 本数据集已完成自动清洗,移除了以下内容: - 空或缺失的回复 - 长度不足10个字符的回复 - 拒答类回复(如"问题不完整""无法解决"等) - 无实质内容的回复 - 仅重复问题的回复 ## 清洗报告 - **原始条目数:** 3,305 - **清洗后有效条目数:** 2,160 - **剔除条目数:** 1,145(占比34.6%) - **数据集字段:** ['id', 'problem', 'thinking', 'solution', 'difficulty', 'category', 'timestamp', 'hash'] ### 拒答统计详情 | 拒答原因 | 数量 | 占比 | |--------|------:|---:| | 拒答/问题不完整 | 1,090 | 33.0% | | 空回复 | 18 | 0.5% | | 无实质内容回复 | 17 | 0.5% | | 问题过短 | 14 | 0.4% | | 回复过短 | 6 | 0.2% | ## 使用方式 python from datasets import load_dataset ds = load_dataset("crownelius/Opus-4.6-Reasoning-3000x", split="train") *清洗完成时间:2026年2月12日21:49 UTC* 本数据集的制作成本极高,约合500加元。随着后续模型的发布,若您能关注或提供小额资助以支持更多同类数据集的制作,我们将不胜感激。 - 资助链接:https://ko-fi.com/abcuo --- ## 统计信息 | 指标 | 数值 | |--------|-------| | 总提示词Token(Token) | 0 | | 总补全Token(Token) | 1,636,368 | | 总Token(Token) | 1,636,368 | | 总成本 | 40.91美元(USD) | | 平均对话轮次 | 1.00 | | 平均工具调用次数 | 0.00 | | 单条数据平均Token(Token)数 | 757.58 | *成本估算基于Claude Opus 4.6在[OpenRouter](https://openrouter.ai)平台的定价标准(输入Token每百万5美元,输出Token每百万25美元)*

提供机构:
illume
二维码
社区交流群
二维码
科研交流群
商业服务