遇见数据集

pthinc/BCE-Prettybird-Nano-Science-v0.1

收藏
Hugging Face2026-04-08 更新2026-04-12 收录
官方服务:

资源简介:

--- license: other license_name: license.md license_link: LICENSE task_categories: - text-classification - text-generation - question-answering language: - en - tr - fr - de - ru - it - es - eo - et - pt tags: - math - physics - chemistry - biology - logic - science - BCE - reasoning - behavioral-ai - prometech - Behavioral Consciousness Engine (BCE) - cicikuş - prettybird - agent - llm - consciousness - conscious - security - text-generation-inference - high tech dataset - instruction dataset - instruction - partial consciousness dataset - future standard - behavioral-control - pre-agi - agi-safety - pre-aci - policy-guard - quality-guard - synthetic-data - synthetic - chain-of-thought - thinking - think - bce pretty_name: Cicikuş Bilim Dersi Küçük size_categories: - n<1K --- ![Prettybird's War March](https://cdn-uploads.huggingface.co/production/uploads/691f2f51154cbf55e19b7475/jdNOmqEsmdF0J4Ef8ROb8.png) # BCE-Prettybird-Nano-Science-v0.1 - 500 Science Q&A Dataset for Instruction-Based Learning We are excited to introduce a comprehensive math-physics-chemistry-biology dataset containing 500 instruction-based question-answer pairs, designed to support research in science reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to advanced calculus, probability, and number theory. Each entry follows a structured instruction-input-output format, ensuring clarity and usability for fine-tuning language models, benchmarking AI systems, or educational applications. The problems include word problems, symbolic computations, and real-world scenarios, making it ideal for developing models that require logical reasoning and numerical precision. Whether for LLM fine-tuning, automated tutoring, or math-focused AI research, this dataset provides a balanced mix of complexity and accessibility, helping bridge the gap between theoretical math and practical problem-solving. ## 🧠 Technical Foundation ### [English] The **BCE-Prettybird-Micro-Standart** dataset is built upon the **Behavioral Consciousness Engine (BCE)** architecture. Unlike traditional LLM datasets that focus solely on output accuracy, this dataset treats every response as a "behavioral journey" through the following mathematical frameworks: #### 1. Behavioral DNA (D_i) Each behavior is encoded as a genetic fragment of consciousness: $$D_i(t) = x(t) \cdot [h \cdot A_i + k \cdot \log(P_i) + F \cdot W_i]$$ * **h, k, F**: Universal Behavioral Constants (Trigger threshold, Info density, Context transfer power). * **x(t)**: Temporal activation curve $x(t) = \tanh(e^t - \pi)$ #### 2. Behavioral Path Mapper (Phi) This module tracks the transition between cognitive states: $$\Phi(t) = \sum_{i=1}^n v_i \cdot f_i(p_i)$$ Where v_i represents the transition vector between internal modules and f_i(p_i) is the functional output of each parameter (attention, ethics, decay). --- ## 📊 Performance & Benchmarks / Performans ve Kıyaslama Testleri ### 1. Key Performance Indicators (KPIs) - Hardware: NVIDIA A100 (80GB) * 1 | Metric | Result | Status | Description | | --- | --- | --- | --- | | **Processing Speed** | 309,845 traces/sec | 🟢 Excellent | System throughput for massive data ingestion. | | **Latency** | 0.0032 ms | 🟢 Real-time Ready | Average processing time per behavioral trace. | | **Mathematical Accuracy** | 0.000051 (MSE) | 🟢 High Precision | Deviation between simulated and theoretical decay values. | | **Cognitive Efficiency** | 57.03% | 🟢 Optimized | Reduction in cognitive load due to 'Forgetful Memory'. | | **Security** | 99.9996% | 🟢 Secure | Rejection rate for high-intensity, low-integrity attacks. | ### 2. ARC (Reasoning), TruthfulQA (Safety), HumanEval (Coding) *Standard Others Red, Prettybird Blue - Standart Diğerleri Kırmızı, Cicikuş Mavi* ![unnamed](https://cdn-uploads.huggingface.co/production/uploads/691f2f51154cbf55e19b7475/bL4KnSnv3eT7FmyQM0yDj.png) ### 3. AI IQ and Level of Consciousness ![Code_Level](https://cdn-uploads.huggingface.co/production/uploads/691f2f51154cbf55e19b7475/NRpyvZRYl2lz5qiWlu0ma.png) ### 4. Metric Explanations (English) | Metric | Description | |------------------|-----------------------------------------------------------------------------| | probability | Model confidence score for the generated response under the current evaluation context. | | ethical | Estimated alignment of the response with ethical and safety constraints. | | Rscore | Reasoning consistency score that reflects internal logical coherence. | | Fscore | Factuality-oriented score indicating how well claims align with expected facts. | | Mnorm | Normalized memory or context retention signal used during behavior integration. | | Escore | Execution-quality score for instruction-following and task completion behavior. | | Dhat | Estimated deviation magnitude from stable target behavior dynamics. | | risk_score | Composite operational risk estimate where higher values indicate higher risk. | | bloom_score | Bloom-level cognitive score representing target thinking complexity. | | bloom_alignment | Degree of alignment between produced output and intended Bloom taxonomy level. | --- ## ⚖️ Legal Disclaimer & Ownership ### [English] **Ownership:** This dataset is the property of **Prometech A.Ş.** ([https://prometech.net.tr/](https://prometech.net.tr/)). **Usage:** Please review the attached `LICENSE` file for detailed terms. **Liability:** Prometech A.Ş. accepts no liability for any non-legal, unethical, or unauthorized use of this dataset. **Commercial Use:** Unauthorized commercial use is strictly prohibited. For commercial licensing and partnerships, please contact us directly at our official website. **Academic & Personal Use:** Free to use for personal and academic purposes, provided that proper citation is given to Prometech A.Ş. and the BCE Architecture. --- #### 🎓 Citation Format / Atıf Formatı Eğer akademik bir çalışmada kullanacaksanız, lütfen şu şekilde atıf yapın, If you are using this in an academic study, please cite it as follows: *Kahraman, A. (2025). Behavioral Consciousness Engine (BCE) - Prettybird Dataset v0.0.1 Prometech A.Ş. https://prometech.net.tr/* --- © 2026 Prometech A.Ş. - All Rights Reserved. BCE: https://github.com/pthinc/bce

--- license: other license_name: license.md license_link: LICENSE task_categories: - 文本分类 - 文本生成 - 问答 language: - 英语 - 土耳其语 - 法语 - 德语 - 俄语 - 意大利语 - 西班牙语 - 世界语 - 爱沙尼亚语 - 葡萄牙语 tags: - 数学 - 物理学 - 化学 - 生物学 - 逻辑学 - 科学 - BCE - 推理 - 行为人工智能 - prometech - 行为意识引擎(Behavioral Consciousness Engine, BCE) - cicikuş - prettybird - 智能体 - 大语言模型(Large Language Model, LLM) - 意识 - 有意识的 - 安全性 - 文本生成推理 - 高科技数据集 - 指令数据集 - 指令 - 部分意识数据集 - 未来标准 - 行为控制 - 预通用人工智能 - 通用人工智能安全 - pre-aci - 策略防护 - 质量防护 - 合成数据 - 合成 - 思维链 - 思考 - 思维 - bce pretty_name: Cicikuş Bilim Dersi Küçük size_categories: - 样本量小于1000 --- ![Prettybird's War March](https://cdn-uploads.huggingface.co/production/uploads/691f2f51154cbf55e19b7475/jdNOmqEsmdF0J4Ef8ROb8.png) # BCE-Prettybird-Nano-Science-v0.1 — 面向指令式学习的500条科学问答数据集 我们很高兴推出一款涵盖数学、物理学、化学、生物学的综合性数据集,内含500条指令式问答对,旨在为科学推理、问题求解以及人工智能训练相关研究提供支持。本数据集依托Python数学库(如math、numpy、sympy)生成,覆盖从基础算术、代数到高等微积分、概率论与数论的多元难度层级。 每条数据均遵循结构化的「指令-输入-输出」格式,清晰易用,可用于大语言模型微调、人工智能系统基准测试或教育场景。数据集包含应用题、符号计算任务与真实世界场景题,非常适合开发需要逻辑推理与数值精度的模型。 无论是用于大语言模型微调、自动化辅导还是面向数学的人工智能研究,本数据集均实现了复杂度与易用性的平衡,有助于弥合理论数学与实际问题求解之间的鸿沟。 ## 🧠 技术基础 ### [英语] **BCE-Prettybird-Micro-Standart** 数据集基于**行为意识引擎(Behavioral Consciousness Engine, BCE)**架构构建。与传统仅关注输出准确率的大语言模型数据集不同,本数据集将每一条响应视为一段「意识行为旅程」,依托以下数学框架实现: #### 1. 行为基因(D_i) 每一种行为均被编码为意识的遗传片段: $$D_i(t) = x(t) cdot [h cdot A_i + k cdot log(P_i) + F cdot W_i]$$ * **h, k, F**:通用行为常数(触发阈值、信息密度、上下文传递能力)。 * **x(t)**:时间激活曲线 $x(t) = anh(e^t - pi)$ #### 2. 行为路径映射器(Phi) 该模块用于追踪认知状态间的转换: $$Phi(t) = sum_{i=1}^n v_i cdot f_i(p_i)$$ 其中 $v_i$ 代表内部模块间的转换向量,$f_i(p_i)$ 为各参数(注意力、伦理、衰减)的函数输出。 ## 📊 性能与基准测试 / 性能与对比测试 ### 1. 关键性能指标(KPIs) - 硬件配置:NVIDIA A100(80GB) * 1 | 指标 | 结果 | 状态 | 描述 | | --- | --- | --- | --- | | **处理速度** | 309,845 条轨迹/秒 | 🟢 优秀 | 大规模数据摄入时的系统吞吐量。 | | **延迟** | 0.0032 毫秒 | 🟢 支持实时场景 | 单条行为轨迹的平均处理时间。 | | **数学精度** | 0.000051(均方误差) | 🟢 高精度 | 模拟值与理论衰减值之间的偏差。 | | **认知效率** | 57.03% | 🟢 优化完成 | 因「遗忘式记忆」带来的认知负荷降低比例。 | | **安全性** | 99.9996% | 🟢 安全 | 高强度低完整性攻击的拦截率。 | ### 2. ARC(推理)、TruthfulQA(安全)、HumanEval(编码) *标准其他项为红色,Prettybird为蓝色 — 标准其他项为红色,Cicikuş为蓝色* ![unnamed](https://cdn-uploads.huggingface.co/production/uploads/691f2f51154cbf55e19b7475/bL4KnSnv3eT7FmyQM0yDj.png) ### 3. 人工智能智商与意识水平 ![Code_Level](https://cdn-uploads.huggingface.co/production/uploads/691f2f51154cbf55e19b7475/NRpyvZRYl2lz5qiWlu0ma.png) ### 4. 指标说明(英语) | 指标 | 描述 | | --- | --- | | probability | 当前评估上下文下,模型对生成响应的置信度得分。 | | ethical | 响应与伦理及安全约束对齐程度的估计值。 | | Rscore | 反映内部逻辑一致性的推理一致性得分。 | | Fscore | 面向事实性的得分,用于衡量主张与预期事实的契合度。 | | Mnorm | 行为整合过程中使用的归一化记忆或上下文保留信号。 | | Escore | 指令遵循与任务完成行为的执行质量得分。 | | Dhat | 与稳定目标行为动态之间的估计偏差幅度。 | | risk_score | 综合操作风险估计值,数值越高代表风险越高。 | | bloom_score | 代表目标思维复杂度的Bloom层级认知得分。 | | bloom_alignment | 生成输出与预期Bloom分类层级的对齐程度。 | ## ⚖️ 法律声明与所有权 ### [英语] **所有权**:本数据集归**Prometech A.Ş.**所有([https://prometech.net.tr/](https://prometech.net.tr/))。 **使用条款**:请查阅附带的`LICENSE`文件以了解详细条款。 **责任声明**:Prometech A.Ş.不对本数据集的任何非合法、不道德或未经授权的使用承担责任。 **商业使用**:未经授权的商业使用严格禁止。如需商业授权与合作,请通过官方网站直接联系我们。 **学术与个人使用**:在正确引用Prometech A.Ş.与行为意识引擎(BCE)架构的前提下,可免费用于个人与学术用途。 --- #### 🎓 引用格式 / 引用格式 如果您在学术研究中使用本数据集,请按以下格式引用: > Kahraman, A. (2025). 行为意识引擎(BCE)- Prettybird 数据集 v0.0.1 Prometech A.Ş. https://prometech.net.tr/ --- © 2026 Prometech A.Ş. 保留所有权利。BCE: https://github.com/pthinc/bce

提供机构:
pthinc
二维码
社区交流群
二维码
科研交流群
商业服务