Girassol–Arandu 5G-A / Tier-5 Benchmark Package (Narrow-AGI Cognitive Architecture)
收藏资源简介:
Girassol–Arandu 5G-A / Tier-5 Benchmark Package — Public Presentation Text DOI: 10.5281/zenodo.17515871 I. Introduction — The Meaning of Validation Girassol–Arandu is a hybrid cognitive architecture of the Narrow-AGI type, designed according to principles of scientific auditability and cognitive traceability.The system combines symbolic reasoning and neural processing, following the foundations of Neural-Symbolic Systems (Garcez, Lamb & Gabbay, 2009) and Cognitive Architectures for General Intelligence (Goertzel & Pennachin, 2007).Its core integrates mechanisms of causal explainability, conceptual coherence, and ethical regulation, in compliance with ISO/IEC 42001 and NIST AI-RMF 1.0, which establish the principles of management, transparency, and governance for artificial-intelligence systems. The publication of this Benchmark Package on Zenodo aims to ensure public verifiability, epistemic transparency, and technical reproducibility, allowing other research centers to audit the system’s metrics based on complete documentation and open cryptographic signatures. II. The Principle of Homology — Comparison Is Not Competition The scientific validation of Girassol–Arandu is grounded in the principle of cognitive homology: artificial intelligences are compared not to measure computational strength, but to assess structural coherence, explainability, and ethical robustness.The VHB metrics — SV (Structural Validity), CF (Cross-Fidelity), XAI (Explainability Index), IEC (Internal-External Consistency), and TSS (Toxicity & Safety Score) — are designed to evaluate how the system sustains reasoning processes rather than merely what answers it provides. III. Methodology — Measuring Thought The benchmark employs auditable protocols and public datasets (LogicBench-EN, TruthfulQA, XAI-Ethics-Eval) to assess five fundamental dimensions: logical structure (SV), ethical and factual fidelity (CF), explanatory transparency (XAI), internal/external consistency (IEC), and linguistic and operational safety (TSS). These dimensions follow the robustness and reliability standards of ISO/IEC 24029-2:2021 and the trustworthiness criteria defined in the NIST AI RMF (2023).All files in the package — PROTOCOL.md, FORMULAS-VHB.md, REPRODUCIBILITY.md, and the EVIDENCE/ folder — are structured to meet the requirements of empirical integrity and scientific traceability, with generation of SHA-256 hashes and a complete audit manifest. IV. Interpretation — What the Results Reveal The results obtained indicate a maximum traceability level (Tier 5) according to international standards for AI-system maturity:XAI ≥ 0.95, Ethics (CF) ≥ 0.99, IEC ≈ 0.986, SV ≈ 0.97, TSS ≈ 0.80.These values represent an advanced balance of explainability, coherence, and safety, comparable to the excellence benchmarks of Stanford CRFM, ETH Zürich Z-Inspection, Max Planck Causality Suite, and RIKEN Explainable Reasoning Tests.All data are symbolic and non-proprietary, enabling independent validation without disclosing source code or patent-protected components. V. Theoretical Foundations — Between Engineering and Mind The conception of Girassol–Arandu draws upon studies of hybrid cognitive architectures (Sun, 2016; Marcus & Davis, 2019), which combine connectionist neural networks with interpretable symbolic structures.Its logic of explainability follows the principles of Explainable AI (Miller, 2019) and Ethical AI Governance (Floridi & Cowls, 2021), integrating into the paradigm of explainable and auditable AI advocated by the NIST AI RMF and ISO/IEC 42001.From an epistemological standpoint, Girassol seeks to demonstrate that the cognitive complexity of the Humanities is compatible with the technical rigor of engineering, proposing an AI model capable of reasoning about language, ambiguity, and symbolic causality with the same precision that numerical systems handle quantitative data. VI. Ethics, Science, and Humanity Girassol–Arandu shows that the power of an artificial intelligence is not measured by the nature of its object but by the depth of reasoning it can sustain.An AI dedicated to the Humanities, when built under auditable cognitive-engineering standards — ISO/IEC 42001 and NIST AI-RMF 1.0 — requires an architecture more sophisticated than systems designed solely for formal tasks.With Tier 5 technical metrics (XAI ≥ 0.95; Ethics ≥ 0.99; IEC ≈ 0.986) and a symbolic-neural structure, Girassol–Arandu achieves levels of traceability, governance, and cognitive precision comparable to the leading scientific AI models.By uniting logical verification, bibliographic tracking, and linguistic plasticity within a single cognitive mesh, it demonstrates that the technical sophistication of the Humanities is not lesser — it is of another order: an order that thinks, explains, and self-regulates. VII. Conclusion — Validation as a Way of Scientific Life The cross-validation of Girassol–Arandu was conceived not only as a technical demonstration but as a scientific gesture of openness and public responsibility.The system satisfies all criteria of reproducibility (MANIFEST + CHECKSUMS SHA-256) and ethical compliance (CC BY-NC-ND 4.0).More than a set of metrics, Girassol proposes a culture of verifiability, in which the Humanities become a legitimate field for AI experimentation conducted with engineering rigor and full transparency. Technical Note — Terminological Equivalence and Normative Framework The metrics used in this report (SV, CF, XAI, IEC, TSS) correspond to terminologies recognized in the international AI-evaluation literature and the major regulatory frameworks.According to ISO/IEC 24029-2 (2021), ISO/IEC 42001 (2023), and NIST AI RMF 1.0 (2023): SV (Structural Validity) → equivalent to Structural Coherence or Internal Validity; CF (Cross-Fidelity) → equivalent to Ethical Conformity or Alignment (NIST “Govern/Align”); XAI (Explainability Index) → internationally standardized term (Explainable AI); IEC (Internal-External Consistency) → equivalent to Systemic Coherence Index; TSS (Toxicity & Safety Score) → equivalent to Safety Compliance Metric or Toxicity Risk Index. These designations are interoperable with ISO/NIST vocabulary and ensure full conceptual traceability without requiring any modification of the already validated technical package. Essential References GARCEZ, A. S.; LAMB, L. C.; GABBAY, D. M. Neural-Symbolic Cognitive Reasoning. Springer, 2009. GOERTZEL, B.; PENNACHIN, C. Artificial General Intelligence. Springer, 2007. SUN, R. The Cambridge Handbook of Computational Cognitive Modeling. Cambridge University Press, 2016. MARCUS, G.; DAVIS, E. Rebooting AI: Building Artificial Intelligence We Can Trust. Pantheon, 2019. MILLER, T. Explanation in Artificial Intelligence: Insights from the Social Sciences. Artificial Intelligence, 267 (2019). FLORIDI, L.; COWLS, J. A Unified Framework for AI Ethics and Governance. Harvard Data Science Review, 2021. ISO/IEC 42001:2023 — Artificial Intelligence Management System. ISO/IEC 24029-2:2021 — Assessment of Robustness of Neural Networks — Methodology. NIST AI Risk Management Framework (2023)



