遇见数据集

Verification artefact package for "After the Euclidean Highway: Hyperbolic Expert AI as the Next Innovation"

收藏
Zenodo2026-07-23 更新2026-08-01 收录
官方服务:

资源简介:

This deposit contains the corrected, self-contained verification artefact package for the HySAT (Hyperbolic Structure-Aware Training) placement principle developed across six expert small-language-model projects (17,954,911 cumulative training samples; approximately 317,000 cumulative optimizer steps; six completed HySAT training runs; zero NaN events in those completed runs). The evidence is disclosed at three explicit tiers. First, automated gates in verify_claims.py reproduce the H2 manifold-drift result (3,474/3,474 recorded steps within 1e-3; 81.95% within the stricter 1e-4 diagnostic threshold) and the H5 batch-size pair-formation gate (five evaluable b >= 8 rows; zero falsifiers). Second, released traces, evaluation aggregates, interval records, trainer states, hyperparameter configurations, matched ablations, and controlled toy/mechanism tests support direct inspection of the corresponding claims. Third, manifest and research-access records identify claims whose production artefacts or historical per-run terminal records were not retained publicly. The historical negative-result record is stated as 17 documented HyperLoRA-family training collapses, comprising NaN termination or unrecoverable divergence. The package does not claim that all 17 terminal events are independently observable as NaN in deposited per-run traces. One preserved early AdmitBrain E2 trajectory directly records an unrecoverable divergence plateau through step 13,000 (total loss approximately 4,288–4,339) and contains no textual NaN/Inf in the retained portion. A later E2 trainer-state record is a separate stabilized run. The included crash_reproduce.py is a controlled mechanism test, not a reconstruction of every historical collapse. Version 1.4 removes material associated with a rejected related manuscript, replaces overbroad crash/NaN language with evidence-matched collapse terminology, corrects the AdmitBrain E2 lineage, distinguishes automated verification from direct inspection and manifest/access-protocol evidence, and adds a correction notice plus manuscript-alignment record. These corrections do not change the reported dataset scope, the six completed-run zero-NaN result, the matched-ablation results, or the H2/H5 automated outcomes.

提供机构:
Zenodo
创建时间:
2026-07-23
二维码
社区交流群
二维码
科研交流群
商业服务