遇见数据集

R8 (Cognitive Risk Analyzer) v1.9 public data bundle: scoring outputs, annotation labels, and corpus metadata

收藏
Zenodo2026-09-30 更新2026-10-01 收录
官方服务:

资源简介:

R8 (Cognitive Risk Analyzer): An Exploratory Lexical Framework for Approximating Cognitive Manipulation Risk in Japanese Text R8 is an exploratory, dictionary-based framework that approximates cognitive manipulation risk in Japanese-language text across 12 theoretically grounded categories, producing a single composite score, the Cognitive Manipulation Index (CMI). The theoretical framework integrates Japanese and Western accounts of group influence, operationalizing both as the detection of linguistic signals targeting System 1 cognitive processing. Calibrated against a 205-document corpus (196 with valid CMI > 0), the standard detection mode (HIGH ≥ 41) yields Precision of 97.6%, Recall of 34.7%, and F1 of 51.2% for HIGH-risk binary classification. These are calibration metrics measuring agreement between the classifier and the author's own labels — not performance against a validated ground truth. Their interpretation is bounded by reliability: a two-model LLM pilot (Claude Sonnet 4.6, n = 196, κ = 0.241; Gemini 2.5 Flash, n = 192, κ = 0.111) and a preliminary two-rater human assessment (Cohen's κ = 0.094 and 0.197) all fall in the insufficient tier (κ < 0.35) of the pre-specified framework. The present data do not distinguish whether this reflects model or rater limitations or the structural difficulty of the annotation criteria; consolidating the criteria and establishing reproducible reliability is deferred to Phase 2, before any performance claim is advanced. High precision indicates that documents flagged HIGH tend to be so labelled by the author; it does not establish that R8 detects manipulation in any validated sense. This deposit contains the scoring data and document metadata, released under CC BY 4.0. The annotation criteria and source code are publicly available at https://github.com/takahiro-oss/r8-lexical-analyzer, the code under the MIT License and the documentation under CC BY 4.0. Original texts are not redistributed, and source URLs are withheld for every document; the released metadata does not identify the corpus documents. Use of Artificial Intelligence. Large language models were involved in this work in two distinct roles. As objects of measurement, Claude Sonnet 4.6 (Anthropic, 2026) and Gemini 2.5 Flash (Google, 2026) were evaluated as candidate annotators; their labels are reported as data. As tools, Claude (Anthropic) was used to implement the scoring engine and analysis scripts (including the script that computes the reported κ values), to perform the computations, to draft and revise the manuscript under the author's review, for Japanese–English translation, and for literature search; two further models, Claude Fable (Anthropic, 2026) and Gemini (Google, 2026), were each asked to report suspected errors in drafts of the manuscript, which the author adjudicated against the frozen data. Every reported value was recomputed from a read-only, SHA256-verified snapshot of the corpus and recorded in a claim-level evidence ledger; values not reproducible from the frozen data were removed. Scripts that require the corpus texts themselves - including the script behind the values reported in Section 5.8 - cannot be run by a reader. The hypotheses, risk categories, dictionary entries, corpus labels, and interpretations were determined by the author; no AI tool determined any of these, and no AI tool is an author of this work. Version note. This version corrects CODEBOOK.md (the scope of supplementary_table_s3.csv, the cluster-level derivations in Section 4.4, and the release statement in Section 8) and carries a regenerated MANIFEST.md; the other 23 files are byte-identical to the previous version (10.5281/zenodo.21928851). R8(Cognitive Risk Analyzer):日本語テキストの認知的操作リスクを近似する探索的語彙フレームワーク R8 は、日本語テキストの認知的操作リスクを12の理論的カテゴリにわたって近似する、辞書ベースの探索的フレームワークであり、単一の合成スコア Cognitive Manipulation Index (CMI) を算出する。理論的枠組みは、集団の影響に関する日本と西洋の論考を統合し、両者を、システム1の認知処理を標的とする言語的シグナルの検出として操作化(operationalize)したものである。 205文書のコーパス(うち有効 CMI > 0 は196件)に対する較正では、標準検出モード(HIGH ≥ 41)で HIGH リスク二値分類の Precision 97.6%、Recall 34.7%、F1 51.2% を得た。これらは分類器と著者自身のラベルとの一致を測る較正指標であり、検証済みの真値に対する性能ではない。その解釈は信頼性によって制約される。二モデルの LLM パイロット(Claude Sonnet 4.6, n = 196, κ = 0.241/Gemini 2.5 Flash, n = 192, κ = 0.111)および二評定者による予備的な人間評価(Cohen's κ = 0.094・0.197)は、いずれも事前規定した枠組みの不十分な水準(κ < 0.35)に位置する。現在のデータは、これがモデルや評定者の限界によるものか、アノテーション基準の構造的な難しさによるものかを判別しない。基準の統合と再現可能な信頼性の確立は Phase 2 に委ね、性能に関する主張はそれ以前には行わない。高い Precision は、HIGH と判定された文書が著者によっても HIGH とラベルされやすいことを示すにとどまり、R8 が検証された意味で操作を検出することを立証するものではない。 本デポジットはスコアリングデータおよび文書メタデータを収録し、CC BY 4.0 の下で公開する。アノテーション基準とソースコードは https://github.com/takahiro-oss/r8-lexical-analyzer にて公開されている(コードは MIT License、ドキュメントは CC BY 4.0)。原文テキストは再配布せず、全文書について source URL は非公開とする。公開されたメタデータからコーパス文書を特定することはできない。 人工知能の利用について。 本研究には大規模言語モデルが二つの異なる役割で関与している。測定対象としては、Claude Sonnet 4.6 (Anthropic, 2026) および Gemini 2.5 Flash (Google, 2026) をアノテーター候補として評価し、そのラベルはデータとして報告する。道具としては、Claude (Anthropic) を、スコアリングエンジンおよび分析スクリプト(報告 κ 値を計算するスクリプトを含む)の実装、計算処理の遂行、著者の点検を前提とした原稿の起草・改訂、日英翻訳、および文献検索に使用した。さらに Claude Fable (Anthropic, 2026) および Gemini (Google, 2026) には、それぞれ原稿の草稿の誤りの疑いを報告させ、著者が凍結データに対して判定した。報告した全ての数値は、読み取り専用の SHA256 照合済みコーパススナップショットから再計算し、主張単位の evidence ledger に記録した。凍結データから再現できなかった値は削除した。コーパステキスト自体を必要とするスクリプト(5.8 節の報告値を産出したものを含む)は、読者が実行することはできない。本研究における仮説、リスクカテゴリ、辞書項目、コーパスのラベル、および解釈は著者が決定した。これらを決定した AI ツールは存在せず、いかなる AI ツールも本研究の著者ではない。 版の注記。本版は CODEBOOK.md(supplementary_table_s3.csv の対象範囲、4.4 節のクラスタ単位の導出、8 節の公開範囲の記述)を訂正し、再生成した MANIFEST.md を収録する。その他の23ファイルは前の版(10.5281/zenodo.21928851)とバイト単位で同一である。

提供机构:
Zenodo
创建时间:
2026-09-30
二维码
社区交流群
二维码
科研交流群
商业服务