遇见数据集

e: Chinese AI Trust Index (CATI) Dataset: YouTube Audience Discourse on DeepSeek — Geopolitical Distrust vs Technical Trust Corpus (n=1,346, April 2026)

收藏
Zenodo2026-04-23 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains 1,346 YouTube comments collected viathe YouTube Data API v3 (April 2026) across five discoursedomains targeting naturalistic consumer responses toDeepSeek — the Chinese open-source AI model releasedJanuary 2025 — covering both geopolitical distrust andtechnical trust dimensions: (1) DeepSeek AI privacy data security concerns China 2025(2) Chinese AI DeepSeek vs American AI trust western 2025(3) DeepSeek AI dangerous China spying national security(4) DeepSeek AI reaction honest review should I use it(5) DeepSeek AI better than ChatGPT impressive technology The corpus is the primary empirical dataset for theChinese AI Trust Index (CATI) — a novel computationalmetric measuring the ratio of Geopolitical Distrust toTechnical Trust in naturalistic consumer discourse aboutChinese AI systems in Western digital environments. CATI = Geopolitical Distrust / (Geopolitical Distrust + Technical Trust)1.0 = pure Geopolitical Distrust0.5 = Contested zone — balanced discourse0.0 = pure Technical Trust Three signal types: GEOPOLITICAL DISTRUST (GD): national security, data sovereignty, China surveillance fears → "spying", "ban", "CCP", "national security risk" TECHNICAL TRUST (TT): performance, open source respect, capability recognition → "open source", "impressive", "better than ChatGPT" GEOPOLITICAL ANXIETY (GA): modifier signal — epistemic uncertainty about risk without clear distrust position → "worried about", "should I use it", "concerned" KEY FINDINGS:- Total corpus: n=1,346 comments, ~50 videos- Categorised (CATI computed): 200 (14.9%)- Mean CATI: 0.5363 — Contested zone, marginally Geopolitical Distrust dominant- GD density: 0.7091/100 tokens- TT density: 0.5454/100 tokens- GA density: 0.0180/100 tokens (minimal)- Geopolitical Distrust cluster: 104 (52.0%)- Technical Trust cluster: 89 (44.5%)- Contested: 7 (3.5%)- Geopolitical Anxiety: 6 (3.0%) DOMAIN-LEVEL CATI GRADIENT:- Privacy/data security concerns: CATI=0.8264 (highest GD) GD=6.63/100 | TT=0.95/100 — privacy frame maximises distrust- vs American AI trust: CATI=0.7321 GD=8.16/100 | TT=1.68/100 — geopolitical comparison dominant- National security/spying: CATI=0.7121 GD=7.78/100 | TT=2.16/100 — security frame sustains distrust- Honest review/should I use it: CATI=0.3571 GD=2.47/100 | TT=6.10/100 — pragmatic frame favours TT- Better than ChatGPT/technology: CATI=0.2576 (lowest GD) GD=1.74/100 | TT=4.67/100 — technical frame inverts to TT CRITICAL INVERSION FINDING (mirrors ASAI):When consumers frame DeepSeek through privacy/securitylens → CATI=0.8264 (Geopolitical Distrust dominant).When framed through technical performance lens →CATI=0.2576 (Technical Trust dominant). Same AI system,opposite trust orientations depending on discourse frame. Most-liked Technical Trust comment (20,290 likes):"If the release of a single open source AI model isenough to crash the value of these companies..."→ Technical capability recognised independently of origin. Most-liked Geopolitical Distrust (14,701 likes):"Open AI convincing the world AI is expensive, thenChina's like, nah it's a coupon deal..." → Ironicframing that blends admiration with geopolitical unease. Files:- deepseek_videos.csv: ~50 unique videos metadata- deepseek_comments.csv: 1,346 raw comments- deepseek_results.csv: CATI scores + signal annotation Method: YouTube Data API v3. Python 3.12, April 2026.

提供机构:
Zenodo
创建时间:
2026-04-23
二维码
社区交流群
二维码
科研交流群
商业服务