遇见数据集

Minimal Dataset and Code for: Social friction vs. Cognitive efficiency: A comparative analysis of help-seeking behaviors in human communities and generative AI

收藏
Figshare2026-02-26 更新2026-04-28 收录
官方服务:

资源简介:

This repository contains the minimal dataset and author-generated Python code required to replicate the findings of the study: "Social friction vs. Cognitive efficiency: A comparative analysis of help-seeking behaviors in human communities and generative AI," submitted to PLOS ONE.The replication package includes:matched_corpus_N30000.csv: The complete, propensity-score-matched dataset consisting of 30,000 dialogue turns from LMSYS-Chat-1M (Human-AI) and Stack Exchange (Human-Human).analysis_and_figures_code.py: A comprehensive Python script used for text preprocessing, linguistic feature extraction (Hedges, Politeness markers, Vulnerability metrics, etc.), statistical analysis, and high-resolution figure generation.All data processing follows standard computational sociolinguistic protocols to ensure full transparency and reproducibility of the research.

本仓库包含复现该研究成果所需的极简数据集与作者原创Python代码,该研究论文为"Social friction vs. Cognitive efficiency: A comparative analysis of help-seeking behaviors in human communities and generative AI",已提交至PLOS ONE。本复现包包含以下内容: 1. `matched_corpus_N30000.csv`:完整的倾向得分匹配(propensity-score-matched)数据集,涵盖来自LMSYS-Chat-1M(人机对话)与Stack Exchange(人人对话)的30000轮对话回合。 2. `analysis_and_figures_code.py`:一款功能完备的Python脚本,可用于文本预处理、语言特征提取(含模糊限制语(Hedges)、礼貌标记(Politeness markers)、脆弱性指标(Vulnerability metrics)等)、统计分析以及高分辨率图表生成。 所有数据处理流程均遵循标准计算社会语言学(computational sociolinguistic)规范,以保障研究的完全透明性与可复现性。

创建时间:
2026-02-26
二维码
社区交流群
二维码
科研交流群
商业服务