hermion-qmsum-interaction-dynamics
收藏资源简介:
Hermion QMSum Interaction Dynamics 数据集是一个结构化交互动态数据集,从QMSum会议语料库中35个真实企业产品团队提取。其核心创新在于生成过程完全不依赖语义内容分析:原始会议转录文本通过开源的Hermion分类器在本地编码为数学信号,仅这些信号被发送至基于物理模型的交互智能引擎处理,确保了隐私并避免了文本传输或存储。这构成了首个无需访问语义内容、从真实企业会议数据中提取的结构化交互动态数据集,输出具有确定性、结构化和可复现的特点。数据集包含三个CSV文件,分别对应三项纵向研究:研究1(35行)分析管理角色与设计角色在团队完整项目周期中的交互;研究2(70行)分别考察项目经理和市场营销角色与整个团队的交互,衡量各执行角色对团队的结构性影响;研究3(280行)涵盖每个团队中八种以执行角色为核心的配对组合,包括从执行二人组到个体角色配对,再到执行角色与整个团队的交互。数据集共包含31个字段,详细描述了交互的多个方面,包括团队标识、配对类型、项目阶段、会话数量、交互进展百分比、峰值进展、动量状态、兴趣水平、拓扑模式与标签、主要阻力类型与构成比例(结构阻力、不对称阻力)、多种风险评分(幽灵风险、拒绝风险、停滞风险)、参与者积极信号比例、退缩行为模式、信号份额、下一步行动建议、不匹配检测与描述、恢复能力、分组标签、处理的信号数量以及经过窗口算法过滤后的消息数量。关键发现揭示了项目经理在交互动态中的结构性放大效应、项目经理与用户界面设计师配对产生最高平均峰值、团队交互峰值通常出现在开发阶段、所有团队均表现出以结构阻力为主的普遍模式、近半数团队在衰退事件前可检测到跨参与者不匹配信号,以及团队组合交互产生的非线性和耦合效应。数据集采用纵向设计,将每个团队的所有会话按时间顺序(从启动到开发再到结论)串联,形成一个连续的交互弧。数据处理采用了窗口算法,在隔离参与者配对时,仅保留原始对话流中对方组附近特定范围内的消息,以聚焦于真实的交换而非所有消息。整个处理流水线完全开源,可复现。数据集适用于交互动力学、团队智能、组织行为、会议智能、群体智能、多参与者系统以及基于物理模型的专业团队分析等研究任务。
The Hermion QMSum Interaction Dynamics dataset contains structured interaction dynamics data extracted from 35 real-world enterprise product teams in the QMSum meeting corpus. Its core innovation lies in its generation process, which does not rely on semantic content analysis: the original meeting transcripts are locally encoded into mathematical signals using the open-source Hermion classifier, and only these signals are sent to a physics-model-based interaction intelligence engine for processing, ensuring privacy and avoiding text transmission or storage. This constitutes the first structured interaction dynamics dataset extracted from real enterprise meeting data without accessing semantic content, with outputs that are deterministic, structured, and reproducible, contrasting with unstructured LLM-based methods. The dataset includes three CSV files corresponding to three longitudinal studies: Study 1 (35 rows) analyzes the interaction between management and design roles across the teams full project cycle; Study 2 (70 rows) examines the interaction of project managers and marketing roles with the entire team, measuring the structural impact of each executive role on the team; Study 3 (280 rows) covers eight executive role-centric pairing combinations per team, ranging from executive dyads to individual role pairings to executive roles interacting with the entire team. The dataset contains 31 fields detailing various aspects of interactions, including team identifier, pairing type, project phase, session count, interaction progress percentage, peak progress, momentum state, interest level, topological patterns and labels, primary resistance types and composition ratios (structural resistance, asymmetric resistance), multiple risk scores (ghost risk, rejection risk, stagnation risk), participant positive signal ratio, retreat behavior patterns, signal share, next action suggestions, mismatch detection and description, resilience, grouping labels, number of processed signals, and message count filtered by window algorithm. Key findings reveal the structural amplification effect of project managers in interaction dynamics, the highest average peak from project manager and UI designer pairings, team interaction peaks typically occurring in the development phase, a universal pattern of structural resistance dominance across all teams, detectable cross-participant mismatch signals before decline events in nearly half of the teams, and nonlinear and coupling effects from team combination interactions. The dataset employs a longitudinal design, concatenating all sessions of each team in chronological order (from initiation to development to conclusion) to form a continuous interaction arc. Data processing uses a window algorithm that, when isolating participant pairings, retains only messages within a specific range near the other group in the original dialogue flow to focus on genuine exchanges rather than all messages. The entire processing pipeline is fully open-source and reproducible. The dataset is suitable for research tasks such as interaction dynamics, team intelligence, organizational behavior, meeting intelligence, swarm intelligence, multi-participant systems, and physics-model-based professional team analysis.
Hermion QMSum Interaction Dynamics 数据集详情
数据集概览
该数据集包含从 35个真实企业产品团队 的 QMSum 会议语料中提取的结构化交互动态信息,共涵盖 385次智能运行、3项研究,全程未读取任何消息内容。数据通过 Hermion 基于物理的交互智能引擎处理生成,每条记录均为实时 API 调用的结果,具有确定性、结构化和可复现的特点。
- 许可证: CC BY 4.0
- 任务类别: 文本分类、其他
- 语言: 英语
- 数据规模: n<1K
- 标签: interaction-dynamics、team-intelligence、organizational-behavior、meeting-intelligence、group-intelligence、physics-based、multi-actor、professional-teams、qmsum、longitudinal
数据创新点
这是首个在不访问语义内容的情况下,从真实企业会议数据中提取交互动态的结构化数据集。原始转录文本在研究者本机通过 Hermion 开源分类器转换为编码数学信号,仅信号传输至智能引擎,全程未传输或存储消息文本。相比基于 LLM 的方法(产生非结构化、非确定性、不可比较的输出),本数据集具有完全的可复现性。
数据结构
数据集包含三个 CSV 文件,对应三项研究:
| 文件 | 行数 | 内容描述 |
|---|---|---|
summary_study1_longitudinal.csv |
35行 | 管理层(项目经理+市场营销)vs 设计组(工业设计师+用户界面),覆盖团队完整纵向时间弧 |
summary_study2_longitudinal.csv |
70行 | 项目经理单独 vs 团队(35次运行)、市场营销单独 vs 团队(35次运行),衡量各高管角色对团队整体的结构性影响 |
summary_study3_longitudinal.csv |
280行 | 每个团队8种以高管锚定的配对组合,涵盖高管二元组、个体角色配对及高管vs全团队等所有组合 |
字段说明
数据集包含 31个字段,核心字段包括:
- 标识字段:
meeting_id(项目团队标识)、run_label(配对类型)、phase(阶段,均为 longitudinal) - 交互进程:
progression_pct、peak_pct(峰值进度)、momentum(动量状态:progressing/declining/critical_decline/gliding/stalled/normalizing) - 拓扑与阻力:
topology_mode、topology_label、primary_resistance(主要阻力类型)、dominant_component(主导阻力成分)、structural_share、asymmetric_share - 风险评估:
ghosting_risk(消失风险)、rejection_risk(拒绝风险)、stall_risk(停滞风险),均为0-1评分 - 行为信号:
self_positive_ratio、other_positive_ratio、self_withdrawing、other_withdrawing、self_signal_share - 预测与诊断:
next_move(推荐行动)、other_next_move(预测对方行动)、mismatch_detected、mismatch_desc、recovery_capacity - 元数据:
session_count、signal_count、prepared_messages
关键发现
- 项目经理结构性放大效应: 项目经理产生的峰值动态平均比同等市场营销配对高 +10.5个百分点,在 PM/UI 二元组中差距最大(+13.1pp)。
- PM/UI 二元组表现最优: 项目经理与用户界面设计师的直接关系产生所有配对中最高的平均峰值(25.5%),高于高管配对及 PM vs 全团队。
- 峰值生命周期: 团队在项目结束时,峰值与当前进度之间平均相差 4.2倍。峰值集中在开发阶段(平均5.6%,高于启动阶段的4.9%和结论阶段的5.1%)。
- 普遍的结构性摩擦: 100%的团队(35个)以结构性摩擦为主要阻力类型,结构性份额平均达 87.6%。
- 衰退前信号: 48% 的团队在衰退事件前检测到可识别的跨角色错配,模式一致表现为管理层倾向收尾/推进交易,设计层倾向退出。
- 非线性耦合: 团队 IS1007 中,PM vs 设计组合的峰值达85%,而 PM vs 工业设计师单独仅33%、PM vs 用户界面单独仅45%,组合解锁了单一配对无法产生的峰值。
方法论
- 数据来源: QMSum 产品领域,35个团队均采用一致的4角色标注(项目经理、市场营销、工业设计师、用户界面)。
- 纵向设计: 每个团队的所有会话按时间顺序拼接(启动→开发→结论),形成连续交互弧,会话数3-4个,每团队总轮次915-3,936,无需采样。
- 窗口算法: 配对中角色被隔离时,仅保留原始流中与对方组位置相差 ±N 内的消息,N = min(max(groupSpacingA, groupSpacingB), 20),用于隔离真实交流。
- 隐私保护: 原始 QMSum 转录在本地处理,无消息文本传输至 Hermion 引擎,分类器在研究者机器上将消息编码为数学信号。
- 可复现性: 完整开源流水线位于
https://github.com/hermionai/hermion-research,可凭 Hermion API 密钥复现全部结果。
引用与相关资源
- 研究流水线:
https://github.com/hermionai/hermion-research - 完整论文:
https://github.com/hermionai/hermion-research/tree/main/research/qmsum-product-teams - SSRN 预印本:
https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7167718 - 开源分类器:
https://github.com/hermionai/hermion-classifier-os - Hermion 官网:
https://hermionai.xyz - 引擎试用:
https://hermionai.xyz/try - API 密钥获取:
https://hermionai.xyz/get-started - 底层 QMSum 数据集:
https://github.com/Yale-LILY/QMSum(单独许可)




