HereSay Voice AI Classifications Dataset (2026-Q2-v2)
收藏资源简介:
The HereSay Voice AI Classifications Dataset contains aggregate classifications and per-conversation labels derived from anonymized conversations between humans and a friendly general-purpose AI voice bot on the HereSay platform (heresay.live). This release (2026-Q2-v2) includes 10 classification dimensions across 54 conversations (2334 total turns), covering: OpenAI's Asking/Doing/Expressing trichotomy, voice-specific facets (mic test, practice session, casual chat, emotional support, information seeking), per-turn sentiment (VADER), question-type and pronoun usage distributions, conversation-arc transitions (Sankey-ready), and time-of-day patterns. Raw conversation text is NOT included. Only aggregate stats and per-conversation/per-turn labels. Conversation IDs are irreversible SHA-256 hashes; PII has been redacted via Microsoft Presidio with coreferent token replacement. Headline finding: Voice AI conversations in this sample skew heavily toward "Expressing" (social/emotional chat) compared to text-based AI baselines. Where OpenAI's analysis of 1.1M ChatGPT messages (NBER WP 34255) found 11% Expressing / 40% Doing / 49% Asking, our sample shows roughly 63% Expressing / 22% Doing / 15% Asking. Voice and text appear to be fundamentally different mediums for AI interaction. The dataset is published under the Open Data Commons Attribution License v1.0 (ODC-BY 1.0), matching the precedent set by WildChat (Allen Institute for AI). See LICENSE.txt inside the archive. Limitations: sample size is small (n=54). HereSay users are not representative of the general population. Voice data is not directly comparable to text-based studies without medium adjustment. Classifiers are heuristic regex matching, not LLM-mediated. Full caveats in METHODOLOGY.md. Canonical download (registered users): https://heresay.live/dataset/voice-ai. Companion blog post: What People Ask AI Voice Bots.



