遇见数据集

Dataset for: Language of Persuasion and Misrepresentation in Business Communication: A Textual Detection Approach

收藏
Zenodo2025-08-04 更新2026-05-26 收录
官方服务:

资源简介:

Annotated Corpus of Business Communications: A Multidimensional Resource for Deception Detection Research Introduction This manuscript presents a meticulously curated dataset comprising 4,848 annotated business communication texts, designed to address a critical lacuna in domain-generalizable deception detection systems. As identified by Baden et al. (2022), existing frameworks often fail to capture the nuanced linguistic patterns of persuasive and misleading discourse across heterogeneous business contexts. Our corpus bridges this gap through systematic collection, expert annotation, and rigorous preprocessing of authentic business communications, thereby enabling robust computational analysis of deceptive language strategies. Methodological Framework The corpus encompasses textual artifacts drawn from ecologically valid business communication ecosystems, including: Marketing emails and social media advertisements (Facebook, Twitter/X) Financial reports and fraudulent corporate statements Newspaper articles and LinkedIn professional communications YouTube advertisements and international business reports Marketing criteria documentation This multi-source approach ensures comprehensive representation of communicative strategies across organizational hierarchies, industries, and stakeholder interactions. The diversity of sources mitigates sampling bias and enhances the corpus's ecological validity for real-world applications. Three domain experts with complementary specializations (business communication, theoretical linguistics, and computational linguistics) conducted the annotation process using a dual-coder methodology: Independent Annotation: Each text underwent parallel evaluation by two experts Consensus Resolution: Discrepancies were resolved through deliberative consensus or adjudication by a third expert Theoretical Grounding: Annotations incorporated established deception indicators including: Uncertainty expressions and hedging devices Extreme positive sentiment inflation Non-immediate linguistic markers Strategic framing techniques (Larcker & Zakolyukina, 2012; Craig et al., 2013) The annotation framework synthesized seminal deception detection paradigms (Zhou & Zhang, 2008; Humpherys et al., 2011) while adapting them to contemporary business communication contexts. The corpus exhibits a deliberately balanced distribution across three target classes: Class Instances Percentage Factual 1,980 40.8% Persuasive 1,479 30.5% Misleading 1,389 28.7% Total 4,848 100% This distribution addresses the persistent challenge of class imbalance in deception detection tasks (Glockner et al., 2022), facilitating robust model training without artificial weighting interventions. We implemented a methodical five-stage normalization protocol to enhance computational tractability while preserving semantic integrity: Case Standardization: Uniform conversion to lowercase to ensure lexical consistency URL Excision: Systematic removal of hyperlinks using regex patterns (http§+|www§+|https§+) Social Media Artifact Elimination: Removal of platform-specific markers (@mentions, #hashtags) Non-alphanumeric Character Filtering: Elimination of punctuation and numerical characters Whitespace Normalization: Standardization of inter-token spacing This pipeline effectively addresses the "curse of dimensionality" in textual analysis while maintaining the semantic features essential for deception detection. Dataset Architecture The corpus is provided in two complementary formats: Primary Corpus (raw_dataset.csv): Structure: Text | Label | Aspect Content: Original unmodified texts with expert annotations Processed Corpus (cleaned_dataset.csv): Structure: Text | Label | Aspect | clean_text Content: Preprocessed texts alongside original annotations The Aspect dimension provides granular categorization across six business communication domains: Product marketing Financial reporting Investment claims Health/environmental claims Risk disclosures Legal disclosure Research Applications This corpus enables multifaceted investigations in computational linguistics and business communication studies: Deception Detection Systems: Training and evaluation of machine learning models for automated identification of misleading business communications Linguistic Pattern Analysis: Empirical examination of lexical, syntactic, and pragmatic markers across factual, persuasive, and deceptive discourse Cross-domain Validation: Assessment of model generalizability across diverse business communication contexts Ethical Communication Frameworks: Development of guidelines for transparent business communication practices Educational Applications: Pedagogical resources for business ethics and professional communication training Ethical Considerations We acknowledge the profound ethical implications of automated deception detection technologies. Researchers utilizing this corpus must consider: Potential biases in annotation and source representation Societal consequences of false positives/negatives in applied systems Privacy preservation in business communication analysis Cultural variations in deceptive communication markers Conclusion This annotated corpus represents a significant methodological contribution to computational linguistics and business communication research. Its rigorous construction, balanced design, and ecological validity provide an essential resource for developing theoretically grounded and empirically validated deception detection systems. By making this dataset available to the research community, we aim to foster interdisciplinary collaboration and advance the development of responsible, transparent communication technologies in business contexts.

提供机构:
Sayem Hossen
创建时间:
2025-08-04
二维码
社区交流群
二维码
科研交流群
商业服务