遇见数据集

SlangTrack (ST) Dataset

收藏
Zenodo2025-02-05 更新2026-05-26 收录
官方服务:

资源简介:

The SlangTrack (ST) Dataset is a novel, meticulously curated resource aimed at addressing the complexities of slang detection in natural language processing. This dataset uniquely emphasizes words that exhibit both slang and non-slang contexts, enabling a binary classification system to distinguish between these dual senses. By providing comprehensive examples for each usage, the dataset supports fine-grained linguistic and computational analysis, catering to both researchers and practitioners in NLP. Key Features: Unique Words: 48,508 Total Tokens: 310,170 Average Post Length: 34.6 words Average Sentences per Post: 3.74 These features ensure a robust contextual framework for accurate slang detection and semantic analysis. Significance of the Dataset: Unified Annotation: The dataset offers consistent annotations across the corpus, achieving high Inter-Annotator Agreement (IAA) to ensure reliability and accuracy. Addressing Limitations: It overcomes the constraints of previous corpora, which often lacked differentiation between slang and non-slang meanings or did not provide illustrative examples for each sense. Comprehensive Coverage: Unlike earlier corpora that primarily supported dictionary-style entries or paraphrasing tasks, this dataset includes rich contextual examples from historical (COHA) and contemporary (Twitter) sources, along with multiple senses for each target word. Focus on Dual Meanings: The dataset emphasizes words with at least one slang and one dominant non-slang sense, facilitating the exploration of nuanced linguistic patterns. Applicability to Research: By covering both historical and modern contexts, the dataset provides a platform for exploring slang's semantic evolution and its impact on natural language processing. Target Word Selection: The target words were carefully chosen to align with the goals of fine-grained analysis. Each word in the dataset: It coexists in the slang SD wordlist and the Corpus of Historical American English (COHA). Has between 2 and 8 distinct senses, including both slang and non-slang meanings. Was cross-referenced using trusted resources such as: Green's Dictionary of Slang Urban Dictionary Online Slang Dictionary Oxford English Dictionary Features at least one slang and one dominant non-slang sense. Excludes proper nouns to maintain linguistic relevance and focus. Data Sources and Collection: 1. Corpus of Historical American English (COHA): Historical examples were extracted from the cleaned version of COHA (CCOHA). Data spans the years 1980–2010, capturing the evolution of target words over time. 2. Twitter: Twitter was selected for its dynamic, real-time communication, offering rich examples of contemporary slang and informal language. For each target word, 1,000 examples were collected from tweets posted between 2010–2020, reflecting modern usage. Dataset Scope: The final dataset comprises ten target words, meeting strict selection criteria to ensure linguistic and computational relevance. Each word: Demonstrates semantic diversity, balancing slang and non-slang senses. Offers robust representation across both historical (COHA) and modern (Twitter) contexts. The SlangTrack Dataset serves as a public resource, fostering research in slang detection, semantic evolution, and informal language processing. Combining historical and contemporary sources provides a comprehensive platform for exploring the nuances of slang in natural language. Data Statistics: The table below provides a breakdown of the total number of instances categorized as slang or non-slang for each target keyword in the SlangTrack (ST) Dataset. Keyword Non-slang Slang Total BMW 1083 14 1097 Brownie 582 382 964 Chronic 1415 270 1685 Climber 520 122 642 Cucumber 972 79 1051 Eat 2462 561 3023 Germ 566 249 815 Mammy 894 154 1048 Rodent 718 349 1067 Salty 543 727 1270 Total 9755 2907 12662 Sample Texts from the Dataset: The table below provides examples of sentences from the SlangTrack (ST) Dataset, showcasing both slang and non-slang usage of the target keywords. Each example highlights the context in which the target word is used and its corresponding category. Example Sentences Target Keyword Category Today, I heard, for the first time, a short scientific talk given by a man dressed as a rodent...! An interesting experience. Rodent Slang On the other. Mr. Taylor took food requests and, with a stern look in his eye, told the children to stay seated until he and his wife returned with the food. The children nodded attentively. After the adults left, the children seemed to relax, talking more freely and playing with one another. When the parents returned, the kids straightened up again, received their food, and began to eat, displaying quiet and gracious manners all the while. Eat Non-Slang Greater than this one that washed between the shores of Florida and Mexico. He balanced between the breakers and the turning tide. Small particles of sand churned in the waters around him, and a small fish swam against his leg, a momentary dark streak that vanished in the surf. He began to swim. Buoyant in the salty water, he swam a hundred meters to a jetty that sent small whirlpools around its barnacle rough pilings. Salty Non-Slang Mom was totally hating on my dance moves. She's so salty. Salty Slang **Licenses** The SlangTrack (ST) dataset is built using a combination of licensed and publicly available corpora. To ensure compliance with licensing agreements, all data has been extensively preprocessed, modified, and anonymized while preserving linguistic integrity. The dataset has been randomized and structured to support research in slang detection without violating the terms of the original sources. The **original authors and data providers retain their respective rights**, where applicable. We encourage users to **review the licensing agreements** included with the dataset to understand any potential usage limitations. While some source corpora, such as **COHA, require a paid license and restrict redistribution**, our processed dataset is **legally shareable and publicly available** for **research and development purposes**.

俚语追踪(SlangTrack, ST)数据集是一项新颖且经过精心构建的资源,旨在解决自然语言处理(Natural Language Processing, NLP)领域中俚语检测的复杂问题。该数据集的独特之处在于重点收录同时存在俚语语境与非俚语语境的词汇,可支持二分类系统对这两种语义进行区分。通过为每种用法提供详尽示例,该数据集可支撑细粒度的语言学与计算分析,面向自然语言处理领域的研究人员与从业者。 核心特征: 唯一词汇量:48508 总Token数:310170 平均单帖长度:34.6个词 平均单帖句子数:3.74 上述特征可为精准俚语检测与语义分析提供稳固的语境支撑框架。 数据集的研究价值: 统一标注:本数据集为整个语料库提供了一致的标注,实现了较高的标注者间一致性(Inter-Annotator Agreement, IAA),确保了数据集的可靠性与准确性。 解决现有局限:它弥补了此前语料库的不足——此前的语料库往往无法区分俚语与非俚语语义,或未为每种语义提供示例。 覆盖范围全面:与此前仅支持词典式词条或释义改写任务的语料库不同,本数据集包含来自历史语料(美国历史语料库,Corpus of Historical American English, COHA)与当代语料(Twitter)的丰富语境示例,且为每个目标词汇提供多种语义。 聚焦双语义:本数据集重点收录同时具备至少一种俚语语义与主流非俚语语义的词汇,便于探索细微的语言学模式。 研究适用性:通过覆盖历史与现代语境,该数据集为探究俚语的语义演化及其对自然语言处理的影响提供了研究平台。 目标词汇遴选: 目标词汇经过严格遴选,以契合细粒度分析的研究目标。数据集中的每个词汇均满足以下条件: 1. 同时收录于俚语SD词表与美国历史语料库(Corpus of Historical American English, COHA); 2. 拥有2至8种不同语义,涵盖俚语与非俚语含义; 3. 通过以下权威资源进行交叉验证:《格林俚语词典》《城市词典》《在线俚语词典》《牛津英语词典》; 4. 至少包含一种俚语语义与一种主流非俚语语义; 5. 排除专有名词,以确保研究的语言学相关性与聚焦性。 数据来源与采集: 1. 美国历史语料库(Corpus of Historical American English, COHA): 历史语料取自经过清洗的COHA版本(CCOHA),采集时段为1980年至2010年,可捕捉目标词汇随时间的语义演化。 2. Twitter: Twitter以其动态的实时交流特性,可提供丰富的当代俚语与非正式语言示例。本数据集为每个目标词汇采集了2010年至2020年间发布的1000条推文作为现代语料。 数据集范围: 最终数据集包含10个目标词汇,均经过严格遴选以确保其语言学与计算学相关性。每个目标词汇均满足: - 具备语义多样性,平衡俚语与非俚语语义; - 在历史(COHA)与现代(Twitter)语境中均有充分的样本覆盖。 俚语追踪数据集作为公开资源,可推动俚语检测、语义演化及非正式语言处理领域的研究。结合历史与当代语料的构建方式,为探索自然语言中俚语的细微语义差异提供了全面的研究平台。 数据统计: 下表展示了俚语追踪(ST)数据集中每个目标关键词的俚语与非俚语标注实例总数: | 关键词 | 非俚语实例数 | 俚语实例数 | 总实例数 | |-------|-------------|-----------|---------| | BMW | 1083 | 14 | 1097 | | Brownie | 582 | 382 | 964 | | Chronic | 1415 | 270 | 1685 | | Climber | 520 | 122 | 642 | | Cucumber | 972 | 79 | 1051 | | Eat | 2462 | 561 | 3023 | | Germ | 566 | 249 | 815 | | Mammy | 894 | 154 | 1048 | | Rodent | 718 | 349 | 1067 | | Salty | 543 | 727 | 1270 | | 总计 | 9755 | 2907 | 12662 | 数据集示例文本: 下表展示了俚语追踪(ST)数据集中的例句,呈现了目标关键词的俚语与非俚语用法,每个示例均标注了目标词汇的使用语境与对应类别。 | 示例句子 | 目标关键词 | 类别 | |---------|-----------|-----| | 今日我首次聆听了一场由装扮成啮齿动物的男士带来的简短科学讲座……!这是一段有趣的经历。 | Rodent | 俚语 | | 另一方面,泰勒先生记录了用餐需求,眼神严肃地告诉孩子们待在座位上,直到他和妻子取回食物。孩子们专注地点头。大人离开后,孩子们似乎放松下来,自由交谈并互相玩耍。当父母返回时,孩子们重新坐直,接过食物并全程展现出安静得体的举止。 | Eat | 非俚语 | | 这片海域介于佛罗里达与墨西哥海岸之间。他在碎浪与涨潮间保持平衡。周围海水中翻搅着细碎的沙粒,一条小鱼游过他的腿,转瞬即逝的深色身影消失在浪花中。他开始游泳,在咸咸的海水中漂浮,游了一百米抵达一座码头,码头附着藤壶的粗糙桩柱周围形成了小型漩涡。 | Salty | 非俚语 | | 妈妈完全吐槽了我的舞蹈动作,她太刻薄了。 | Salty | 俚语 | 许可证: 俚语追踪(ST)数据集整合了授权与公开可用的语料库构建而成。为遵守授权协议,所有数据均经过深度预处理、修改与匿名化处理,同时保留了语言学完整性。数据集经过随机化与结构化处理,可支持俚语检测相关研究,且未违反原始数据源的使用条款。 原作者与数据提供者保留各自的相关权利(如适用)。我们鼓励用户查阅数据集附带的授权协议,以了解潜在的使用限制。部分源语料库(如COHA)需要付费授权并限制再分发,但本数据集的处理版本可合法共享并公开获取,用于研究与开发用途。

提供机构:
Zenodo
创建时间:
2024-10-15
二维码
社区交流群
二维码
科研交流群
商业服务