遇见数据集

HuNeBR: A Dataset of Annotated Humorous Transcriptions from YouTube Shorts by Northeastern Brazilian Comedians

收藏
Zenodo2026-05-04 更新2026-05-26 收录
官方服务:

资源简介:

HuNeBR is a dataset containing 475 humor-based texts transcribed from YouTube Shorts featuring comedians from Brazil’s Northeast, sourced from videos published between April 10, 2022, and September 9, 2024. The material was compiled by selecting well-known regional comedians, identified through their media exposure and public acclaim. YouTube Shorts were chosen as the source due to their concise format, which facilitates quicker processing. Transcriptions were initially produced using automated tools and subsequently refined through manual editing. Each entry includes metadata such as the performance context (e.g., podcast, stand-up), the comedian's state of origin, notable cultural references, punchlines, and a multi-label classification across eight humor styles (including fun, benevolent humor, nonsense, wit, irony, sarcasm, satire, and cynicism). Additionally, each text is accompanied by an in-depth explanation of its comedic elements. The annotation process adhered to a thorough, multi-phase protocol grounded in established academic frameworks. It involved a lead annotator and six independent reviewers working across three stages: initial annotation, crossed review, and final adjustments based on collective input. This process took place over three months (January–March 2025), ensuring high levels of precision and consistency through structured cross-checking and expert evaluation. The final dataset is presented in a structured CSV format with 17 columns, providing a robust foundation for linguistic, sociocultural, and computational studies of humor in Brazilian Portuguese. The folders listed below correspond to the data collection, annotation, and review phases. The final folder contains the main dataset (brazilian_ne_annotated_humorous_texts.csv), along with a PDF document that describes the columns present in all stages.

HuNeBR是一个包含475条幽默文本的数据集,这些文本转录自巴西东北部喜剧演员的YouTube短视频(YouTube Shorts),采集自2022年4月10日至2024年9月9日发布的相关视频。本次素材遴选通过媒体曝光度与公众认可度,筛选出知名的地域喜剧演员并取用其作品。选择YouTube短视频作为数据源,是因其格式简洁,可提升处理效率。转录工作先通过自动化工具生成初稿,随后经人工编辑完成润色。 每条数据条目均包含如下元数据:表演场景(例如播客、单口喜剧)、喜剧演员的出身州、典型文化指代、笑点,以及覆盖8种幽默风格的多标签分类(包括趣味、善意幽默、无厘头、机智、反讽、挖苦、讽刺与犬儒主义)。此外,每条文本还附带对其喜剧元素的深度解析。 标注流程遵循基于成熟学术框架制定的多阶段严谨规范,由1名主标注员与6名独立评审员分三阶段完成:初始标注、交叉审核,以及基于集体意见的最终调整。该流程历时3个月(2025年1月至3月),通过结构化交叉校验与专家评估,确保了极高的标注精度与一致性。 最终数据集以结构化CSV格式存储,共包含17列,可为巴西葡萄牙语幽默的语言学、社会文化学与计算学研究提供坚实的研究基础。 以下列出的文件夹对应数据采集、标注与审核阶段。最终文件夹包含主数据集(brazilian_ne_annotated_humorous_texts.csv),以及一份说明各阶段字段含义的PDF文档。

提供机构:
Zenodo
创建时间:
2025-05-27
二维码
社区交流群
二维码
科研交流群
商业服务