遇见数据集

CBC/Corporati

收藏
Figshare2016-01-06 更新2026-04-08 收录
官方服务:

资源简介:

The Corporate Blogging Corpus (CBC/Corporati) was assembled between early 2006 and late 2007 as part of my dissertation on the style and pragmatics of corporate blogs (Puschmann, 2010). A list of 137 English-language company blogs were selected and categorized by their function and place in the organization (e.g. PR, marketing, company leadership). The blog posts were harvested over a period of 1.5 years using PHP, MySQL and MagpieRSS. Part of speech tagging was also performed, but this data is excluded from this data set. See the included README file and Puschmann (2010; available at http://blog.ynada.com/368) for a detailed analysis of the corpus. NOTE TO THE ADMIN: this data belongs into the category Social Science/Linguistics (presently missing). Please add if possible.

企业博客语料库(Corporate Blogging Corpus,CBC/Corporati)于2006年初至2007年末期间汇编完成,系笔者针对企业博客风格与语用学展开的学位论文(Puschmann, 2010)的组成部分。研究团队选取了137个英语企业博客,并按照其功能及在企业组织中的定位进行分类,例如公共关系(Public Relations,PR)类博客、市场营销类博客以及企业领导层博客。本数据集的博客文章采集工作耗时1.5年,通过PHP、MySQL及MagpieRSS工具完成。研究过程中还开展了词性标注(Part of speech tagging)工作,但该部分标注数据未被纳入当前数据集。如需了解该语料库的详细分析内容,请参阅随附的README文件以及Puschmann(2010)的相关研究成果(可通过链接http://blog.ynada.com/368获取)。致管理员提示:本数据集应归类至社会科学/语言学(Social Science/Linguistics)类目(当前该类目尚未添加),如条件允许请予以补充。

创建时间:
2011-12-30
二维码
社区交流群
二维码
科研交流群
商业服务