five

Coronavirus Corpus

收藏
DataCite Commons2023-04-27 更新2025-04-16 收录
下载链接:
https://dataverse.ucla.edu/citation?persistentId=doi:10.25346/S6/6WMNQU
下载链接
链接失效反馈
官方服务:
资源简介:
The Coronavirus Corpus contains about 1.5 billion words of data in approximately 1.9 million texts from Jan 2020 - Dec 2022, and it is designed to be the definitive record of the social, cultural, and economic impact of the coronavirus (COVID-19) during this time. The corpus allows you to see the frequency of words and phrases month by month and even day by day since January 2020, such as social distancing, flatten the curve, WORK * home, Zoom, Wuhan, hoard*, toilet paper, curbside, pandemic, reopen, defy, anti-mask*. Access to material is limited to UCLA graduate students and faculty. Undergraduates please use the standard web interface for the corpora: https://www.english-corpora.org/corona/

冠状病毒语料库(Coronavirus Corpus)包含约15亿词量的数据,涵盖2020年1月至2022年12月间的近190万篇文本,旨在构建该时段内冠状病毒(COVID-19)对社会、文化及经济影响的权威档案。该语料库支持按月度乃至逐日查询2020年1月以来的词汇与短语使用频次,覆盖社交距离(social distancing)、拉平疫情曲线(flatten the curve)、居家办公(WORK * home)、Zoom、武汉、囤积类词汇(hoard*)、卫生纸、路边服务(curbside)、大流行病(pandemic)、重新开放(reopen)、违抗(defy)、反口罩类词汇(anti-mask*)等主题。本语料库仅对加州大学洛杉矶分校(UCLA)的研究生与教职员工开放,本科生请使用标准语料库网页界面:https://www.english-corpora.org/corona/
提供机构:
UCLA Dataverse
创建时间:
2023-04-14
搜集汇总
数据集介绍
main_image_url
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务