遇见数据集

tvarchive Dataset

收藏
Zenodo2021-07-05 更新2026-05-28 收录
数据链接:
官方服务:

资源简介:

The <code>tvarchive</code> dataset contains word-frequency and other non-consumptive-use data about 1,205,844 English-language transcriptions of U.S. television news broadcasts. The documents were scraped from the Internet Archive's TV News Archive, which includes automatic captions of select U.S. news broadcasts since 2009. While the complete TV News Archive contains over 2.2 million transcripts, WE1S researchers were only able to collect about 1.2 million documents containing complete transcripts. The full TV News Archive includes transcripts from 33 networks and hundreds of shows. Unlike other WE1S datasets, the <code>tvarchive</code> dataset was not collected using keyword searches for specific terms (i.e., documents containing the word "humanities"). <em>(See WE1S Research Materials Overview for the relation between the project's "datasets" and "collections.")</em>

提供机构:
Zenodo
创建时间:
2021-07-03
二维码
社区交流群
二维码
科研交流群
商业服务