stanforddams/biglocal
收藏资源简介:
ICE新闻稿数据集是一个结构化数据集,包含965篇由美国移民和海关执法局(ICE)发布的新闻稿,这些新闻稿从ice.gov网站抓取而来。该数据集由斯坦福大学的Big Local News创建,为记者提供了一个机器可读的记录,涵盖了ICE关于执法行动的公开声明。与ICE的行政数据(仅记录逮捕、驱逐和转移,但缺乏个人细节)不同,新闻稿通常包含个人姓名、年龄、国籍、指控以及以叙事形式描述的执法行动具体细节。该数据集以结构化字段捕获这些文本信息,并附带提取的元数据,适用于命名实体识别、信息提取和问责新闻业。语料库主要集中在2025年和2026年(965条记录中的956条),反映了该时期ICE执法活动和新闻稿发布量的激增。数据集语言为英语,许可为美国政府作品(公共领域)。
A structured dataset of 965 press releases published by U.S. Immigration and Customs Enforcement (ICE), scraped from ice.gov. Created by Big Local News at Stanford University, this dataset gives journalists a machine-readable record of ICEs own public statements about enforcement actions. ICE administrative data logs arrests, removals, and transfers but includes very little detail about individuals. Press releases are different: ICE routinely names people, lists ages, nationalities, and charges, and describes the specifics of each action in narrative form. This dataset captures that text in structured fields alongside extracted metadata, making it useful for named entity recognition, information extraction, and accountability journalism. The corpus is heavily skewed toward 2025 and 2026 (956 of 965 records), reflecting the surge in ICE enforcement activity and press release volume during that period. Language(s) (NLP): English. License: U.S. Government Works (public domain).




