遇见数据集

Digital Platform Work Repository

收藏
Zenodo2026-07-31 更新2026-08-01 收录
官方服务:

资源简介:

This research includes a n=272 (en 122, es 104, pt 41, 5 bilingual) corpus documenting the gap between how data annotation platforms represent the work and how the workers describe it in actuality. The corpus merges three components: 1) an English-language corpus of platform marketing claims and worker testimony, 2) a Spanish- and Portuguese-language corpus capturing the recruitment ecosystem and worker discourse of Latin American annotation economy, and 3) an original survey of n=8 Latin American annotators. Sources span corporate marketing materials and recruitment messages, worker review platforms, investigative journalism in three languages, legal filings including three lawsuits against Scale AI and a US Department of Labor investigation, independent audits, and peer-reviewed scholarship. The analysis of this survey represented a census approximation, because public-facing platform claims are finite and enumerable for platforms sampled, but the numerator is a sample drawn from unbounded and unenumerable populations of worker accounts that rise yearly. A second coder independently coded a randomised 10% sub-sample (n=23) on all three variables. The inter-rater agreement was 76.2% for category and 81% for evidence tier, with disagreements resolved through discussion and the resulting rulings propagated across the full corpus, which the first author then verified for consistency.

提供机构:
Zenodo
创建时间:
2026-07-31
二维码
社区交流群
二维码
科研交流群
商业服务