遇见数据集

Readability of U.S. food handler training materials: A natural language processing analysis of worker study guides and the federal Food Code

收藏
Zenodo2026-08-08 更新2026-08-13 收录
官方服务:

资源简介:

This archive contains the analysis code and computed results for the study "Readability of U.S. Food Handler Training Materials: A Natural Language Processing Analysis of Worker Study Guides and the Federal Food Code" by Morris Brako and Anirudh R. Naig (Iowa State University). This is version 2 and it supersedes version 1. Version 1 accompanied the original submission and contains values that have since been corrected. It remains available for the record but should not be used. The differences are substantive rather than cosmetic: the corpus has grown from 6 documents to 20, the statistical design has moved from the sentence to the document as the unit of analysis, and an extraction cap that restricted the 668-page Food Code to roughly 18 percent of its length has been removed. The study uses a natural language processing pipeline to measure the reading difficulty of 20 U.S. food safety documents: the 2022 FDA Food Code, 14 worker-facing food handler study guides spanning 11 jurisdictions in eight states, and five Spanish-language counterparts that are deposited but excluded from analysis, because English-calibrated readability formulas are not valid on Spanish text. For each English document it computes Flesch-Kincaid Grade Level, the Gunning Fog Index, SMOG, New Dale-Chall, lexical density and full sentence-length distributions. Document-level values are tested against plain-language targets using exact one-sample Wilcoxon signed-rank tests with bootstrap confidence intervals under both percentile and bias-corrected and accelerated methods. The comparison between the federal code and the worker guides is reported descriptively only, because the corpus contains a single federal document and document type is therefore confounded with document identity. Results are reported under three extraction rule sets, so that the dependence of each quantity on artifact-filtering choices is visible rather than hidden in a single reported figure. The archive also contains a complete audit of every sentence in the corpus segmented at 80 tokens or more, 417 in total, each classified as genuine connected prose or as one of eight artifact types, with the classification features recorded for every sentence. The archive includes 35 files organised into eight groups: retrieval, the jurisdiction census, extraction and analysis, validation, the three rule sets, inference, tables and figures, and an automated consistency check of reported values against the computed outputs. The analysed source documents are public records issued by U.S. federal, state and county authorities and are not redistributed here, since redistribution depends on each issuing authority's licence. The retrieval log records for every attempt the requested URL, UTC timestamp, HTTP status, final URL after redirects, content type, payload size, SHA-256 digest, PDF structural validity and page count, so a reader who retrieves a document independently can confirm byte identity with the copy analysed here. Note that several agency URLs that served these documents during the study period now return HTTP 404 while mirrors continue to serve byte-identical files; use the digests rather than the URLs. See README.md for full instructions.

提供机构:
Zenodo
创建时间:
2026-08-08
二维码
社区交流群
二维码
科研交流群
商业服务