A Confound-Annotated Multi-Institution Curriculum Corpus with Parsed Prerequisite Logic
收藏资源简介:
This dataset is a machine-readable curriculum corpus assembled from the published catalogs of thirteen universities, centered on the Arabian Gulf (United Arab Emirates, Saudi Arabia, Qatar, Kuwait, Bahrain) and extended by the American University of Beirut and the California Institute of Technology as comparators. It comprises 32 catalog editions, 51,858 course records, and 39,709 prerequisite relations, released as two CSV tables (courses.csv, prereq_edges.csv) and a JSON summary (datasets.json), together with a data dictionary and a reference implementation of the analysis code. The dataset has three distinguishing properties. First, prerequisites are parsed into conjunctive-normal form, so the alternatives a catalog expresses with the word "or" are preserved as boolean structure rather than flattened into a list of mandatory codes. Second, one institution is covered by an eleven-edition panel spanning 2015-2016 to 2025-2026, which supports course-level longitudinal analysis without recourse to web archives. Third, every record is annotated with measurement-confound fields (notation drift, selective disclosure, subject-code renumbering, and prerequisite-operator ambiguity), each exposed as a filterable predicate so that diachronic and cross-institution analyses can control them rather than fall prey to them. The data support research on curricular structure, prerequisite-network complexity, academic advising and course planning, accreditation analysis, and the natural-language processing of course descriptions, and they cover a region that has been largely absent from existing curriculum datasets.




