遇见数据集

xLEARN: Event-level xAPI learning experience data with academic context

收藏
Zenodo2026-04-24 更新2026-05-26 收录
官方服务:

资源简介:

xLEARN is a large-scale, longitudinal dataset of event-level xAPI learning experience data with academic context in higher education. The dataset was derived from the university's Learning Management System operated within the Yaşar University Open and Distance Learning Application and Research Center and integrated with institutional records from the Student Information System. The dataset contains 18,114,648 event-level xAPI statements collected between 2019 and 2022. It combines fine-grained behavioral interaction records with contextual academic information, including pseudonymized student attributes, course metadata, instructor information, and final course grades. The dataset covers 9,905 users, including 9,358 students and 547 instructors, across 4,431 course offerings. It also includes 8 standardized xAPI verb types, 37 xAPI activity types, and 114,152 final grade records. xLEARN was prepared to support reproducible, process-oriented learning analytics and educational data mining research. The interaction data are provided in their original event-level form, without aggregation, feature engineering, or derived behavioral indicators. Each xAPI statement preserves the actor, verb, object, timestamp, context, and associated metadata, including the complete nested JSON representation in the full_statement field. Contextual data are distributed as separate CSV files and can be linked to interaction records through consistent pseudonymized identifiers. The dataset is organized as a structured data package containing Learning Record Store data, Student Information System data, a data dictionary, validation scripts, validation results, a README file, and license information. The main data files include lrs_statements.csv, lrs_agents.csv, lrs_verbs.csv, sis_students.csv, sis_courses.csv, sis_instructors.csv, and sis_final_grades.csv. These files can be linked using shared keys such as account_name, site_id, and verb_id. All personal identifiers were pseudonymized before release. The dataset uses UTF-8 encoded CSV files, ISO 8601 timestamps with timezone information, and xAPI version 1.0.3. Grades are reported on a 0–100 scale, and NULL values in grade fields indicate withdrawn students or cases where a final grade was not applicable. A structured multi-step validation process was applied to assess data quality. The validation procedures covered duplicate detection, referential integrity, temporal consistency, missing values, data sanity checks, xAPI verb–activity consistency, cross-source consistency between LRS and SIS data, record completeness, and user mapping. The validation/ directory includes deterministic scripts and a shell script for reproducing all reported validation checks. The dataset represents a lossless capture of available LMS events. It does not include learning activities occurring outside the LMS, and timestamps should be interpreted as recorded system interactions rather than direct measures of cognitive engagement or task duration. Since no derived features are provided, researchers are expected to define their own preprocessing, aggregation, feature construction, and modeling procedures according to their research aims. The xLEARN dataset is shared under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0). Users are free to share and adapt the dataset for any purpose, including research and teaching, provided that appropriate credit is given, a link to the license is provided, and any changes made are indicated.

提供机构:
Zenodo
创建时间:
2026-04-24
二维码
社区交流群
二维码
科研交流群
商业服务