ops-data
收藏资源简介:
ELLNS(English Language Learning in Naturalistic Settings)是一个大规模英语语言学习数据集,专门用于研究英语作为第二语言(ESL)学习者在自然环境中的学习过程。数据集收集自真实的在线语言学习平台,时间跨度为2013年至2018年,共包含来自9个不同国家的15,323名学习者的约1,000万条学习记录。数据内容涵盖学习者进行的各种学习活动,包括听力理解、阅读理解、词汇练习、语法练习等多种任务类型,每条记录详细记录了学习事件的具体信息:学习者执行的学习活动、使用的学习材料、学习成果(正确或错误)、响应时间以及时间戳。数据集提供了丰富的元数据信息:学习者特征方面包括年龄、性别、国家、母语、自我报告的英语水平等;学习材料特征包括材料难度等级、材料类型(如文章、对话、练习)、主题分类等;此外还包含会话标识符,用于追踪连续的学习会话。数据以JSON格式组织,每条记录代表一个独立的学习事件,便于机器处理和分析。该数据集特别适用于研究自我调节学习、个性化学习路径、学习行为模式分析、英语习得过程以及跨文化学习差异等研究领域,具体应用任务包括学习行为建模与预测、学习者特征分析、学习成果预测模型开发、个性化学习推荐系统构建、不同学习者群体间的比较研究等。所有数据均已进行匿名化处理以保护学习者隐私,且数据完全来自真实学习环境,反映了自然状态下的学习过程,而非受控实验环境。
ELLNS (English Language Learning in Naturalistic Settings) is a large-scale English language learning dataset specifically designed for studying the learning processes of English as a Second Language (ESL) learners in naturalistic settings. The dataset is collected from real online language learning platforms, spanning from 2013 to 2018, and includes approximately 10 million learning records from 15,323 learners across 9 different countries. The data content covers various learning activities conducted by learners, including listening comprehension, reading comprehension, vocabulary exercises, grammar exercises, and other task types. Each record details specific information about a learning event: the learning activity performed by the learner, the learning materials used, learning outcomes (correct or incorrect), response time, and timestamps. The dataset provides rich metadata: learner characteristics include age, gender, country, native language, self-reported English proficiency, etc.; learning material features include material difficulty level, material type (e.g., articles, dialogues, exercises), topic classification, etc.; it also includes session identifiers for tracking continuous learning sessions. The data is organized in JSON format, with each record representing an independent learning event, facilitating machine processing and analysis. This dataset is particularly suitable for research areas such as self-regulated learning, personalized learning paths, learning behavior pattern analysis, English acquisition processes, and cross-cultural learning differences. Specific application tasks include learning behavior modeling and prediction, learner characteristic analysis, learning outcome prediction model development, personalized learning recommendation system construction, and comparative studies among different learner groups. All data has been anonymized to protect learner privacy, and the data is entirely derived from real learning environments, reflecting natural learning processes rather than controlled experimental settings.
数据集名称
ops-data
许可证信息
- 许可证类型:其他(other)
- 许可证名称:ellns-1
- 许可证链接:https://huggingface.co/datasets/hmnsyrd/ops-data/blob/main/LICENSE
说明
该数据集详情页面未提供关于数据集的描述、用途、构成、规模、来源等具体信息,仅包含上述许可证相关内容。




