EnglishTense: A large scale English texts dataset categorized into three categories: Past, Present, Future tenses.
收藏资源简介:
he EnglishTense dataset is a comprehensive collection of English sentences meticulously categorized based on their tense: Past, Present, and Future. The dataset comprises a total of 13,316 annotated sentences with three distinct tense types: 4,621 in the present tense, 3,851 in the past tense, and 4,844 in the future tense. This dataset is designed to facilitate research and development in natural language processing (NLP) and computational linguistics, particularly for English, a widely spoken language in the world. With applications spanning tense detection, text classification, language modeling, and educational tools, EnglishTense is a valuable resource for the NLP community, facilitating advancements in temporal analysis and robust NLP model development.
EnglishTense 数据集是一套基于时态精心分类的全面英语语句集合,涵盖过去时、现在时与将来时三大类别。该数据集共包含13,316条标注语句,涵盖三类明确的时态类别:现在时4,621条、过去时3,851条、将来时4,844条。本数据集旨在推动自然语言处理(Natural Language Processing,简称NLP)与计算语言学领域的研究与开发,尤其针对全球使用广泛的英语这一语种。其应用场景涵盖时态检测、文本分类、语言建模与教育工具开发等领域,EnglishTense 是NLP社区的宝贵资源,可助力时序分析与鲁棒NLP模型开发的技术进步。



