CCAE(Corpus of Chinese-based Asian Englishes)是由北京科技大学创建的多品种语料库,包含六种基于中文的亚洲英语变体,总计3.4亿词,来源于44.8万份网络文档。该数据集旨在为亚洲英语(尤其是中文英语)研究提供首个公开可访问的资源,支持特定语言模型的构建和下游任务,如语言变体识别和词汇变异识别。创建过程中,研究团队通过定制的数据收集和清洗流程确保数据质量,同
This dataset provides digitized copies of the tables of contents of the two major Soviet journals for Oriental Studies (Vostokovedenie): Narody Azii i Afriki (Народы Азии и Африки, Peoples of Asia and