A 160B bilingual long-text dataset with 3 categories: holistic, aggregated and chaotic long texts.(万卷长文是一个160B 的双语长文本数据集,分为 3 类:整体长文本、聚合长文本和混沌长文本)