GSMA/Telco-Common-Corpus
收藏资源简介:
Telco Common Corpus (TCC) 是一个包含约100亿标记的完全开放、免费许可的电信知识集合,涵盖科学文献、专利、开放数据和开放网络项目,每个文档都经过许可证和来源验证。该数据集源于GSMA为电信行业推动AI发展的努力,旨在解决当前模型在真实电信任务(如网络管理)中的不足。数据来源多样,包括3GPP预备文档、RFC规范、IEEE开放获取论文、美国及欧盟专利、维基百科等,总计10,034,365,675个标记。数据处理使用内部管道进行高级解析,并主要针对英语文档。
Telco Common Corpus (TCC) is a ten billion tokens collection of fully open, free licensed telecommunications knowledge (scientific literature, patents, open data, and open-web projects) with licence and provenance verified at a document-level. It stems from GSMAs effort to make AI work for the telecom sector, addressing shortcomings in current models for real telecom tasks. The dataset features diverse sources including 3GPP preparatory documents, RFC specifications, IEEE open access papers, US and EU patents, Wikipedia, and more, totaling 10,034,365,675 tokens. Processing involves advanced parsing of documents and is primarily in English.




