DLT-Corpus
收藏资源简介:
DLT-Corpus是由伦敦大学学院等机构构建的分布式账本技术领域最大规模文本集合,包含29.8亿Tokens的跨学科数据。该数据集整合科学文献(3.7万篇)、美国专利(4.9万项)和社交媒体(2200万条)三大来源,通过语义检索和领域过滤确保内容相关性,其关键词密度达到通用语料的8.7倍。该资源支持技术创新扩散分析、市场情绪追踪等应用,为区块链领域的自然语言处理研究提供了首个综合性文本基础设施。
DLT-Corpus is the largest-scale text collection in the distributed ledger technology (DLT) field, constructed by institutions including University College London (UCL). It contains 2.98 billion tokens of interdisciplinary data, integrating three major sources: 37,000 scientific papers, 49,000 U.S. patents, and 22 million social media posts. Content relevance is ensured via semantic retrieval and domain filtering, with its keyword density reaching 8.7 times that of general-purpose corpora. This resource supports applications such as technological innovation diffusion analysis and market sentiment tracking, providing the first comprehensive text infrastructure for natural language processing research in the blockchain domain.



