BeanCounter
收藏资源简介:
BeanCounter是由芝加哥大学创建的一个大规模、低毒性的商业文本数据集,包含超过1590亿个Tokens,主要来源于企业的公开披露文件。数据集通过EDGAR系统收集,经过清洗和去重处理,确保了数据的高质量和事实性。BeanCounter的应用领域主要集中在金融领域,旨在通过提供低毒性、高质量的数据来训练更优的金融领域语言模型,减少模型生成有害内容的风险。
BeanCounter is a large-scale, low-toxicity commercial text dataset developed by the University of Chicago. It contains over 159 billion tokens, primarily sourced from publicly disclosed corporate filings. The dataset is collected via the EDGAR system, and undergoes cleaning and deduplication processes to ensure high data quality and factual accuracy. Its main application focuses on the financial domain, aiming to train superior financial domain language models by providing low-toxicity, high-quality data, thereby reducing the risk of the models generating harmful content.



