登录后查看消息通知
搜索
常见问题
消息
登录
首页
/
数据集
/
14 million word corpus (.txt)
14 million word corpus (.txt)
收藏
kaggle
2023-01-15 更新
2024-03-07 收录
文本数据
语言分析
数据链接:
https://www.kaggle.com/datasets/luisgasparcordeiro/14-million-word-corpus-txt
数据链接
链接失效反馈
官方服务:
问题咨询
购买咨询
在线客服
NEW
资源简介:
Collection of news, reports and books in text format
应用场景:
创建时间:
2023-01-15
相关数据集
spsither/tibetan_monolingual_A_merged_123_lines
藏文
文本数据
--- dataset_info: features: - name: text dtype: string splits: - name: train num_bytes: 46558973496 num_examples: 209309183 - name: test num_bytes: 1629513756 num_example
Hugging Face
2024-04-24 更新
24
0
dgambettaphd/D_gen8_run1_llama2-7b_sciabs_doc1000_real32_synt96_vuw
文本数据
机器学习
该数据集包含1000个样本,每个样本有一个唯一的ID和对应的文档内容。数据集分为一个训练集,总大小为429457字节。下载大小为209087字节,数据集总大小为429457字节。默认配置下的数据文件路径为data/train-*。
Hugging Face
2024-12-18 更新
6
0
1817_AA00047511_TXT.zip
历史文献
文本数据
Extracted text of 1817 issues of Barbados Mercury in TXT format.
DataCite Commons
2021-09-28 更新
5
0
M1keR/pgbooks-nl
自然语言处理
文本数据
该数据集包含文本和token数量两个字段,提供了训练集划分,共有950个样本,数据集大小为347,428,673字节,下载大小为211,944,305字节。
Hugging Face
2025-10-23 更新
7
0
ngoc2018/twitter_dataset_1717631161
社交媒体分析
文本数据
--- dataset_info: features: - name: id dtype: string - name: tweet_content dtype: string - name: user_name dtype: string - name: user_id dtype: string - name: created_at
Hugging Face
2024-06-05 更新
5
0
© 2023-2026 上海数据发展科技有限责任公司 版权所有
沪ICP备17003045号-15
沪公网安备31010402336585号
热门搜索
社区交流群
科研交流群
商业服务
数据资源
寻源服务
数据采集
标注服务
数据产品
代理销售
数据领域
凭证登记
数据产品
介绍推广