登录后查看消息通知
搜索
常见问题
消息
登录
首页
/
数据集
/
tokenizer and word embeddings
tokenizer and word embeddings
收藏
Figshare
2018-06-28 更新
2026-04-29 收录
文本分词
词向量表示
数据链接:
https://figshare.com/articles/dataset/tokenizer_and_word_embeddings/6713837
数据链接
链接失效反馈
官方服务:
问题咨询
购买咨询
在线客服
NEW
资源简介:
Word tokenizer and precomputed word embeddings
应用场景:
创建时间:
2018-06-28
相关数据集
Tristan/t5-small-october-wikipedia-2022-tokenized-512
自然语言处理
文本分词
--- dataset_info: features: - name: input_ids sequence: int32 - name: attention_mask sequence: int8 - name: special_tokens_mask sequence: int8 splits: - name: train num_byt
Hugging Face
2023-01-02 更新
16
0
GoogleNews-vectors-negative300
词向量表示
自然语言处理
GoogleNews-vectors-negative300
Figshare
2023-06-29 更新
7
0
hac541309/polyglot-ko-tokenizer-corpus
韩语处理
文本分词
--- language: ko dataset_info: features: - name: text dtype: string splits: - name: train num_bytes: 17909910180 num_examples: 11808255 download_size: 9384042407 dataset_size:
Hugging Face
2023-07-13 更新
6
0
tokenizer and word embeddings
文本分词
词向量表示
Word tokenizer and precomputed word embeddings
Figshare
2018-06-28 更新
8
0
GoogleNews-vectors-negative300 2.bin
词向量表示
自然语言处理模型
for word embedding
Figshare
2018-12-10 更新
7
0
© 2023-2026 上海数据发展科技有限责任公司 版权所有
沪ICP备17003045号-15
沪公网安备31010402336585号
热门搜索
社区交流群
科研交流群
商业服务
数据资源
寻源服务
数据采集
标注服务
数据产品
代理销售
数据领域
凭证登记
数据产品
介绍推广