数据链接:
官方服务:
资源简介:
Trained word2vec model
应用场景:
创建时间:
2022-11-18
相关数据集
Quick-Thought Vectors
Quick-Thought Vectors 数据集包含了一系列用于文本表示的向量,这些向量是通过Quick-Thought模型生成的。该模型通过预测句子是否紧随另一个句子来学习句子级别的表示,从而捕捉句子间的语义关系。数据集主要用于自然语言处理任务,如句子相似度计算、文本分类等。
github.com200
matthewspring/puck-gemini-flash-en-nano-codec-dataset
--- dataset_info: features: - name: text dtype: string - name: nano_layer_1 sequence: int64 - name: nano_layer_2 sequence: int64 - name: nano_layer_3 sequence: int64 - name
Hugging Face2025-12-04 更新140
MA-tokenweights/wikitext-103-raw-v2-tf-idf-wordlevel
--- dataset_info: features: - name: article_id dtype: int64 - name: source_split dtype: string - name: source_line_start dtype: int64 - name: source_line_end dtype: int64 -
Hugging Face2026-03-18 更新50
tasksource/lexcomp-nc-attributes
--- license: apache-2.0 language: - en --- https://github.com/vered1986/lexcomp/tree/master ``` @article{shwartz-dagan-2019-still, title = "Still a Pain in the Neck: Evaluating Text Representation
Hugging Face2023-06-02 更新190
Fraser/mnist-text-default
MNIST dataset adapted to a text-based representation. This allows testing interpolation quality for Transformer-VAEs. System is heavily inspired by Matthew Rayfield's work https://youtu.be/Z9K3cwSL6
Hugging Face2021-02-22 更新100



