相关数据集
formermagic/github_python_1m
--- annotations_creators: - found language_creators: - found language: - py license: - mit multilinguality: - monolingual size_categories: - 100K<n<1M source_datasets: - original task_categories: - se
Hugging Face2022-10-21 更新100
joaosanches/cleaned_tedtalks_total
--- dataset_info: features: - name: pt dtype: string - name: pt-br dtype: string splits: - name: train num_bytes: 66045930 num_examples: 314702 download_size: 41381774 da
Hugging Face2024-01-30 更新80
mathews5546/mathewzvk-starter-llama2-1k
--- dataset_info: features: - name: text dtype: string splits: - name: train num_bytes: 1654448 num_examples: 1000 download_size: 966692 dataset_size: 1654448 configs: - config
Hugging Face2024-03-13 更新40
LIACC/Emakhuwa-loanwords-detection
--- license: mit language: - pt - vmw task_categories: - text-classification --- # Detecting Loanwords in Emakhuwa Paper: **Detecting Loanwords in Emakhuwa: An Extremely Low-Resource {B}antu Languag
Hugging Face2024-08-05 更新80
nafisneehal/trialbrain_baseline_measures_instruct_v1
该数据集包含62,270个训练样本,每个样本包含instruction、input和output三个字符串类型的特征。数据集总大小为211,029,841字节,下载大小为96,199,241字节。数据文件路径为data/train-*。
Hugging Face2024-12-19 更新60



