遇见数据集

buddhist-nlp/pali-english

收藏
Hugging Face2023-05-07 更新2024-03-04 收录
官方服务:

资源简介:

--- dataset_info: features: - name: input_text dtype: string - name: target_text dtype: string - name: file_name dtype: string splits: - name: train num_bytes: 34632454.0 num_examples: 132151 - name: validation num_bytes: 2063756.0 num_examples: 7832 - name: test num_bytes: 2049351.0 num_examples: 7832 - name: test_500 num_bytes: 124892.0 num_examples: 499 - name: validation_500 num_bytes: 132892.0 num_examples: 499 download_size: 21840989 dataset_size: 39003345.0 --- # Dataset Card for "pali-english" [More Information needed](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)

提供机构:
buddhist-nlp
原始信息汇总

数据集概述

数据集特征

  • input_text:数据类型为字符串。
  • target_text:数据类型为字符串。
  • file_name:数据类型为字符串。

数据集分割

  • train:包含132,151个样本,总大小为34,632,454字节。
  • validation:包含7,832个样本,总大小为2,063,756字节。
  • test:包含7,832个样本,总大小为2,049,351字节。
  • test_500:包含499个样本,总大小为124,892字节。
  • validation_500:包含499个样本,总大小为132,892字节。

数据集大小

  • 下载大小:21,840,989字节。
  • 数据集总大小:39,003,345.0字节。
二维码
社区交流群
二维码
科研交流群
商业服务