dianavdavidson/IndicVoices-Hinglish-two
收藏资源简介:
--- dataset_info: features: - name: audio dtype: audio - name: text dtype: string - name: duration dtype: float64 - name: lang dtype: string - name: samples dtype: int64 - name: verbatim dtype: string - name: normalized dtype: string - name: speaker_id dtype: string - name: scenario dtype: string - name: task_name dtype: string - name: gender dtype: string - name: age_group dtype: string - name: job_type dtype: string - name: qualification dtype: string - name: area dtype: string - name: district dtype: string - name: state dtype: string - name: occupation dtype: string - name: verification_report dtype: string - name: unsanitized_verbatim dtype: string - name: unsanitized_normalized dtype: string - name: unsanitized_no_noise_inds dtype: string - name: unsanitized_no_latin dtype: string - name: hinglish_mixed_scripts dtype: string - name: hinglish_mixed_script_lowercase dtype: string - name: english_words dtype: string - name: count_english_words dtype: int64 - name: count_dev_words_with_dupes dtype: int64 - name: count_dev_words_no_dupes dtype: int64 - name: ratio_hindi_words dtype: float64 - name: ratio_english_words dtype: float64 - name: ratio_english_words_range dtype: string - name: hindi_words dtype: string - name: unique_hindi_words sequence: string - name: unique_hindi_words_count dtype: int64 - name: unique_english_words sequence: string - name: unique_english_words_count dtype: int64 splits: - name: train num_bytes: 38741373506.008 num_examples: 371224 - name: test num_bytes: 409109200.326 num_examples: 4606 download_size: 38376573722 dataset_size: 39150482706.334 configs: - config_name: default data_files: - split: train path: data/train-* - split: test path: data/test-* ---



