phongnt251199/vietnam_heritage_wiki_chunks_v2
收藏资源简介:
--- dataset_info: features: - name: combined_text dtype: string - name: original_content dtype: string - name: metadata struct: - name: group dtype: string - name: level dtype: int64 - name: path dtype: string - name: section dtype: string - name: topic dtype: string - name: group dtype: string - name: image_urls list: string - name: best_image_url dtype: string - name: match_score dtype: float64 splits: - name: train num_bytes: 178398756 num_examples: 64124 download_size: 102351503 dataset_size: 178398756 configs: - config_name: default data_files: - split: train path: data/train-* ---
This dataset contains 64,124 training examples, each with features including: combined_text (combined text string), original_content (original content string), metadata (a structure with group, level, path, section, and topic fields), group (group string), image_urls (list of image URLs), best_image_url (best image URL string), and match_score (matching score float). The dataset may be used for text analysis, multimodal learning, or content matching tasks, but its specific purpose is not explicitly stated in the README.




