h-gajdov/vezilka-crowd-sourcing-20260527-023047
收藏资源简介:
--- dataset_info: features: - name: id dtype: string - name: text dtype: string - name: source dtype: string - name: chunk dtype: int64 - name: topic dtype: string - name: description dtype: string - name: file_type dtype: string - name: dialect dtype: 'null' - name: page dtype: int64 splits: - name: train num_bytes: 431948 num_examples: 367 download_size: 165405 dataset_size: 431948 configs: - config_name: default data_files: - split: train path: data/train-* ---
This dataset contains 367 training examples with a total size of 431948 bytes and a download size of 165405 bytes. Features include id (string type), text (text content, string type), source (origin, string type), chunk (chunk number, integer type), topic (topic, string type), description (description, string type), file_type (file type, string type), dialect (dialect, null type), and page (page number, integer type). The data is only provided in a train split and is suitable for text analysis or natural language processing tasks, but specific purposes and background are not explicitly described in the README.




