ardauzunoglu/smollm2_v17grpo_c4_lowq_200m2b_subsample20m_grpo_prompt
收藏资源简介:
--- dataset_info: features: - name: row_idx dtype: int64 - name: sample_idx dtype: int64 - name: text dtype: string - name: finish_reason dtype: string - name: stop_reason dtype: string - name: document dtype: string - name: prompt_token_count dtype: int64 splits: - name: train num_bytes: 352901282 num_examples: 100018 download_size: 213567221 dataset_size: 352901282 configs: - config_name: default data_files: - split: train path: data/train-* ---
This dataset is a text processing dataset containing 100,018 training samples with a total size of approximately 352.9 MB. It features row index (row_idx), sample index (sample_idx), text content (text), finish reason (finish_reason), stop reason (stop_reason), document identifier (document), and prompt token count (prompt_token_count), suitable for text generation, analysis, or language model-related tasks.




