ChatQA2Seg
收藏资源简介:
# Dataset Card for ChatQA2Seg <!-- Provide a quick summary of the dataset. --> The segmented version of [a filtered version of ChatQA2](https://huggingface.co/datasets/Seerkfang/ChatQA2-Long-SFT-data-long_sft_train_filtered). We use the [segmenter](https://modelscope.cn/models/SyonLi/Qwen3-4B-Instruct-2507-Segmenter) with recursion depth 2 and threshold value 0.2 and 0.4 for each level of recursion respectively. This dataset has been used as the training dataset in the paper [Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation](https://arxiv.org/abs/2605.15913). ## Dataset Details ### Dataset Description The fields in the cut_item: - txt_marker: The text string with inserted candidate cut points. - chunk_id: The segmentation boundaries for each chunk. - chunk_plain_text: The text content of each chunk. - cut_prob: The segmentation probabilities given by the segmenter. - threshold: The employed threshold for segmentation. If you find this dataset useful, please cite: ``` @article{li2026towards, title={Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation}, author={Li, Shuaiyi and Zhang, Zhisong and Wang, Yan and Zhu, Lei and Ma, Dongyang and Deng, Chenlong and Deng, Yang and Lam, Wai}, journal={arXiv preprint arXiv:2605.15913}, year={2026} } ```



