QuixiAI/fp_pt
收藏资源简介:
该数据集包含从本地JSONL文件转换而来的文本记录。每个源记录都有一个文本字段,在数据集中作为`text`列公开。数据转换过程中进行了清理:解码了HTML字符引用(如`'`、`&`、`<`和`"`);通过保守的文本清理移除了转录阶段标签、常见的YouTube样板文本、控制/替换字符、PDF百分比空格伪影以及原始LaTeX `tabular`包装器;并对常见提取伪影(如`A.Connella`、`ISSN:2153`和`text.Next`)进行了标点间距规范化。
This dataset consists of text records converted from local JSONL files. Each source record includes a text field, which is exposed as the `text` column within the dataset. Cleanup was conducted throughout the data conversion workflow: HTML character references (e.g., `'`, `&`, `<`, and `"`) were decoded; transcription-stage tags, common YouTube boilerplate text, control/replacement characters, PDF percentage-space artifacts, and original LaTeX `tabular` wrappers were removed via conservative text cleaning; and punctuation spacing was standardized for common extraction artifacts such as `A.Connella`, `ISSN:2153`, and `text.Next`.




