遇见数据集

T-REx Bite 1.0

收藏
Zenodo2025-04-07 更新2026-05-26 收录
官方服务:

资源简介:

T-REx Bite adapts T-REx to the next-token-prediction paradigm by ensuring that, in each text snippet, the subject appears before the object. This alignment mimics real-world scenarios in which a model sees partial information (the subject and some context) and must then predict or "retrieve" the missing object. To keep snippets manageable within the limited context windows of smaller language models (such as GPT-2), each snippet is capped at 512 characters. We further require that the snippet explicitly mentions both subject and object, does not start a new sentence at the object, and is linked to the corresponding subgraph in T-REx Star. By applying these constraints, we obtain about 6.4 million short "bites" for training, 0.92 million for testing, and 0.75 million for validation. Each bite is a compact piece of Wikipedia text that retains the original clarity and diversity of T-REx but is tailored to ensure a direct subject--object alignment. This structure lets researchers readily evaluate how well an LLM completes the object token(s) given the preceding subject. The dataset naturally accommodates multi-token objects under modern sub-word tokenization, removing the single-token assumptions of LAMA-like methods.

提供机构:
Zenodo
创建时间:
2025-04-07
二维码
社区交流群
二维码
科研交流群
商业服务