thinkPy/americasnlp_2025_guarani_analysis
收藏资源简介:
该数据集包含用于文本转换或编辑任务的数据,主要字段包括ID、源文本(Source)、变更描述(Change)和目标文本(Target)。数据集针对多个大语言模型(如SmolLM3-3B、Llama系列、Qwen系列等)提供了分词后的源文本和目标文本(Tokenized_Source和Tokenized_Target),以及对应的词元计数(Source_Token_Count和Target_Token_Count)、生育度量(Source_Fertility_Metric和Target_Fertility_Metric)、词元距离(Token_Distance)和归一化词元距离(Token_Distance_Normalized)。此外,还包含源文本和目标文本的形态学词元(Morphological_Tokens_Source和Morphological_Tokens_Target)。数据集分为训练集(178个示例)、开发集(79个示例)和测试集(364个示例),适用于自然语言处理任务,如文本改写、编辑距离分析或多模型分词比较。
This dataset contains data for text transformation or editing tasks, with main fields including ID, source text (Source), change description (Change), and target text (Target). The dataset provides tokenized source and target texts (Tokenized_Source and Tokenized_Target) for multiple large language models (e.g., SmolLM3-3B, Llama series, Qwen series), along with corresponding token counts (Source_Token_Count and Target_Token_Count), fertility metrics (Source_Fertility_Metric and Target_Fertility_Metric), token distance (Token_Distance), and normalized token distance (Token_Distance_Normalized). Additionally, it includes morphological tokens for source and target texts (Morphological_Tokens_Source and Morphological_Tokens_Target). The dataset is divided into training set (178 examples), development set (79 examples), and test set (364 examples), suitable for natural language processing tasks such as text paraphrasing, edit distance analysis, or multi-model tokenization comparison.




