CPM-Bench V1: Cost-Per-Meaning Benchmark for Multilingual Token Efficiency in Large Language Models
收藏资源简介:
CPM-Bench V1 is a multilingual Cost-Per-Meaning benchmark for measuring token efficiency across languages and tokenizers in large language model workflows. This dataset release contains 3,000 controlled English prompt-output pairs across 40 practical LLM task categories, translated into 20 target languages. The benchmark includes 60,000 prompt-language rows and 120,000 tokenizer-expanded analysis rows. Prompts are translated and back-translated with NLLB-200 distilled 1.3B, and semantic preservation is checked using LaBSE sentence embeddings. The dataset reports input, output, and total Token Efficiency Ratio (TER) relative to English for the Qwen3 14B Base tokenizer and the NLLB tokenizer. Under Qwen3, mean total TER ranges from 0.94x for Chinese to 8.64x for Punjabi, while the NLLB tokenizer ranges from about 0.93x to 1.57x. The uploaded file, cpm_bench_v1_data_assets.zip, contains the canonical prompt-language dataset, tokenizer-expanded master file, translation and back-translation files, translated reference outputs, LaBSE similarity scores, language/category summaries, human review sheets, generated charts, status metadata, and output README. Important files inside the bundle include 00_DATASET_MAIN_NO_TOKENIZER_DUPLICATES.csv as the canonical prompt-language dataset, 00_MASTER_FOR_REVIEW.csv as the tokenizer-expanded analysis file, 05B_overall_language_summary.csv as the language-level summary, and 05_summary_by_language_category_tokenizer.csv as the category-language-tokenizer summary. CPM-Bench V1 uses controlled reference outputs to isolate tokenizer-level cost from natural model-generation variability.



