3gpp-tsg
收藏资源简介:
# 3GPP Working Group Classification (Benchmark Dataset) ## Dataset Summary This dataset benchmarks a model’s ability to **classify excerpts from 3GPP technical documents** into the correct **3GPP Working Group (WG)**. Each example is presented in an **instruction-following** format, where the model must output a JSON object containing exactly one label from a predefined WG list. The goal is to evaluate **telecom-domain understanding** (beyond keyword matching) by requiring the model to infer which WG (e.g., RAN*, SA*, CT*) would own the referenced technical content. ## Supported Tasks - **Text Classification**: Assign one label among 3GPP working groups (multi-class classification). - **Instruction Following**: Constrained-format JSON output. ## Labels (Working Groups) The dataset uses the following label set: `CT1, CT3, CT4, CT6, RAN1, RAN2, RAN3, RAN4, RAN5, RAN_AH1, SA1, SA2, SA3, SA4, SA5, SA6` ## Data Format Each record includes: - `instruction` (string): Task instruction to the model. - `input` (string): Prompt containing the WG label set and the extracted 3GPP text to classify. - `output` (string): Gold label in JSON format: `{"WORKING GROUP": "<WG>"}`. - `file_name` (string): Source identifier for the excerpt (e.g., 3GPP contribution file name). ### Example ```json { "instruction": "As a distinguished expert in telecommunication domain ...", "input": "Classify the following text ... You MUST select ONE working group name from this list: {...}\n\n###TEXT:\n{...}\n", "output": "{\"WORKING GROUP\": \"SA4\"}", "file_name": "S4-230942.txt" } ``` ## Intended Uses - **Benchmarking telecom-aware LLMs** (e.g., TelecomGPT-style models). - **Comparing general-purpose LLMs vs telecom-adapted models** to measure domain specialization gains. - **Testing RAG grounding quality on standards excerpts**, using classification accuracy as a sanity check for domain understanding. - **Instruction-tuning for structured classification outputs**, particularly constrained JSON responses. ## How to Use ### Load Locally (Raw JSON) ```python import json with open("3gpp_class_eval.json", "r", encoding="utf-8") as f: data = json.load(f) print(len(data)) print(data[0].keys()) ``` Here it is properly formatted for your dataset card — clean, copy-ready: ### Load as a Hugging Face Dataset If you upload the file as `3gpp_class_eval.json` in your dataset repository: ```python from datasets import load_dataset ds = load_dataset("KU-DFI/3gpp-tsg", data_files="3gpp_class_eval.json") print(ds) print(ds["train"][0]) ``` ## Citation ```bibtex @dataset{tsg_telecomgpt, title = {Telecomgpt: A framework to build telecom-specific large language models}, author = {Hang Zou, Qiyang Zhao, Yu Tian, Lina Bariah, Faouzi Bader, Thierry Lestable, Merouane Debbah}, year = {2025}, publisher = {IEEE Transactions on Machine Learning in Communications and Networking}, url = {https://ieeexplore.ieee.org/abstract/document/11097898} }



