cpnlab/LSM-Tokenized-Full
收藏资源简介:
该数据集是大型频谱模型(LSM)——令牌化射频数据集,旨在将原始无线频谱测量数据转换为离散的令牌序列,以连接无线信号处理和大语言模型(LLMs)。数据集基于论文《大型频谱模型(LSMs):通过令牌化RF数据的仅解码器Transformer驱动的频谱活动预测》引入。它包含约21 TB原始RF数据,覆盖33个子GHz频段(54 MHz至990 MHz),约84亿个令牌,词汇量为128个令牌,存储为1字节整数以实现快速I/O和低内存占用。每个样本为固定长度令牌序列(292个令牌),包括元数据令牌(36个)和频谱令牌(256个),用于频谱预测任务:输入前128个频谱令牌,预测后128个令牌。数据生成流程从原始IQ数据开始,经过STFT(频谱图)、最大池化(时间)、修剪均值(稳健平滑)和256×256频谱图,最终令牌化。数据集具有高时间分辨率(约3.91毫秒)和频率分辨率(78.125 kHz),适用于长序列建模,并包含蜂窝活动、广播信号和环境噪声模式等多样频谱环境。它兼容GPT风格、LLaMA、Mistral等基于Transformer的模型,以及LSTM、RNN和生成方法如扩散模型。数据集设计用于AI驱动的频谱智能研究,如动态频谱接入(DSA)和6G无线系统。
This dataset is a tokenized radio frequency (RF) dataset tailored for Large Spectrum Models (LSMs), which aims to convert raw wireless spectrum measurement data into discrete token sequences to bridge wireless signal processing and Large Language Models (LLMs). This dataset was introduced in the paper titled "Large Spectrum Models (LSMs): Decoder-only Transformer-driven Spectrum Activity Prediction via Tokenized RF Data". It contains approximately 21 TB of raw RF data, covering 33 sub-GHz frequency bands ranging from 54 MHz to 990 MHz, with around 8.4 billion tokens, a vocabulary size of 128 tokens, and is stored as 1-byte integers to enable fast I/O and low memory footprint. Each sample is a fixed-length token sequence of 292 tokens, comprising 36 metadata tokens and 256 spectrum tokens, designed for spectrum prediction tasks: the task takes the first 128 spectrum tokens as input and aims to predict the latter 128 tokens. The data generation pipeline starts from raw IQ data, undergoes Short-Time Fourier Transform (STFT, spectrogram generation), temporal max pooling, trimmed mean processing for robust smoothing, and 256×256 spectrogram formatting, before final tokenization. The dataset features high temporal resolution (approximately 3.91 ms) and frequency resolution (78.125 kHz), making it suitable for long-sequence modeling, and includes diverse spectrum environments such as cellular activity, broadcast signals, and ambient noise patterns. It is compatible with Transformer-based models like GPT-style architectures, LLaMA, and Mistral, as well as LSTM, RNN, and generative methods such as diffusion models. This dataset is designed for AI-driven spectrum intelligence research, such as dynamic spectrum access (DSA) and 6G wireless systems.




