MaFLLM
收藏资源简介:
MaFLLM is dataset designed for fingerprinting and identifying underlying Large Language Models (LLMs) based solely on encrypted network traffic behavior. Collected from API-driven conversations between an agent client and a local LLM inference service running on port, the captures record fine-grained TCP/UDP side-channel metrics—including packet timestamps, packet size distributions, IP metadata, protocols. The dataset encompasses 10 distinct LLM model families (ranging from lightweight models like llama3.2-4b and phi-4-mini-reasoning to larger architectures like qwen2.5-7b and qwen3-14b), categorized across 6 prompt categories (C1–C6) representing different task domains, with over 10 unique prompts per category and 100 individual sample runs per prompt (S001–S100). Each capture file adheres to the systematic naming scheme M{model}-C{category}-P{prompt}-S{sample}.pcap, offering a structured bench for research in encrypted traffic analysis, privacy attacks, and LLM inference classification.



