遇见数据集

Promset: An annoted dataset for translating natural language to PromQl

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

PromSet is an annotated dataset designed to support natural language processing (NLP) research for system monitoring. It is particularly suited to applications involving the training and evaluation of large language models to translate queries expressed in natural language into their equivalent in PromQL, the query language used by the Prometheus monitoring tool. An initial dataset was constructed from the results of our experiments on Prometheus, during which we created a set of queries and their natural language descriptions. We then added additional data by collecting PromQL queries and their descriptions from various web sources. This raw data was curated, reviewed, corrected, and enriched with Gemini, resulting in a high-quality dataset suitable for research and development. The dataset contains a total of 4,350 manually curated pairs, each linking an English description to a corresponding PromQL expression. It is provided in CSV format, with two fields: description (a human-readable query) and promql (its equivalent in PromQL syntax). Each record represents a concrete and practical monitoring scenario, such as metric aggregation, label filtering, or time-based calculations. In many cases, a single PromQL query is associated with multiple English-language descriptions, increasing linguistic variation and enabling more robust model training. By bridging the gap between human-readable instructions and machine-interpretable PromQL syntax, Promset enables the development of intelligent systems capable of automatically understanding and generating monitoring queries. This facilitates the creation of more intuitive observability tools, streamlines DevOps workflows, and opens new avenues in research on natural language-to-code translation.

PromSet是一款带注释的数据集,旨在支撑面向系统监控的自然语言处理(NLP)研究,尤其适配于训练与评估大语言模型(LLM)的相关应用场景,可实现将自然语言表述的查询转换为普罗米修斯(Prometheus)监控工具所采用的查询语言PromQL的等价形式。 初始数据集源自我们在普罗米修斯上开展的实验结果,实验期间我们构建了一组查询语句及其自然语言描述。随后我们通过从各类网络资源中采集PromQL查询语句及其描述信息补充了额外的数据,该原始数据经Gemini整理、审阅、修正并增强后,最终形成了适用于研究与开发的高质量数据集。 本数据集共计包含4350条经人工整理的语句对,每一条均将英文描述与对应的PromQL表达式相关联。数据集以CSV格式提供,包含两个字段:"description"(人类可读的查询语句)与"promql"(其对应的PromQL语法等价形式)。每条记录均对应一个具体且实用的监控场景,例如指标聚合、标签过滤或基于时间的计算。在多数场景下,单条PromQL查询可对应多条英文描述,这提升了语言多样性,有助于训练出鲁棒性更强的模型。 通过弥合人类可读指令与机器可解析PromQL语法之间的鸿沟,PromSet助力开发能够自动理解并生成监控查询的智能系统。这有助于打造更具易用性的可观测性工具,优化DevOps工作流程,并为自然语言转代码翻译领域的研究开辟新方向。

创建时间:
2025-12-12
二维码
社区交流群
二维码
科研交流群
商业服务