CircularBalls/plaincode-cnl-100k
收藏资源简介:
该数据集由Plaincode生成,通过流式处理Python源文件,将每个接受的Python模块/函数投影到当前的受控自然语言表面,并对每种语言强制执行严格的Python往返证明。数据集包含Python源文件以及六种语言的受控自然语言表面:英语、西班牙语、法语、葡萄牙语、普通话(中文)和印地语。每个行只有在所有请求的CNL表面通过Plaincode证明精确反向转换为原始Python字节时才会被接受,确保数据的准确性和一致性。数据集模式包括行ID、Python源代码信息、各语言的CNL文本、往返恢复的Python代码、证明字典、严格语言精确性检查以及语义家族标签等字段。生成统计显示有86,550个接受行,分布在多个数据分片中。
This dataset is generated by Plaincode, which stream-processes Python source files, projects each accepted Python module/function onto current controlled natural language (CNL) surface forms, and enforces strict Python round-trip verification for each language. The dataset contains Python source files and CNL surface forms in six languages: English, Spanish, French, Portuguese, Mandarin (Chinese), and Hindi. Each row is accepted only when all requested CNL surface forms are proven by Plaincode to be accurately reversibly converted back to the original Python bytes, ensuring data accuracy and consistency. The dataset schema includes fields such as row ID, Python source code information, CNL texts in each language, round-trip recovered Python code, verification dictionaries, strict language accuracy checks, and semantic family tags. Generation statistics show that there are 86,550 accepted rows distributed across multiple data shards.




