F2STRANS Benchmark
收藏资源简介:
F2STRANS数据集是一个全新的代码翻译基准,它包含了最新的源代码、广泛的测试用例和手动注释的真实翻译。数据集涵盖了五种编程语言:C、C++、Go、Java和Python,每个问题选择两个代码解决方案作为源代码,并经过了广泛的测试用例的验证,确保了数据集的全面性和实用性。该数据集旨在帮助研究人员评估代码翻译模型的功能准确性和样式一致性,为代码翻译领域的研究提供了宝贵的资源。
The F2STRANS dataset is a novel code translation benchmark that contains up-to-date source code, extensive test cases, and manually annotated ground-truth translations. The dataset covers five programming languages: C, C++, Go, Java, and Python. For each problem, two code solutions are selected as the source code, which have been validated by extensive test cases to ensure the comprehensiveness and practicality of the dataset. This dataset aims to assist researchers in evaluating the functional accuracy and style consistency of code translation models, providing a valuable resource for research in the field of code translation.

- 1Function-to-Style Guidance of LLMs for Code Translation哈尔滨工业大学(深圳), 华为翻译服务中心(北京), 浙江大学(杭州) · 2025年



