amkyawdev/combined-myanmar-llm-dataset
收藏资源简介:
该数据集是一个多任务、多语言的大规模训练数据集,包含约302万个样本。每个样本包含对话历史(messages)、指令(instruction)、响应(response)、代码片段(code_snippets)、执行结果(execution_result)、错误日志(error_log)等字段。此外,还标注了任务类型(task_type)、使用的框架(framework)、运行时环境(runtime)、数据库(database)、工具(tools_used)、语言(language)、难度(difficulty)、项目规模(project_size)、复杂度评分(complexity_score)、用例(use_case)等元信息。该数据集适用于训练模型执行指令遵循、代码生成、工具调用、多轮对话、环境配置及复杂任务推理等场景。
This dataset is a large-scale multi-task and multi-language training dataset with approximately 3.02 million examples. Each example includes fields such as conversation history (messages), instruction, response, code snippets, execution results, error logs, and more. It also provides metadata including task type, framework, runtime environment, database, tools used, language, difficulty, project size, complexity score, use case, etc. The dataset is suitable for training models for instruction following, code generation, tool calling, multi-turn dialogue, environment setup, and complex task reasoning.



