OpenResearcher/OpenResearcher-Dataset
收藏资源简介:
--- dataset_info: - config_name: seed_42 features: - name: qid dtype: int64 - name: question dtype: string - name: answer dtype: string - name: messages list: - name: channel dtype: string - name: content list: - name: channel_config struct: - name: channel_required dtype: bool - name: valid_channels list: string - name: conversation_start_date dtype: string - name: knowledge_cutoff dtype: string - name: model_identity dtype: string - name: reasoning_effort dtype: string - name: text dtype: string - name: tools struct: - name: browser struct: - name: description dtype: string - name: name dtype: string - name: tools list: - name: description dtype: string - name: name dtype: string - name: parameters struct: - name: properties struct: - name: cursor struct: - name: default dtype: int64 - name: type dtype: string - name: id struct: - name: default dtype: int64 - name: type list: string - name: loc struct: - name: default dtype: int64 - name: type dtype: string - name: num_lines struct: - name: default dtype: int64 - name: type dtype: string - name: pattern struct: - name: type dtype: string - name: query struct: - name: type dtype: string - name: source struct: - name: type dtype: string - name: topn struct: - name: default dtype: int64 - name: type dtype: string - name: view_source struct: - name: default dtype: bool - name: type dtype: string - name: required list: string - name: type dtype: string - name: type dtype: string - name: content_type dtype: string - name: name dtype: string - name: recipient dtype: string - name: role dtype: string - name: latency_s dtype: float64 - name: error dtype: 'null' - name: attempts dtype: int64 - name: status dtype: string - name: chunk_idx dtype: int64 - name: num_chunks dtype: int64 splits: - name: train num_bytes: 1117982919 num_examples: 6102 download_size: 461369938 dataset_size: 1117982919 - config_name: seed_43 features: - name: qid dtype: int64 - name: question dtype: string - name: answer dtype: string - name: messages list: - name: channel dtype: string - name: content list: - name: channel_config struct: - name: channel_required dtype: bool - name: valid_channels list: string - name: conversation_start_date dtype: string - name: knowledge_cutoff dtype: string - name: model_identity dtype: string - name: reasoning_effort dtype: string - name: text dtype: string - name: tools struct: - name: browser struct: - name: description dtype: string - name: name dtype: string - name: tools list: - name: description dtype: string - name: name dtype: string - name: parameters struct: - name: properties struct: - name: cursor struct: - name: default dtype: int64 - name: type dtype: string - name: id struct: - name: default dtype: int64 - name: type list: string - name: loc struct: - name: default dtype: int64 - name: type dtype: string - name: num_lines struct: - name: default dtype: int64 - name: type dtype: string - name: pattern struct: - name: type dtype: string - name: query struct: - name: type dtype: string - name: source struct: - name: type dtype: string - name: topn struct: - name: default dtype: int64 - name: type dtype: string - name: view_source struct: - name: default dtype: bool - name: type dtype: string - name: required list: string - name: type dtype: string - name: type dtype: string - name: content_type dtype: string - name: name dtype: string - name: recipient dtype: string - name: role dtype: string - name: latency_s dtype: float64 - name: error dtype: 'null' - name: attempts dtype: int64 - name: status dtype: string - name: chunk_idx dtype: int64 - name: num_chunks dtype: int64 splits: - name: train num_bytes: 1158931340 num_examples: 6102 download_size: 477917292 dataset_size: 1158931340 - config_name: seed_44 features: - name: qid dtype: int64 - name: question dtype: string - name: answer dtype: string - name: messages list: - name: channel dtype: string - name: content list: - name: channel_config struct: - name: channel_required dtype: bool - name: valid_channels list: string - name: conversation_start_date dtype: string - name: knowledge_cutoff dtype: string - name: model_identity dtype: string - name: reasoning_effort dtype: string - name: text dtype: string - name: tools struct: - name: browser struct: - name: description dtype: string - name: name dtype: string - name: tools list: - name: description dtype: string - name: name dtype: string - name: parameters struct: - name: properties struct: - name: cursor struct: - name: default dtype: int64 - name: type dtype: string - name: id struct: - name: default dtype: int64 - name: type list: string - name: loc struct: - name: default dtype: int64 - name: type dtype: string - name: num_lines struct: - name: default dtype: int64 - name: type dtype: string - name: pattern struct: - name: type dtype: string - name: query struct: - name: type dtype: string - name: source struct: - name: type dtype: string - name: topn struct: - name: default dtype: int64 - name: type dtype: string - name: view_source struct: - name: default dtype: bool - name: type dtype: string - name: required list: string - name: type dtype: string - name: type dtype: string - name: content_type dtype: string - name: name dtype: string - name: recipient dtype: string - name: role dtype: string - name: latency_s dtype: float64 - name: error dtype: string - name: attempts dtype: int64 - name: status dtype: string - name: chunk_idx dtype: int64 - name: num_chunks dtype: int64 splits: - name: train num_bytes: 1108145493 num_examples: 6102 download_size: 456027149 dataset_size: 1108145493 - config_name: seed_45 features: - name: qid dtype: int64 - name: question dtype: string - name: answer dtype: string - name: messages list: - name: channel dtype: string - name: content list: - name: channel_config struct: - name: channel_required dtype: bool - name: valid_channels list: string - name: conversation_start_date dtype: string - name: knowledge_cutoff dtype: string - name: model_identity dtype: string - name: reasoning_effort dtype: string - name: text dtype: string - name: tools struct: - name: browser struct: - name: description dtype: string - name: name dtype: string - name: tools list: - name: description dtype: string - name: name dtype: string - name: parameters struct: - name: properties struct: - name: cursor struct: - name: default dtype: int64 - name: type dtype: string - name: id struct: - name: default dtype: int64 - name: type list: string - name: loc struct: - name: default dtype: int64 - name: type dtype: string - name: num_lines struct: - name: default dtype: int64 - name: type dtype: string - name: pattern struct: - name: type dtype: string - name: query struct: - name: type dtype: string - name: source struct: - name: type dtype: string - name: topn struct: - name: default dtype: int64 - name: type dtype: string - name: view_source struct: - name: default dtype: bool - name: type dtype: string - name: required list: string - name: type dtype: string - name: type dtype: string - name: content_type dtype: string - name: name dtype: string - name: recipient dtype: string - name: role dtype: string - name: latency_s dtype: float64 - name: error dtype: 'null' - name: attempts dtype: int64 - name: status dtype: string - name: chunk_idx dtype: int64 - name: num_chunks dtype: int64 splits: - name: train num_bytes: 1106900749 num_examples: 6102 download_size: 455271833 dataset_size: 1106900749 - config_name: seed_46 features: - name: qid dtype: int64 - name: question dtype: string - name: answer dtype: string - name: messages list: - name: channel dtype: string - name: content list: - name: channel_config struct: - name: channel_required dtype: bool - name: valid_channels list: string - name: conversation_start_date dtype: string - name: knowledge_cutoff dtype: string - name: model_identity dtype: string - name: reasoning_effort dtype: string - name: text dtype: string - name: tools struct: - name: browser struct: - name: description dtype: string - name: name dtype: string - name: tools list: - name: description dtype: string - name: name dtype: string - name: parameters struct: - name: properties struct: - name: cursor struct: - name: default dtype: int64 - name: type dtype: string - name: id struct: - name: default dtype: int64 - name: type list: string - name: loc struct: - name: default dtype: int64 - name: type dtype: string - name: num_lines struct: - name: default dtype: int64 - name: type dtype: string - name: pattern struct: - name: type dtype: string - name: query struct: - name: type dtype: string - name: source struct: - name: type dtype: string - name: topn struct: - name: default dtype: int64 - name: type dtype: string - name: view_source struct: - name: default dtype: bool - name: type dtype: string - name: required list: string - name: type dtype: string - name: type dtype: string - name: content_type dtype: string - name: name dtype: string - name: recipient dtype: string - name: role dtype: string - name: latency_s dtype: float64 - name: error dtype: 'null' - name: attempts dtype: int64 - name: status dtype: string - name: chunk_idx dtype: int64 - name: num_chunks dtype: int64 splits: - name: train num_bytes: 1135517221 num_examples: 6102 download_size: 468875734 dataset_size: 1135517221 - config_name: seed_47 features: - name: qid dtype: int64 - name: question dtype: string - name: answer dtype: string - name: messages list: - name: channel dtype: string - name: content list: - name: channel_config struct: - name: channel_required dtype: bool - name: valid_channels list: string - name: conversation_start_date dtype: string - name: knowledge_cutoff dtype: string - name: model_identity dtype: string - name: reasoning_effort dtype: string - name: text dtype: string - name: tools struct: - name: browser struct: - name: description dtype: string - name: name dtype: string - name: tools list: - name: description dtype: string - name: name dtype: string - name: parameters struct: - name: properties struct: - name: cursor struct: - name: default dtype: int64 - name: type dtype: string - name: id struct: - name: default dtype: int64 - name: type list: string - name: loc struct: - name: default dtype: int64 - name: type dtype: string - name: num_lines struct: - name: default dtype: int64 - name: type dtype: string - name: pattern struct: - name: type dtype: string - name: query struct: - name: type dtype: string - name: source struct: - name: type dtype: string - name: topn struct: - name: default dtype: int64 - name: type dtype: string - name: view_source struct: - name: default dtype: bool - name: type dtype: string - name: required list: string - name: type dtype: string - name: type dtype: string - name: content_type dtype: string - name: name dtype: string - name: recipient dtype: string - name: role dtype: string - name: latency_s dtype: float64 - name: error dtype: 'null' - name: attempts dtype: int64 - name: status dtype: string - name: chunk_idx dtype: int64 - name: num_chunks dtype: int64 splits: - name: train num_bytes: 1179953502 num_examples: 6102 download_size: 487937315 dataset_size: 1179953502 - config_name: seed_48 features: - name: qid dtype: int64 - name: question dtype: string - name: answer dtype: string - name: messages list: - name: channel dtype: string - name: content list: - name: channel_config struct: - name: channel_required dtype: bool - name: valid_channels list: string - name: conversation_start_date dtype: string - name: knowledge_cutoff dtype: string - name: model_identity dtype: string - name: reasoning_effort dtype: string - name: text dtype: string - name: tools struct: - name: browser struct: - name: description dtype: string - name: name dtype: string - name: tools list: - name: description dtype: string - name: name dtype: string - name: parameters struct: - name: properties struct: - name: cursor struct: - name: default dtype: int64 - name: type dtype: string - name: id struct: - name: default dtype: int64 - name: type list: string - name: loc struct: - name: default dtype: int64 - name: type dtype: string - name: num_lines struct: - name: default dtype: int64 - name: type dtype: string - name: pattern struct: - name: type dtype: string - name: query struct: - name: type dtype: string - name: source struct: - name: type dtype: string - name: topn struct: - name: default dtype: int64 - name: type dtype: string - name: view_source struct: - name: default dtype: bool - name: type dtype: string - name: required list: string - name: type dtype: string - name: type dtype: string - name: content_type dtype: string - name: name dtype: string - name: recipient dtype: string - name: role dtype: string - name: latency_s dtype: float64 - name: error dtype: 'null' - name: attempts dtype: int64 - name: status dtype: string - name: chunk_idx dtype: int64 - name: num_chunks dtype: int64 splits: - name: train num_bytes: 1171147444 num_examples: 6102 download_size: 483010306 dataset_size: 1171147444 - config_name: seed_49 features: - name: qid dtype: int64 - name: question dtype: string - name: answer dtype: string - name: messages list: - name: channel dtype: string - name: content list: - name: channel_config struct: - name: channel_required dtype: bool - name: valid_channels list: string - name: conversation_start_date dtype: string - name: knowledge_cutoff dtype: string - name: model_identity dtype: string - name: reasoning_effort dtype: string - name: text dtype: string - name: tools struct: - name: browser struct: - name: description dtype: string - name: name dtype: string - name: tools list: - name: description dtype: string - name: name dtype: string - name: parameters struct: - name: properties struct: - name: cursor struct: - name: default dtype: int64 - name: type dtype: string - name: id struct: - name: default dtype: int64 - name: type list: string - name: loc struct: - name: default dtype: int64 - name: type dtype: string - name: num_lines struct: - name: default dtype: int64 - name: type dtype: string - name: pattern struct: - name: type dtype: string - name: query struct: - name: type dtype: string - name: source struct: - name: type dtype: string - name: topn struct: - name: default dtype: int64 - name: type dtype: string - name: view_source struct: - name: default dtype: bool - name: type dtype: string - name: required list: string - name: type dtype: string - name: type dtype: string - name: content_type dtype: string - name: name dtype: string - name: recipient dtype: string - name: role dtype: string - name: latency_s dtype: float64 - name: error dtype: 'null' - name: attempts dtype: int64 - name: status dtype: string - name: chunk_idx dtype: int64 - name: num_chunks dtype: int64 splits: - name: train num_bytes: 1120741628 num_examples: 6102 download_size: 461516938 dataset_size: 1120741628 - config_name: seed_50 features: - name: qid dtype: int64 - name: question dtype: string - name: answer dtype: string - name: messages list: - name: channel dtype: string - name: content list: - name: channel_config struct: - name: channel_required dtype: bool - name: valid_channels list: string - name: conversation_start_date dtype: string - name: knowledge_cutoff dtype: string - name: model_identity dtype: string - name: reasoning_effort dtype: string - name: text dtype: string - name: tools struct: - name: browser struct: - name: description dtype: string - name: name dtype: string - name: tools list: - name: description dtype: string - name: name dtype: string - name: parameters struct: - name: properties struct: - name: cursor struct: - name: default dtype: int64 - name: type dtype: string - name: id struct: - name: default dtype: int64 - name: type list: string - name: loc struct: - name: default dtype: int64 - name: type dtype: string - name: num_lines struct: - name: default dtype: int64 - name: type dtype: string - name: pattern struct: - name: type dtype: string - name: query struct: - name: type dtype: string - name: source struct: - name: type dtype: string - name: topn struct: - name: default dtype: int64 - name: type dtype: string - name: view_source struct: - name: default dtype: bool - name: type dtype: string - name: required list: string - name: type dtype: string - name: type dtype: string - name: content_type dtype: string - name: name dtype: string - name: recipient dtype: string - name: role dtype: string - name: latency_s dtype: float64 - name: error dtype: 'null' - name: attempts dtype: int64 - name: status dtype: string - name: chunk_idx dtype: int64 - name: num_chunks dtype: int64 splits: - name: train num_bytes: 1139271069 num_examples: 6102 download_size: 470031205 dataset_size: 1139271069 - config_name: seed_51 features: - name: qid dtype: int64 - name: question dtype: string - name: answer dtype: string - name: messages list: - name: channel dtype: string - name: content list: - name: channel_config struct: - name: channel_required dtype: bool - name: valid_channels list: string - name: conversation_start_date dtype: string - name: knowledge_cutoff dtype: string - name: model_identity dtype: string - name: reasoning_effort dtype: string - name: text dtype: string - name: tools struct: - name: browser struct: - name: description dtype: string - name: name dtype: string - name: tools list: - name: description dtype: string - name: name dtype: string - name: parameters struct: - name: properties struct: - name: cursor struct: - name: default dtype: int64 - name: type dtype: string - name: id struct: - name: default dtype: int64 - name: type list: string - name: loc struct: - name: default dtype: int64 - name: type dtype: string - name: num_lines struct: - name: default dtype: int64 - name: type dtype: string - name: pattern struct: - name: type dtype: string - name: query struct: - name: type dtype: string - name: source struct: - name: type dtype: string - name: topn struct: - name: default dtype: int64 - name: type dtype: string - name: view_source struct: - name: default dtype: bool - name: type dtype: string - name: required list: string - name: type dtype: string - name: type dtype: string - name: content_type dtype: string - name: name dtype: string - name: recipient dtype: string - name: role dtype: string - name: latency_s dtype: float64 - name: error dtype: 'null' - name: attempts dtype: int64 - name: status dtype: string - name: chunk_idx dtype: int64 - name: num_chunks dtype: int64 splits: - name: train num_bytes: 1154584409 num_examples: 6102 download_size: 475934015 dataset_size: 1154584409 - config_name: seed_52 features: - name: qid dtype: int64 - name: question dtype: string - name: answer dtype: string - name: messages list: - name: channel dtype: string - name: content list: - name: channel_config struct: - name: channel_required dtype: bool - name: valid_channels list: string - name: conversation_start_date dtype: string - name: knowledge_cutoff dtype: string - name: model_identity dtype: string - name: reasoning_effort dtype: string - name: text dtype: string - name: tools struct: - name: browser struct: - name: description dtype: string - name: name dtype: string - name: tools list: - name: description dtype: string - name: name dtype: string - name: parameters struct: - name: properties struct: - name: cursor struct: - name: default dtype: int64 - name: type dtype: string - name: id struct: - name: default dtype: int64 - name: type list: string - name: loc struct: - name: default dtype: int64 - name: type dtype: string - name: num_lines struct: - name: default dtype: int64 - name: type dtype: string - name: pattern struct: - name: type dtype: string - name: query struct: - name: type dtype: string - name: source struct: - name: type dtype: string - name: topn struct: - name: default dtype: int64 - name: type dtype: string - name: view_source struct: - name: default dtype: bool - name: type dtype: string - name: required list: string - name: type dtype: string - name: type dtype: string - name: content_type dtype: string - name: name dtype: string - name: recipient dtype: string - name: role dtype: string - name: latency_s dtype: float64 - name: error dtype: string - name: attempts dtype: int64 - name: status dtype: string - name: chunk_idx dtype: int64 - name: num_chunks dtype: int64 splits: - name: train num_bytes: 1133567180 num_examples: 6102 download_size: 467407950 dataset_size: 1133567180 - config_name: seed_53 features: - name: qid dtype: int64 - name: question dtype: string - name: answer dtype: string - name: messages list: - name: channel dtype: string - name: content list: - name: channel_config struct: - name: channel_required dtype: bool - name: valid_channels list: string - name: conversation_start_date dtype: string - name: knowledge_cutoff dtype: string - name: model_identity dtype: string - name: reasoning_effort dtype: string - name: text dtype: string - name: tools struct: - name: browser struct: - name: description dtype: string - name: name dtype: string - name: tools list: - name: description dtype: string - name: name dtype: string - name: parameters struct: - name: properties struct: - name: cursor struct: - name: default dtype: int64 - name: type dtype: string - name: id struct: - name: default dtype: int64 - name: type list: string - name: loc struct: - name: default dtype: int64 - name: type dtype: string - name: num_lines struct: - name: default dtype: int64 - name: type dtype: string - name: pattern struct: - name: type dtype: string - name: query struct: - name: type dtype: string - name: source struct: - name: type dtype: string - name: topn struct: - name: default dtype: int64 - name: type dtype: string - name: view_source struct: - name: default dtype: bool - name: type dtype: string - name: required list: string - name: type dtype: string - name: type dtype: string - name: content_type dtype: string - name: name dtype: string - name: recipient dtype: string - name: role dtype: string - name: latency_s dtype: float64 - name: error dtype: 'null' - name: attempts dtype: int64 - name: status dtype: string - name: chunk_idx dtype: int64 - name: num_chunks dtype: int64 splits: - name: train num_bytes: 1024473567 num_examples: 6102 download_size: 420595363 dataset_size: 1024473567 - config_name: seed_54 features: - name: qid dtype: int64 - name: question dtype: string - name: answer dtype: string - name: messages list: - name: channel dtype: string - name: content list: - name: channel_config struct: - name: channel_required dtype: bool - name: valid_channels list: string - name: conversation_start_date dtype: string - name: knowledge_cutoff dtype: string - name: model_identity dtype: string - name: reasoning_effort dtype: string - name: text dtype: string - name: tools struct: - name: browser struct: - name: description dtype: string - name: name dtype: string - name: tools list: - name: description dtype: string - name: name dtype: string - name: parameters struct: - name: properties struct: - name: cursor struct: - name: default dtype: int64 - name: type dtype: string - name: id struct: - name: default dtype: int64 - name: type list: string - name: loc struct: - name: default dtype: int64 - name: type dtype: string - name: num_lines struct: - name: default dtype: int64 - name: type dtype: string - name: pattern struct: - name: type dtype: string - name: query struct: - name: type dtype: string - name: source struct: - name: type dtype: string - name: topn struct: - name: default dtype: int64 - name: type dtype: string - name: view_source struct: - name: default dtype: bool - name: type dtype: string - name: required list: string - name: type dtype: string - name: type dtype: string - name: content_type dtype: string - name: name dtype: string - name: recipient dtype: string - name: role dtype: string - name: latency_s dtype: float64 - name: error dtype: 'null' - name: attempts dtype: int64 - name: status dtype: string - name: chunk_idx dtype: int64 - name: num_chunks dtype: int64 splits: - name: train num_bytes: 1122591364 num_examples: 6102 download_size: 463334159 dataset_size: 1122591364 - config_name: seed_55 features: - name: qid dtype: int64 - name: question dtype: string - name: answer dtype: string - name: messages list: - name: channel dtype: string - name: content list: - name: channel_config struct: - name: channel_required dtype: bool - name: valid_channels list: string - name: conversation_start_date dtype: string - name: knowledge_cutoff dtype: string - name: model_identity dtype: string - name: reasoning_effort dtype: string - name: text dtype: string - name: tools struct: - name: browser struct: - name: description dtype: string - name: name dtype: string - name: tools list: - name: description dtype: string - name: name dtype: string - name: parameters struct: - name: properties struct: - name: cursor struct: - name: default dtype: int64 - name: type dtype: string - name: id struct: - name: default dtype: int64 - name: type list: string - name: loc struct: - name: default dtype: int64 - name: type dtype: string - name: num_lines struct: - name: default dtype: int64 - name: type dtype: string - name: pattern struct: - name: type dtype: string - name: query struct: - name: type dtype: string - name: source struct: - name: type dtype: string - name: topn struct: - name: default dtype: int64 - name: type dtype: string - name: view_source struct: - name: default dtype: bool - name: type dtype: string - name: required list: string - name: type dtype: string - name: type dtype: string - name: content_type dtype: string - name: name dtype: string - name: recipient dtype: string - name: role dtype: string - name: latency_s dtype: float64 - name: error dtype: 'null' - name: attempts dtype: int64 - name: status dtype: string - name: chunk_idx dtype: int64 - name: num_chunks dtype: int64 splits: - name: train num_bytes: 1122857864 num_examples: 6100 download_size: 462379167 dataset_size: 1122857864 - config_name: seed_56 features: - name: qid dtype: int64 - name: question dtype: string - name: answer dtype: string - name: messages list: - name: channel dtype: string - name: content list: - name: channel_config struct: - name: channel_required dtype: bool - name: valid_channels list: string - name: conversation_start_date dtype: string - name: knowledge_cutoff dtype: string - name: model_identity dtype: string - name: reasoning_effort dtype: string - name: text dtype: string - name: tools struct: - name: browser struct: - name: description dtype: string - name: name dtype: string - name: tools list: - name: description dtype: string - name: name dtype: string - name: parameters struct: - name: properties struct: - name: cursor struct: - name: default dtype: int64 - name: type dtype: string - name: id struct: - name: default dtype: int64 - name: type list: string - name: loc struct: - name: default dtype: int64 - name: type dtype: string - name: num_lines struct: - name: default dtype: int64 - name: type dtype: string - name: pattern struct: - name: type dtype: string - name: query struct: - name: type dtype: string - name: source struct: - name: type dtype: string - name: topn struct: - name: default dtype: int64 - name: type dtype: string - name: view_source struct: - name: default dtype: bool - name: type dtype: string - name: required list: string - name: type dtype: string - name: type dtype: string - name: content_type dtype: string - name: name dtype: string - name: recipient dtype: string - name: role dtype: string - name: latency_s dtype: float64 - name: error dtype: 'null' - name: attempts dtype: int64 - name: status dtype: string - name: chunk_idx dtype: int64 - name: num_chunks dtype: int64 splits: - name: train num_bytes: 1150224058 num_examples: 6102 download_size: 474643458 dataset_size: 1150224058 - config_name: seed_57 features: - name: qid dtype: int64 - name: question dtype: string - name: answer dtype: string - name: messages list: - name: channel dtype: string - name: content list: - name: channel_config struct: - name: channel_required dtype: bool - name: valid_channels list: string - name: conversation_start_date dtype: string - name: knowledge_cutoff dtype: string - name: model_identity dtype: string - name: reasoning_effort dtype: string - name: text dtype: string - name: tools struct: - name: browser struct: - name: description dtype: string - name: name dtype: string - name: tools list: - name: description dtype: string - name: name dtype: string - name: parameters struct: - name: properties struct: - name: cursor struct: - name: default dtype: int64 - name: type dtype: string - name: id struct: - name: default dtype: int64 - name: type list: string - name: loc struct: - name: default dtype: int64 - name: type dtype: string - name: num_lines struct: - name: default dtype: int64 - name: type dtype: string - name: pattern struct: - name: type dtype: string - name: query struct: - name: type dtype: string - name: source struct: - name: type dtype: string - name: topn struct: - name: default dtype: int64 - name: type dtype: string - name: view_source struct: - name: default dtype: bool - name: type dtype: string - name: required list: string - name: type dtype: string - name: type dtype: string - name: content_type dtype: string - name: name dtype: string - name: recipient dtype: string - name: role dtype: string - name: latency_s dtype: float64 - name: error dtype: 'null' - name: attempts dtype: int64 - name: status dtype: string - name: chunk_idx dtype: int64 - name: num_chunks dtype: int64 splits: - name: train num_bytes: 1131677120 num_examples: 6102 download_size: 466987581 dataset_size: 1131677120 configs: - config_name: seed_42 data_files: - split: train path: seed_42/train-* - config_name: seed_43 data_files: - split: train path: seed_43/train-* - config_name: seed_44 data_files: - split: train path: seed_44/train-* - config_name: seed_45 data_files: - split: train path: seed_45/train-* - config_name: seed_46 data_files: - split: train path: seed_46/train-* - config_name: seed_47 data_files: - split: train path: seed_47/train-* - config_name: seed_48 data_files: - split: train path: seed_48/train-* - config_name: seed_49 data_files: - split: train path: seed_49/train-* - config_name: seed_50 data_files: - split: train path: seed_50/train-* - config_name: seed_51 data_files: - split: train path: seed_51/train-* - config_name: seed_52 data_files: - split: train path: seed_52/train-* - config_name: seed_53 data_files: - split: train path: seed_53/train-* - config_name: seed_54 data_files: - split: train path: seed_54/train-* - config_name: seed_55 data_files: - split: train path: seed_55/train-* - config_name: seed_56 data_files: - split: train path: seed_56/train-* - config_name: seed_57 data_files: - split: train path: seed_57/train-* license: mit --- <div style="display: flex; align-items: center; justify-content: center; gap: 8px;"> <img src="imgs/or-logo1.png" style="height: 84px; width: auto;"> <img src="imgs/openresearcher-title.svg" style="height: 84px; width: auto;"> </div> <div align="center"> <a href="https://arxiv.org/abs/2603.20278"><img src="https://img.shields.io/badge/arXiv-B31B1B?style=for-the-badge&logo=arXiv&logoColor=white" alt="Blog"></a> <a href="https://huggingface.co/papers/2603.20278"><img src="https://img.shields.io/badge/Paper-FFD966?style=for-the-badge&logo=huggingface&logoColor=ffffff" alt="Model"></a> <!-- <a href="https://huggingface.co/papers/2603.20278"><img src="https://img.shields.io/badge/arXiv-B31B1B?style=for-the-badge&logo=arXiv&logoColor=white" alt="Blog"></a> --> <a href="https://x.com/zhuofengli96475/status/2036475211063648414"><img src="https://img.shields.io/badge/Twitter-000000?style=for-the-badge&logo=X&logoColor=white" alt="Blog"></a> <!-- <a href="https://boiled-honeycup-4c7.notion.site/OpenResearcher-A-Fully-Open-Pipeline-for-Long-Horizon-Deep-Research-Trajectory-Synthesis-2f7e290627b5800cb3a0cd7e8d6ec0ea?source=copy_link"><img src="https://img.shields.io/badge/Blog-4285F4?style=for-the-badge&logo=google-chrome&logoColor=white" alt="Blog"></a> --> <a href="https://github.com/TIGER-AI-Lab/OpenResearcher"><img src="https://img.shields.io/badge/Github-181717?style=for-the-badge&logo=github&logoColor=white" alt="Blog"></a> <a href="https://huggingface.co/datasets/OpenResearcher/OpenResearcher-Dataset"><img src="https://img.shields.io/badge/Dataset-FFB7B2?style=for-the-badge&logo=huggingface&logoColor=ffffff" alt="Dataset"></a> <a href="https://huggingface.co/OpenResearcher/Nemotron-3-Nano-30B-A3B"><img src="https://img.shields.io/badge/Model-FFD966?style=for-the-badge&logo=huggingface&logoColor=ffffff" alt="Model"></a> <a href="https://huggingface.co/spaces/OpenResearcher/OpenResearcher"><img src="https://img.shields.io/badge/Demo-F97316.svg?style=for-the-badge&logo=gradio&logoColor=white" alt="Demo"></a> <!-- <a href="https://wandb.ai/dongfu/nano-v3-sft-search"><img src="https://img.shields.io/badge/WandB%20Logs-48B5A3?style=for-the-badge&logo=weightsandbiases&logoColor=white" alt="WandB Logs"></a> --> <a href="https://huggingface.co/datasets/OpenResearcher/OpenResearcher-Eval-Logs/tree/main"><img src="https://img.shields.io/badge/Eval%20Logs-755BB4?style=for-the-badge&logo=google-sheets&logoColor=white" alt="Eval Logs"></a> </div> </div> <div align="center" style="padding: 10px 0 -4px; display: flex; align-items: center; justify-content: center; gap: 16px;"> <div style="width: 60px; height: 2px; background: linear-gradient(90deg, transparent, #E24B4A);"></div> <span style="font-size: 22px; font-weight: 600; color: #E24B4A;">Adopted by NVIDIA's Nemotron family of models!</span> <div style="width: 60px; height: 2px; background: linear-gradient(90deg, #E24B4A, transparent);"></div> </div> <p align="center"> 🤗 <a href="https://huggingface.co/collections/TIGER-Lab/openresearcher" target="_blank">HuggingFace</a> |<img src="imgs/slack.png" width="14px" style="display:inline;"> <a href="https://join.slack.com/t/openresearcher/shared_invite/zt-3p0r32cky-PqtZkVjjWIAI14~XwcRMfQ" target="_blank">Slack</a> | <img src="imgs/wechat.svg" width="14px" style="display:inline;"> <a href="https://github.com/TIGER-AI-Lab/OpenResearcher/blob/main/assets/imgs/wechat_group.jpg" target="_blank">WeChat</a> </p> ## Overview **OpenResearcher** is a fully open agentic large language model (30B-A3B) designed for **long-horizon deep research** scenarios. It achieves an impressive **54.8%** accuracy on [BrowseComp-Plus](https://huggingface.co/spaces/Tevatron/BrowseComp-Plus), surpassing performance of `GPT-4.1`, `Claude-Opus-4`, `Gemini-2.5-Pro`, `DeepSeek-R1` and `Tongyi-DeepResearch`. It also demonstrates **leading performance** across a range of deep research benchmarks, including BrowseComp, GAIA, WebWalkerQA, and xbench-DeepSearch. We **fully open-source** the training and evaluation recipe—including data, model, training methodology, and evaluation framework for everyone to progress deep research. ## OpenResearcher Training Dataset Our training dataset consists of **96K** high-quality long-horizon DeepResearch trajectories with **100+ turns** generated by GPT-OSS-120B using its [native browser tools](https://docs.vllm.ai/projects/recipes/en/latest/OpenAI/GPT-OSS.html#usage:~:text=Limitation%20section%20below.-,Tool%20Use,-%C2%B6). To enable scalable and cost-efficient data generation, we deploy a self-hosted search engine over carefully constructed ~11B-token [corpus](https://huggingface.co/datasets/OpenResearcher/OpenResearcher-Corpus) , completely eliminating reliance on external search APIs. ## Format Each row in the dataset contains the following fields: - **qid (int64)**: A unique identifier for each question or task. - **question (string)**: The original deepresearch question compiled from [MiroVerse](https://huggingface.co/datasets/miromind-ai/MiroVerse-v0.1). - **answer (string)**: The final answer to the question. - **messages (list)**: A list of messages representing the GPT-OSS 120B deep research trajectory, including intermediate reasoning steps, tool calls, observations, and model responses throughout the problem-solving process. ## Citation ```bibtex @article{li2026openresearcher, title={{OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis}}, author={Li, Zhuofeng and Jiang, Dongfu and Ma, Xueguang and Zhang, Haoxiang and Nie, Ping and Zhang, Yuyu and Zou, Kai and Xie, Jianwen and Zhang, Yu and Chen, Wenhu}, journal={arXiv preprint arXiv:2603.20278}, year={2026} } ```
数据集信息: - 配置名称:seed_42 特征: - 名称:qid,数据类型:int64 - 名称:question,数据类型:字符串(string) - 名称:answer,数据类型:字符串(string) - 名称:messages,数据类型:列表(list),包含子项: - 名称:channel,数据类型:字符串(string) - 名称:content,数据类型:列表(list),包含子项: - 名称:channel_config,数据类型:结构体(struct),包含字段: - 名称:channel_required,数据类型:布尔型(bool) - 名称:valid_channels,数据类型:字符串列表(list<string>) - 名称:conversation_start_date,数据类型:字符串(string) - 名称:knowledge_cutoff,数据类型:字符串(string) - 名称:model_identity,数据类型:字符串(string) - 名称:reasoning_effort,数据类型:字符串(string) - 名称:text,数据类型:字符串(string) - 名称:tools,数据类型:结构体(struct),包含字段: - 名称:browser,数据类型:结构体(struct),包含字段: - 名称:description,数据类型:字符串(string) - 名称:name,数据类型:字符串(string) - 名称:tools,数据类型:列表(list),包含子项: - 名称:description,数据类型:字符串(string) - 名称:name,数据类型:字符串(string) - 名称:parameters,数据类型:结构体(struct),包含字段: - 名称:properties,数据类型:结构体(struct),包含字段: - 名称:cursor,数据类型:结构体(struct),包含字段: - 名称:default,数据类型:int64 - 名称:type,数据类型:字符串(string) - 名称:id,数据类型:结构体(struct),包含字段: - 名称:default,数据类型:int64 - 名称:type,数据类型:字符串列表(list<string>) - 名称:loc,数据类型:结构体(struct),包含字段: - 名称:default,数据类型:int64 - 名称:type,数据类型:字符串(string) - 名称:num_lines,数据类型:结构体(struct),包含字段: - 名称:default,数据类型:int64 - 名称:type,数据类型:字符串(string) - 名称:pattern,数据类型:结构体(struct),包含字段: - 名称:type,数据类型:字符串(string) - 名称:query,数据类型:结构体(struct),包含字段: - 名称:type,数据类型:字符串(string) - 名称:source,数据类型:结构体(struct),包含字段: - 名称:type,数据类型:字符串(string) - 名称:topn,数据类型:结构体(struct),包含字段: - 名称:default,数据类型:int64 - 名称:type,数据类型:字符串(string) - 名称:view_source,数据类型:结构体(struct),包含字段: - 名称:default,数据类型:布尔型(bool) - 名称:type,数据类型:字符串(string) - 名称:required,数据类型:字符串列表(list<string>) - 名称:type,数据类型:字符串(string) - 名称:type,数据类型:字符串(string) - 名称:content_type,数据类型:字符串(string) - 名称:name,数据类型:字符串(string) - 名称:recipient,数据类型:字符串(string) - 名称:role,数据类型:字符串(string) - 名称:latency_s,数据类型:float64 - 名称:error,数据类型:空值(null) - 名称:attempts,数据类型:int64 - 名称:status,数据类型:字符串(string) - 名称:chunk_idx,数据类型:int64 - 名称:num_chunks,数据类型:int64 划分集: - 名称:train,数据字节数:1117982919,样本数量:6102 下载大小:461369938,数据集总大小:1117982919 其余配置(seed_43至seed_57)的结构与seed_42完全一致,仅数据字节数、下载大小、数据集总大小及样本数量存在细微差异,具体如下: - 配置名称:seed_43:训练集数据字节数1158931340,样本数6102,下载大小477917292,数据集总大小1158931340 - 配置名称:seed_44:训练集数据字节数1108145493,样本数6102,下载大小456027149,数据集总大小1108145493 - 配置名称:seed_45:训练集数据字节数1106900749,样本数6102,下载大小455271833,数据集总大小1106900749 - 配置名称:seed_46:训练集数据字节数1135517221,样本数6102,下载大小468875734,数据集总大小1135517221 - 配置名称:seed_47:训练集数据字节数1179953502,样本数6102,下载大小487937315,数据集总大小1179953502 - 配置名称:seed_48:训练集数据字节数1171147444,样本数6102,下载大小483010306,数据集总大小1171147444 - 配置名称:seed_49:训练集数据字节数1120741628,样本数6102,下载大小461516938,数据集总大小1120741628 - 配置名称:seed_50:训练集数据字节数1139271069,样本数6102,下载大小470031205,数据集总大小1139271069 - 配置名称:seed_51:训练集数据字节数1154584409,样本数6102,下载大小475934015,数据集总大小1154584409 - 配置名称:seed_52:训练集数据字节数1133567180,样本数6102,下载大小467407950,数据集总大小1133567180 - 配置名称:seed_53:训练集数据字节数1024473567,样本数6102,下载大小420595363,数据集总大小1024473567 - 配置名称:seed_54:训练集数据字节数1122591364,样本数6102,下载大小463334159,数据集总大小1122591364 - 配置名称:seed_55:训练集数据字节数1122857864,样本数6100,下载大小462379167,数据集总大小1122857864 - 配置名称:seed_56:训练集数据字节数1150224058,样本数6102,下载大小474643458,数据集总大小1150224058 - 配置名称:seed_57:训练集数据字节数1131677120,样本数6102,下载大小466987581,数据集总大小1131677120 所有配置的数据文件均按照`split: train`划分,路径格式为`{config_name}/train-*`。 本数据集采用MIT开源许可证。 --- <div align="center" style="display: flex; align-items: center; justify-content: center; gap: 8px;"> <img src="imgs/or-logo1.png" style="height: 84px; width: auto;"> <img src="imgs/openresearcher-title.svg" style="height: 84px; width: auto;"> </div> <div align="center"> <a href="https://arxiv.org/abs/2603.20278">arXiv 论文</a> | <a href="https://huggingface.co/papers/2603.20278">Hugging Face 论文</a> | <a href="https://x.com/zhuofengli96475/status/2036475211063648414">Twitter 动态</a> | <a href="https://github.com/TIGER-AI-Lab/OpenResearcher">GitHub 仓库</a> | <a href="https://huggingface.co/datasets/OpenResearcher/OpenResearcher-Dataset">数据集主页</a> | <a href="https://huggingface.co/OpenResearcher/Nemotron-3-Nano-30B-A3B">模型主页</a> | <a href="https://huggingface.co/spaces/OpenResearcher/OpenResearcher">在线演示</a> | <a href="https://huggingface.co/datasets/OpenResearcher/OpenResearcher-Eval-Logs/tree/main">评估日志</a> </div> <div align="center" style="padding: 10px 0; display: flex; align-items: center; justify-content: center; gap: 16px;"> <div style="width: 60px; height: 2px; background: linear-gradient(90deg, transparent, #E24B4A);"></div> <span style="font-size: 22px; font-weight: 600; color: #E24B4A;">已被英伟达(NVIDIA)Nemotron系列模型采用!</span> <div style="width: 60px; height: 2px; background: linear-gradient(90deg, #E24B4A, transparent);"></div> </div> <p align="center"> 🤗 <a href="https://huggingface.co/collections/TIGER-Lab/openresearcher" target="_blank">Hugging Face 合集</a> | <img src="imgs/slack.png" width="14px"> Slack 社区 | <img src="imgs/wechat.svg" width="14px"> <a href="https://github.com/TIGER-AI-Lab/OpenResearcher/blob/main/assets/imgs/wechat_group.jpg" target="_blank">微信社区</a> </p> ## 项目概述 **OpenResearcher** 是一款专为**长周期深度研究**场景设计的完全开源智能体大语言模型(30B-A3B)。其在<a href="https://huggingface.co/spaces/Tevatron/BrowseComp-Plus">BrowseComp-Plus</a>基准测试中实现了54.8%的准确率,超越了`GPT-4.1`、`Claude-Opus-4`、`Gemini-2.5-Pro`、`DeepSeek-R1`以及`Tongyi-DeepResearch`的性能表现。同时,该模型在BrowseComp、GAIA、WebWalkerQA以及xbench-DeepSearch等多项深度研究基准测试中均展现出**领先性能**。我们完全开源了训练与评估流程,包括数据集、模型、训练方法以及评估框架,以推动深度研究领域的发展。 ## OpenResearcher 训练数据集 我们的训练数据集包含由GPT-OSS-120B使用其<a href="https://docs.vllm.ai/projects/recipes/en/latest/OpenAI/GPT-OSS.html#usage:~:text=Limitation%20section%20below.-,Tool%20Use,-%C2%B6">原生浏览器工具(browser tools)</a>生成的96K条高质量长周期深度研究轨迹,每条轨迹包含100+轮次的交互。为实现可扩展且高性价比的数据生成,我们基于精心构建的约110亿Token(Token)语料库<a href="https://huggingface.co/datasets/OpenResearcher/OpenResearcher-Corpus">OpenResearcher-Corpus</a>部署了自研搜索引擎,完全摆脱了对外部搜索API的依赖。 ## 数据格式 数据集的每一行包含以下字段: - **qid(int64)**:每个问题或任务的唯一标识符。 - **question(string)**:源自<a href="https://huggingface.co/datasets/miromind-ai/MiroVerse-v0.1">MiroVerse</a>的原始深度研究问题。 - **answer(string)**:该问题的最终答案。 - **messages(list)**:代表GPT-OSS 120B深度研究轨迹的消息列表,涵盖问题求解过程中的中间推理步骤、工具调用、观测结果以及模型响应等内容。 ## 引用 bibtex @article{li2026openresearcher, title={{OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis}}, author={Li, Zhuofeng and Jiang, Dongfu and Ma, Xueguang and Zhang, Haoxiang and Nie, Ping and Zhang, Yuyu and Zou, Kai and Xie, Jianwen and Zhang, Yu and Chen, Wenhu}, journal={arXiv preprint arXiv:2603.20278}, year={2026} }



