ScholarScope-data
收藏资源简介:
ScholarScope Data 是一个研究会话日志数据集,源自ScholarScope(一款AI驱动的资助与奖学金研究助手)。该数据集记录了为用户资金需求档案匹配并呈现的资助机会信息,包含约40个研究会话中的100多条机会记录。每条记录代表一个由ScholarScope流水线从实时网络资源中提取和排序的资助机会,以JSON Lines格式存储,并自动生成Parquet格式镜像。数据集语言为英语,采用Apache-2.0许可证。 记录字段包括:会话唯一标识符(session_id)、会话创建时间戳(created_at)、用户提交的四字段资金需求档案(user_profile,包含申请人来源国家/地区、教育水平、研究领域、研究/资金目标、资助类型、资金需求、额外资格限制、截止日期偏好等)、规范化搜索查询(query)、机会名称(opportunity_title)、组织(organization)、直接链接(source_url)、截止日期(deadline)、奖励金额(funding_amount)、AI生成的资格说明(eligibility_summary)、匹配度评分(match_score,0-100分)和AI生成的匹配理由摘要(reasoning_summary)。 数据来源于ScholarScope Space,每个研究会话仅记录评分最高的机会。AI生成的字段可能不准确或过时,截止日期可能为滚动或无明确日期;该数据集是操作性的研究日志数据,并非精心策划的基准数据集。
ScholarScope Data is a dataset containing research session logs from ScholarScope, an AI-driven funding and scholarship research assistant. It records information on funding opportunities matched and presented to users based on their funding need profiles. The dataset includes over 100 opportunity records from approximately 40 research sessions, each representing a funding opportunity extracted and ranked by the ScholarScope pipeline from real-time web sources. Data is stored in JSON Lines format with an automatically generated Parquet mirror. The dataset is in English and licensed under Apache-2.0. Each record contains fields such as session unique identifier (session_id), session creation timestamp (created_at), user-submitted four-field funding need profile (user_profile, including key information like applicants country/region, education level, research field, research/funding goals, funding type, funding needs, additional eligibility constraints, deadline preferences, and other preferences), normalized search query (query), opportunity title (opportunity_title), organization (organization), direct link to provider page (source_url), application deadline (deadline), funding amount (funding_amount), AI-generated short eligibility summary (eligibility_summary), match score (match_score, 0-100), and AI-generated reasoning summary (reasoning_summary). Data is sourced from ScholarScope Space, with only the highest-scoring opportunity recorded per research session. AI-generated fields may be inaccurate or outdated; deadlines may be rolling or unspecified; this dataset is operational research log data and not a curated benchmark dataset.





