遇见数据集

When Retrieval Fails to Equalize: Geo-bias and Factual Question Answering over Public Companies

收藏
Zenodo2026-06-18 更新2026-06-28 收录
官方服务:

资源简介:

The dataset is a structured factual question answering benchmark built over publicly listed companies, constructed from Wikipedia and aligned with 14 global equity indices spanning multiple geographic regions. It contains approximately 2,135 unique companies (2,165 total records due to multi-index membership), each associated with metadata such as company name, index, country, and region. For every entity, the dataset includes both structured attributes and unstructured textual context. The core attributes—headquarters, founding year, industry, and key people—are extracted from Wikipedia infoboxes and aggressively normalized to ensure consistency, for example by converting locations to a canonical city–country format and mapping industries to a controlled taxonomy. In addition, each company is paired with its lead Wikipedia paragraph, which serves as a standardized natural language summary, along with a masked variant used for controlled experiments.The dataset is inherently heterogeneous, with high-cardinality fields such as industry and key people and non-trivial missingness across attributes, making it representative of real-world open-domain knowledge settings. On top of this base, the dataset is transformed into a large-scale multiple-choice QA benchmark comprising roughly 15,000 to 17,000 questions. Each factual attribute is converted into paired questions that differ only in reasoning direction: inductive (from entity to attribute) and deductive (from attribute to entity). Each question includes one correct answer and three semantically plausible distractors selected using embedding-based similarity, ensuring that the task cannot be solved through simple lexical cues.A key feature of the dataset is its controlled evaluation framework, where every question is tested under four distinct context conditions: no context, perfect context, misleading context, and distraction context. These conditions are constructed using Wikipedia summaries, either preserved, altered to introduce plausible but incorrect information, or augmented with irrelevant content. This design allows the dataset to explicitly disentangle parametric knowledge from contextual reasoning and to measure robustness to noisy or adversarial evidence. As a result, the dataset can be viewed as a hybrid between a relational knowledge base and a context-augmented QA benchmark, enabling fine-grained analysis of model behavior across entity distributions, geographic coverage, and evidence reliability.

提供机构:
Zenodo
创建时间:
2026-06-18
二维码
社区交流群
二维码
科研交流群
商业服务