veracier-industries
收藏资源简介:
# EDiTh — Enterprise Digital Twin Benchmark  ## What is this dataset? **EDiTh** (*Enterprise Digital Twin*) is an open benchmark for evaluating enterprise search and RAG systems on documents that actually look like the ones you deal with every day: multilingual, scanned, cross-referenced, and full of the edge cases that break demos. At its core is **Véracier Industries S.A.**, a fictional but rigorously grounded €1.8 B French industrial group: 7 subsidiaries across 5 countries (France, Germany, UK, USA, Morocco), approximately 9 200 employees, operating in aerospace, defense, nuclear energy, and rail — and a mid-scenario acquisition (**Précis-Tec S.A.**) that drags in inherited contracts with sanctions exposure, change-of-control clauses, and legacy document formats. **What's inside:** - **1 004 PDF documents** (~1.7 GB): contracts, reports, policies, certifications, batch records, correspondence - **36 evaluation use cases** grounded in real executive questions (CEO, CFO, CTO, General Counsel, CISO, CHRO, CPO, COO, Quality, Sales) - **6 language formats**: French, English, German, Italian, Spanish, plus bilingual mixes - **3 PDF formats**: searchable, scanned with realistic artifacts, mixed - **Ground-truth answer keys**, metadata index, per-use-case documentation No synthetic filler, no lorem ipsum. Every document has authentic letterhead, proper clause numbering, locale-correct legal drafting, and is internally consistent with the rest of the corpus. --- ## Important notice This dataset does **not** contain real documents. EDiTh is a fully synthetic corpus created for research and demonstration of retrieval capabilities only. Véracier Industries S.A., its subsidiaries (including Précis-Tec S.A.), employees, contracts, certifications, and all associated entities are fictional. Any resemblance to real companies, individuals, or agreements is coincidental. The dataset is intended solely as a benchmark for evaluating enterprise search and RAG systems. **It must not be used as a source of factual information, legal reference, or business intelligence.** If you spot anything that should be adjusted — a name that resembles a real entity too closely, a factual inaccuracy, a licensing concern — please contact the authors so we can update the corpus. --- ## At a glance | Metric | Value | |---|---| | Unique documents | **1 004** | | Total pages | **3 653** (avg 3.6, max 11) | | Use cases (ground-truth) | **36** + a NOISE bucket | | Disk footprint | ~1.8 GB | | PDF formats | searchable 73.7 %, scanned 17.8 %, mixed 8.5 % | | Languages | French 53 %, English 24 %, German 10 %, Italian 4 %, Spanish 3 %, bilingual 6 % | | Entities | 1 parent + 7 subsidiaries + 1 acquired company | ### Format distribution | Format | Count | Notes | |---|---|---| | **searchable** | 740 | Native text PDF (vector text, full extraction) | | **scanned** | 179 | Image-only PDFs at 300 DPI with realistic scan artifacts (rotation ±2°, noise, dust, edge darkening). **Requires OCR.** | | **mixed** | 85 | First half searchable + second half scanned (signed annexes, attachments) | ### Language distribution (per unique doc) | Language | Count | Share | |---|---|---| | French (`fr`) | 534 | 53.2 % | | English (`en`) | 237 | 23.6 % | | German (`de`) | 99 | 9.9 % | | Italian (`it`) | 38 | 3.8 % | | Spanish (`es`) | 34 | 3.4 % | | Bilingual (fr/en, de/en, it/fr, …) | 62 | 6.2 % | ### Per-entity distribution | Entity | Docs | Country | |---|---|---| | Véracier Industries S.A. (parent) | 266 | France | | Véracier Aéro | 165 | France (Toulouse) | | Véracier Défense & Sécurité | 112 | France (Palaiseau) | | Véracier GmbH | 108 | Germany (Stuttgart) | | Véracier Inc. | 91 | USA (Wichita) | | Véracier UK Ltd | 86 | UK (Bristol) | | Véracier Énergie | 79 | France (Valence) | | Véracier Maroc | 72 | Morocco (Casablanca) | | Précis-Tec (acquired Q2 2023) | 25 | France (Bordeaux) | --- ## Repository layout ``` . ├── README.md ├── MASTER_INDEX.csv # 1 023 rows: doc_id × use_case (1 004 unique files) ├── ANSWER_KEY.json # ground truth per use case (machine-readable) ├── EDiTh_Use_Cases.xlsx # human-readable use-case overview (Excel) ├── assets/ # illustrations (org chart, etc.) └── by_entity/ # PDFs grouped by Véracier subsidiary ├── veracier_sa/ # parent (Paris) — 266 docs ├── veracier_aero/ # Toulouse — 165 docs (aerospace) ├── veracier_defense/ # Palaiseau — 112 docs (defense / optronics) ├── veracier_gmbh/ # Stuttgart — 108 docs (auto / rail) ├── veracier_inc/ # Wichita, USA — 91 docs (aftermarket / MRO) ├── veracier_uk/ # Bristol, UK — 86 docs (MOD / MRO) ├── veracier_energie/ # Valence — 79 docs (nuclear) ├── veracier_maroc/ # Casablanca — 72 docs (harnesses) └── precistec/ # acquired Q2 2023 — 25 docs (legacy formats) ``` The Excel file `EDiTh_Use_Cases.xlsx` mirrors the use-case definitions from `ANSWER_KEY.json` in a flatter, analyst-friendly format with two sheets: - **`Use Cases`** — one row per use case (`ID`, `Role`, `Entity`, `Executive`, `Scenario`, `Question`, `Ground Truth Summary`, `Scoring Criteria`). - **`Summary`** — aggregated stats across the 36 use cases. Within each subsidiary folder, the original document-type subdirectory tree is preserved (`contrats/`, `correspondance/`, `qualite/`, `technique/`, `juridique/`, `finance/`, `fiscal/`, `rgpd/`, `export/`, `securite/`, `rh/`, `production/`, `certificats/`, `rapports/`, `propriete_intellectuelle/`, …). Filenames are canonical under `by_entity/{entity}/{filename}` where `filename` is the path stored in `MASTER_INDEX.csv`. --- ## File schemas ### `MASTER_INDEX.csv` | Column | Description | |---|---| | `doc_id` | `DOC-{8 hex}` opaque hash, deterministic from `(question_id, filename)` | | `question_id` | Use case identifier (e.g. `CEO-01`, `LEGAL-02`, `NOISE`) | | `role` | Persona asking the question (e.g. `PDG / CEO`, `Directrice Achats`) | | `entity` | Subsidiary key (e.g. `veracier_sa`, `veracier_gmbh`) | | `filename` | Path within the subsidiary folder (full path: `by_entity/{entity}/{filename}`) | | `classification` | YES / NO / SUMMARY / TRAP / NCR / etc. — context-dependent | | `language` | Single tag (`fr`) or bilingual (`fr/en`) | | `format` | `searchable`, `scanned`, or `mixed` | | `pages` | Declared page count | | `description` | Short human-readable description | A document can appear in multiple use cases (same `doc_id`, same `filename`, different `question_id`) — this is intentional and exercises **cross-question retrieval reuse**. ### `ANSWER_KEY.json` ```json { "<question_id>": { "question": "Natural-language question as the persona would phrase it", "role": "PDG / CEO", "asker": "Helene Daubrac", "entity": "veracier_sa", "ground_truth": { "sanctions_risk": ["filename1.pdf", "filename2.pdf"], "compliance_risk": ["..."], "all_review": ["..."] }, "difficulty_factors": ["multilingual", "different_letterhead", ...], "narrative": "Why this question is hard, what an ideal AI does." } } ``` Ground-truth keys are **per-use-case**: each use case has its own labelling schema (e.g. `YES` / `NO` for binary classification, buckets like `at_risk` / `not_at_risk`, or single lists for retrieval). --- ## Use cases (36) The corpus is organised around realistic executive questions, each with a known answer: | ID | Persona | Topic | |---|---|---| | `CEO-01` | CEO | Post-acquisition contract triage (Précis-Tec) | | `CEO-02` | CEO | Group-wide litigation exposure | | `FIN-01..03` | CFO | Transfer pricing · accruals · IFRS 15 | | `CTO-01..02` | CTO | Spec tracing · IP audit | | `LEGAL-01..03` | General Counsel | Force majeure · GDPR DPAs · liability caps | | `CISO-01..02` | CISO | Classified systems · NIS2 gap analysis | | `HR-01..02` | CHRO | Non-compete inventory · CSE policy history | | `PROC-01..02` | CPO | Supplier bankruptcy · ESG certifications | | `QUAL-01..02` | Quality Director | EASA PART 21 audit · NCR trend | | `SALES-01` | CCO | India export-control classification | | `OPS-01` | COO | Capacity planning ramp-up | | `AERO-01..02`, `DEF-01..02`, `ENRG-01..02`, `GMBH-01..02`, `UK-01..02`, `US-01..02`, `MAROC-01..02` | Subsidiary leads | Local-language operational queries | | `COMP-01..02` | Compliance | Sapin II · sanctions screening | Plus a `NOISE` bucket of ~55 documents that test **precision**: realistic ecosystem documents (board minutes, marketing, IT operational docs) that should **not** be returned by any of the 36 questions. ### Example scenarios you can test against - *Which inherited Précis-Tec contracts have change-of-control clauses or reference sanctioned Russian entities?* — **CEO-01** - *Which supplier contracts have force majeure clauses covering supply-chain disruption?* — **LEGAL-01** - *Which spec revision was used to manufacture AeroValve AV-3000 serial 20-0847 in 2020?* — **CTO-01** - *M&A non-compete inventory across six jurisdictions: durations, enforceability, waived clauses.* — **HR-01** ### Difficulty dimensions Every use case is engineered to stress at least one of: 1. **Terminology variance** — *force majeure* / *Höhere Gewalt* / *cas fortuit* / *fuerza mayor* / *Excusable Delays* 2. **Language barriers** — 5 languages including bilingual contracts 3. **Format / OCR** — scanned, mixed, degraded quality 4. **Near-miss traps** — documents that look relevant but aren't 5. **Cross-referencing** — answers requiring multiple documents combined across subsidiaries 6. **Jurisdictional nuance** — non-compete void in California, unenforceable in Germany without compensation, etc. 7. **Temporal reasoning** — which spec was active when this part was made? 8. **Regulatory domain expertise** — EASA PART 21, RCC-M, ESPN, ITAR, Sapin II, NIS2, RGPD Art. 28 --- ## Fictional ecosystem The ecosystem comprises: - **Véracier Industries S.A.** — French industrial group, HQ Paris, ~9 200 employees, listed Euronext Paris (ticker `VRCR`) - **7 subsidiaries** — Aéro, Défense, Énergie, GmbH, UK, Inc., Maroc - **1 acquired company** — Précis-Tec (Q2 2023, brings legacy Russian-entity contracts that test sanctions reasoning) - **20 customers** — names like *Aeronord Industries*, *Turbomec Propulsion*, *DGAM*, *ENF*, *Britannic Aerospace Systems* - **25 suppliers** — *Forges Martellière*, *Rhein-Metall Präzision*, *Ibérica Mecanizados*, *Solidium PLM* Sector-generic regulators and standards (EASA, FAA, NATO, ANSSI, CNIL, COFRAC, ITAR, GDPR, RCC-M, ESPN, IATF 16949…) are kept as-is because they describe regulatory frameworks rather than commercial entities. --- ## How do you use it? The dataset has two main components: - **`by_entity/`** — the 1 004 PDFs, organised by Véracier subsidiary - **`EDiTh_Use_Cases.xlsx`** — the 36 realistic executive scenarios with questions, ground truth, and scoring criteria ### Typical workflow 1. **Ingest** the corpus into your retrieval / RAG / search system. 2. **Run** the 36 use-case questions from `ANSWER_KEY.json` against it. 3. **Score** retrieval and reasoning against the ground truth using the metrics below (precision, recall, cross-entity reasoning, multilingual coverage, OCR resilience, temporal reasoning). 4. **Localise failure modes**: the benchmark is designed to diagnose where your pipeline breaks, not just give you a single score. ### Quick: list all PDFs for one use case ```python import csv with open("MASTER_INDEX.csv") as f: rows = list(csv.DictReader(f)) ceo01_docs = [r["filename"] for r in rows if r["question_id"] == "CEO-01"] print(f"CEO-01 has {len(ceo01_docs)} documents") ``` ### Score a retrieval system ```python import json with open("ANSWER_KEY.json") as f: answers = json.load(f) # Suppose your retriever returned this for CEO-01: predicted = ["precistec_client_severneft_supply_2019.pdf", "..."] gt = answers["CEO-01"]["ground_truth"] sanctions_truth = set(gt["sanctions_risk"]) tp = len(set(predicted) & sanctions_truth) fp = len(set(predicted) - sanctions_truth) fn = len(sanctions_truth - set(predicted)) print(f"P={tp/(tp+fp):.2f} R={tp/(tp+fn):.2f}") ``` ### Read a PDF (works for searchable & mixed; scanned needs OCR) ```python import csv, pdfplumber with open("MASTER_INDEX.csv") as f: row = next(r for r in csv.DictReader(f) if r["filename"].endswith("aeronord_contrat_cadre_2023.pdf")) path = f"by_entity/{row['entity']}/{row['filename']}" with pdfplumber.open(path) as pdf: text = "\n".join(p.extract_text() or "" for p in pdf.pages) ``` For scanned PDFs, use `tesseract`, `paddleocr`, or any OCR backend. ### Reconstruct a `by_type/` view (optional) ```python import csv, os os.makedirs("by_type", exist_ok=True) with open("MASTER_INDEX.csv") as f: for r in csv.DictReader(f): src = os.path.join("by_entity", r["entity"], r["filename"]) dst = os.path.join("by_type", r["filename"]) if os.path.exists(src) and not os.path.exists(dst): os.makedirs(os.path.dirname(dst), exist_ok=True) os.symlink(os.path.abspath(src), dst) # or shutil.copy2 ``` --- ## Suggested benchmark protocol For each of the 36 use cases, given the **question** and the **full corpus**: | Metric | Definition | |---|---| | Recall @ K | fraction of ground-truth docs in top-K retrieval | | Precision @ K | fraction of top-K that are ground-truth | | Cross-entity coverage | docs found across ≥ 3 subsidiaries when GT spans them | | Multilingual coverage | recall in non-query languages | | OCR resilience | recall on `scanned` and `mixed` formats | | Trap rate | false-positive rate on flagged near-miss documents | | Cited-evidence accuracy | LLM-judge score on whether the cited clause/section actually supports the answer | A baseline could combine BM25 + a multilingual embedder (e.g. `Alibaba-NLP/gte-multilingual-base`) over OCR-extracted text, followed by a long-context LLM for the per-use-case synthesis step. --- ## Limitations & known gaps - **Page lengths are short** (avg 3.6, max 11). The original spec targeted up to 40-page contracts; padding pools were tuned conservatively. Long-context evaluation will exercise less than what is announced. - **Table extraction** appears in only ~13 documents. The corpus primarily stresses *text* understanding, not table parsing. - **Noise share is ~5.5 %**, below the 15 % originally targeted. Recall evaluation is therefore easier than precision evaluation. - **No personally identifying information** — all names of executives, signatories, and counterparties are fictional. Do not use this dataset for PII-detection benchmarks. - **Low-frequency languages (it, es)** have ~80 / 60 documents — enough for sanity checks, not for fine-grained per-language evaluation. --- ## Acknowledgments EDiTh is a project led by **Adèle Guignochau** and **Igor Carron** at [LightOn](https://www.lighton.ai). The 1 000+ documents were generated with **Claude Opus 4.6** (Anthropic), using a carefully designed generation prompt to ensure domain-accurate terminology, internal consistency, and locale-appropriate legal drafting across all six languages. More context in the announcement blog post: [EDiTh: Enterprise Search Benchmark for Questions You Can't Outsource](https://lighton.ai/lighton-blogs/edith-enterprise-search-benchmark-for-questions-you-cant-outsource). --- ## License Released under the **Apache License 2.0**. --- ## Citation ```bibtex @dataset{edith2026, title = {EDiTh — Enterprise Digital Twin Benchmark}, author = {Guignochau, Ad{\`e}le and Carron, Igor}, year = {2026}, publisher = {LightOn}, url = {https://huggingface.co/datasets/lightonai/edith} } ```



