H²EPR-Bench

H²EPR-Bench

An evidence-traceable benchmark for event-process reconstruction.

Given the same fixed, multi-source evidence, can a model reconstruct how a complex event actually unfolded? H²EPR-Bench evaluates the process structure behind the facts.

3,000event instances
6domains
26event categories
25,814verified documents
3,000reference EPGs
21LLMs evaluated
H2EPR-Bench connects fixed evidence to event-process graphs, reconstruction evaluation, analysis, and future research directions

Key Findings

Best overall score

53.00 H²EPRScore

The leading system still leaves substantial headroom on complete event-process reconstruction.

Strongest capabilities

Facts and coarse order

Source attribution and stage-order scores exceed 82, showing that models often identify the right material.

Weakest capabilities

Relations remain below 19

Action–outcome, causal, and explicit temporal relations are much harder to reconstruct.

Mechanism bottleneck

9.83% path retention

Mediated mechanism paths are rarely preserved, pinpointing the gap between event facts and event logic.

H²EPR-Bench measures the difference between knowing event facts and reconstructing the process that connects them.

Background & Motivation

Real-world events are not static text objects. They unfold through stages, actors, actions, dependencies, outcomes, and evidence distributed across sources. Evaluating whether a model understands such events therefore requires more than summaries, timelines, or isolated extraction targets.

Comparison between narrative summaries, sequential timelines, and event-process graphs
From simplified event outputs to structured event-process graph reconstruction.
Why process graphs?

Event-process graphs keep stage order, actor interactions, operations, relations, and evidence links visible in the same representation.

This turns event understanding into a structured reconstruction problem rather than a single text-generation target.

H²EPR-Bench treats events as structured process objects. Starting from fixed evidence contexts, each event is represented as a hierarchical heterogeneous Event-Process Graph that makes stage progression, actor interactions, temporal dependencies, action-result logic, and evidence links explicit.

H²EPR-Bench centers on reconstruction, while the EPG representation also creates a common substrate for cross-event analysis, continuation, prediction, and simulation research.

Reconstruction evaluation with evidence, graph recovery, and diagnostic scoring
Event graphs Linked processes Scenario simulation Societal world model
Long-term vision of linked event-process graphs supporting societal world models
Long-term direction: linking structured events into larger market and social-process world models.

Benchmark Construction

Task

Fixed-evidence reconstruction

Each system receives an event descriptor and the same fixed evidence package, then returns a structured Event-Process Graph.

Reference

Expert-finalized references

Agent–expert review, human adjudication, and final sign-off produce one consistent medium-granularity reference EPG per event.

Graph

Hierarchical heterogeneous representation

The target representation connects macro stages, meso episodes, micro-level participants, actions, outcomes, relations, and evidence support.

Scoring

Diagnostic graph matching

Evaluation reports structural, temporal, causal, and evidence fidelity, together with a unified H²EPRScore.

FinMycelium Draft EPG broad process coverage Expert finalization consistent evaluation reference
FinMycelium Draft EPG refined through agent–expert review and human adjudication into an expert-finalized reference EPG

FinMycelium enables graph construction at scale; expert review then aligns scope, process relevance, and evidence support before each reference EPG is finalized.

Dataset Overview

H²EPR-Bench covers 3,000 events from 1629 to 2025 across six domains and twenty-six categories. Its evidence packages retain 25,814 verified documents from 84,693 retrieved source records, with every artifact aligned by a stable event ID.

H²EPR-Bench summary panel with event count, domain mix, category distribution, graph scale, reference EPGs, and evaluated models
Dataset summary panel.
H²EPR-Bench domain distribution
Domain distribution.
Top H²EPR-Bench event categories
Top event categories.
Stage-count distribution in H²EPR-Bench reference EPGs
Stage-count distribution.

Interactive Explorer

The Explorer is the browsing layer for event metadata, stage tables, timeline views, and public graph summaries.

Screenshot of the H2EPR-Bench interactive explorer
Interactive event-process graph browser.
Inspect events without navigating raw files. Filter the catalog, open an event profile, inspect stage rows, view timelines, and jump to public artifacts.
Explore Events

Representative Event Processes

Representative FinMycelium Draft-EPG timelines from the 3,000-event collection.

Results & Diagnostics

Direct reconstruction results expose a consistent gap between evidence retention and structured event-process recovery.

01 Fixed evidence

Each system receives the same event descriptor and evidence context.

02 Graph output

The model returns a hierarchical heterogeneous event-process graph.

03 Schema gate

Every benchmark instance contributes to the official score after schema normalization.

04 Graph matching

Outputs are aligned with expert-finalized reference EPGs.

05 Diagnostics

Scores separate structural, temporal, causal, and evidence fidelity.

21 language models evaluated
3,000 events per system
53.00 best H²EPRScore
14 / 21 systems above 90% valid outputs
Schema validity is not enough. Even mostly valid graph outputs remain weak on causal and mechanism reconstruction.
Model traces

Each card summarizes one direct reconstruction baseline across the official score and diagnostic subscores.

Loading model diagnostics...

Diagnostic views

Instance-level spread anchors the diagnostic panel; aggregate readouts and companion checks sit beside and below it.

Event-level Absolute Fidelity distributions by model
Instance spread. Sorted event-level curves show how much variation remains within each model across the 3,000 benchmark events.
Diagnostic readout Aggregate scores hide the bottleneck.

The mean fidelity profile separates source support from process reconstruction. Causal and mechanism recovery remains the most pronounced bottleneck.

Scatter plot comparing evidence fidelity and process organization
Evidence/process profile. Evidence fidelity is plotted against the mean of structural, temporal, and causal fidelity.
Average diagnostic profile across systems with valid graph output
Subscore profile. Mean four-dimensional fidelity across the 20 systems with valid graph output.
Domain quality dotplot
Domain variation. Quality remains bounded across domains, with visible drops in harder event families.
Failure-mode heatmap
Failure modes. Format, schema, endpoint, evidence-reference, and response-envelope failures vary substantially by system.
Valid-only sensitivity diagnostic
Candidate-output sensitivity. Candidate-output fidelity improves for several systems, but does not close the all-slot reconstruction gap.
Token length versus reconstruction quality companion diagnostic
Token use. Longer outputs are not a substitute for reconstructing event-process structure.
Full numeric table

Sortable all-instance results for 21 direct reconstruction systems across 3,000 events.

System Valid outputs (%) H²EPRScore Absolute Structure Temporal Causal Evidence
Loading leaderboard data...

Citation

@misc{h2eprbench2026,
  title = {H²EPR-Bench: An Evidence-Traceable Benchmark for Event-Process Reconstruction},
  author = {AgenticFinLab},
  year = {2026},
  note = {Forthcoming}
}