H²EPR-Bench

H²EPR-Bench

An evidence-traceable benchmark for event-process reconstruction.

H²EPR-Bench evaluates real-world event understanding across summarization, event-process reconstruction, prediction, and simulation from fixed evidence contexts.

3,000event instances
6domains
26event categories
11,333reference stages
21direct LLM systems
Ongoingexpansion and follow-up studies
Overview of H2EPR-Bench supporting reconstruction evaluation, event-process analysis, prediction, simulation, and longer-term world-model construction

Key Findings

Frontier gap

Even frontier LLMs fall short.

With the same evidence context, strong LLMs still fail to reliably reconstruct complex events as structured Event-Process Graphs.

Weak event logic

Complete outputs can miss real understanding.

LLMs may gather the right facts and write fluent accounts, yet fail to infer the temporal and causal logic that makes an event coherent.

Process bottleneck

Process reasoning is the hard part.

Models mention facts, but lose stage order, actor actions, and action-result logic across long event sources.

Agent-ready substrate

Structured events are easier to use than raw text.

H²EPR-Bench turns scattered evidence into staged event graphs for analysis, prediction, and market/social simulation.

H²EPR-Bench maps real-world events into hierarchical heterogeneous Event-Process Graphs for reconstruction, analysis, prediction, and simulation.

Background & Motivation

Real-world events are not static text objects. They unfold through stages, actors, actions, dependencies, outcomes, and evidence distributed across sources. Evaluating whether a model understands such events therefore requires more than summaries, timelines, or isolated extraction targets.

Comparison between narrative summaries, sequential timelines, and event-process graphs
From simplified event outputs to structured event-process graph reconstruction.
Why process graphs?

Event-process graphs keep stage order, actor interactions, operations, relations, and evidence links visible in the same representation.

This turns event understanding into a structured reconstruction problem rather than a single text-generation target.

H²EPR-Bench treats events as structured process objects. Starting from fixed evidence contexts, each event is represented as a hierarchical heterogeneous Event-Process Graph that makes stage progression, actor interactions, temporal dependencies, action-result logic, and evidence links explicit.

This representation makes the benchmark useful beyond leaderboard scoring. It can evaluate event-process reconstruction, support analysis of recurring event patterns, and provide structured event substrates for prediction and simulation settings, such as masking future stages, continuing partially observed events, or connecting multiple events into larger market or social-process scenarios.

Reconstruction evaluation with evidence, graph recovery, and diagnostic scoring
Event graphs Linked processes Scenario simulation Societal world model
Long-term vision of linked event-process graphs supporting societal world models
Long-term direction: linking structured events into larger market and social-process world models.

Benchmark Construction

Task

Fixed-evidence reconstruction

Each system receives an event descriptor and a fixed public evidence context, then returns a structured Event-Process Graph.

Reference

Medium-granularity Gold

Official scores are computed against gated Gold references. Public artifacts support browsing, presentation, and reuse.

Graph

Hierarchical heterogeneous representation

The target representation covers macro stages, stage-internal structure, participants, actions, relations, temporal anchors, and evidence links.

Scoring

Diagnostic graph matching

Evaluation compares structural fidelity, temporal consistency, action-result logic, and evidence traceability.

Public release catalog · stages · FinalCascade · Gantt views Reference EPGs (Gated) official scoring references
Public artifacts and gated Gold release boundary

The public repository is designed for browsing, reuse, and inspection. Official benchmark scoring is tied to the gated Gold references, so release artifacts remain useful without exposing answer keys in the open dataset.

Dataset Overview

Unified-3000 covers 3,000 events from 1629 to 2025 across six domains and twenty-six categories. Its frozen evidence packages retain 25,814 verified documents from 84,693 candidate source records, with public artifacts aligned by canonical event IDs.

Dataset summary panel for Unified-3000, including event count, domain mix, category distribution, graph scale, reference EPGs, and LLM systems
Dataset summary panel.
Domain distribution for Unified-3000
Domain distribution.
Top event categories in Unified-3000
Top event categories.
Stage-count distribution in Unified-3000 reference EPGs
Stage-count distribution.

Interactive Explorer

The Explorer is the browsing layer for event metadata, stage tables, timeline views, and public graph summaries.

Screenshot of the H2EPR-Bench interactive explorer
Interactive event-process graph browser.
Inspect events without navigating raw files. Filter the catalog, open an event profile, inspect stage rows, view timelines, and jump to public artifacts.
Explore Events

Representative Event Processes

FinMycelium draft-EPG timelines from the Unified-3000 event collection.

Results & Diagnostics

Direct reconstruction results expose a consistent gap between evidence retention and structured event-process recovery.

01 Fixed evidence

Each system receives the same event descriptor and evidence context.

02 Graph output

The model returns a hierarchical heterogeneous event-process graph.

03 Schema gate

Invalid normalized outputs remain part of the official all-instance score.

04 Graph matching

Outputs are compared against gated medium-granularity Gold references.

05 Diagnostics

Scores separate structural, temporal, causal, and evidence fidelity.

21 direct LLM systems
3,000 events per system
53.00 best H²EPRScore
14 / 21 systems above 90% valid outputs
Schema validity is not enough. Even mostly valid graph outputs remain weak on causal and mechanism reconstruction.
Model traces

Each card summarizes one direct reconstruction baseline across the official score and diagnostic subscores.

Loading model diagnostics...

Diagnostic views

Instance-level spread anchors the diagnostic panel; aggregate readouts and companion checks sit beside and below it.

Event-level Absolute Fidelity distributions by model
Instance spread. Sorted event-level curves show how much variation remains within each model across the 3,000 benchmark events.
Diagnostic readout Aggregate scores hide the bottleneck.

The mean fidelity profile separates source support from process reconstruction. Causal and mechanism recovery remains the most pronounced bottleneck.

Scatter plot comparing evidence fidelity and process organization
Evidence/process profile. Evidence fidelity is plotted against the mean of structural, temporal, and causal fidelity.
Average diagnostic profile across systems with valid graph output
Subscore profile. Mean four-dimensional fidelity across the 20 systems with valid graph output.
Domain quality dotplot
Domain variation. Quality remains bounded across domains, with visible drops in harder event families.
Failure-mode heatmap
Failure modes. Format, schema, endpoint, evidence-reference, and response-envelope failures vary substantially by system.
Valid-only sensitivity diagnostic
Candidate-output sensitivity. Candidate-output fidelity improves for several systems, but does not close the all-slot reconstruction gap.
Token length versus reconstruction quality companion diagnostic
Token use. Longer outputs are not a substitute for reconstructing event-process structure.
Full numeric table

Sortable all-instance results for 21 direct reconstruction systems across 3,000 events.

System Valid outputs (%) H²EPRScore Absolute Structure Temporal Causal Evidence
Loading leaderboard data...

Citation

@misc{h2eprbench2026,
  title = {H²EPR-Bench: An Evidence-Traceable Benchmark for Event-Process Reconstruction},
  author = {AgenticFinLab},
  year = {2026},
  note = {Forthcoming}
}