← All work
04 — The archivist / tool / Shahriar LabsOpen source

CH-Bench.

Evidence for agent memory

01 / The idea

A zero-dependency Python evaluation harness for retrieval and answer quality, with adapters for different memory systems.

CH-Bench measures persistent-memory systems through a common adapter instead of presenting an impressive graph as evidence. The selected résumé and local documentation describe retrieval metrics such as recall@k, MRR and nDCG, answer evaluation for correctness, groundedness and abstention, and efficiency reporting. The harness supports LongMemEval, LoCoMo and custom long-horizon recall tasks. These are available evaluation capabilities, not a claim of a winning benchmark result. Its adapter boundary lets a memory implementation be compared using the same task and reporting workflow.

02 / What matters
  • 01

    Retrieval and answer evaluation are separate

  • 02

    Four-method system adapter

  • 03

    Groundedness and abstention checks

  • 04

    Zero-dependency core

Source / scope

Source: the selected September 2026 résumé. Status and scope are stated as documented; no current public release or adoption is inferred.

Read the selected résumé (PDF) ↓
Next / Keep exploring

There’s more.

Have a challenge in this space?

Let’s talk about it ↗
Shihab portraying Goku in an anime-inspired costume, keeping his own facial identity
The other lives of Shihab

01/ 08

The pursuit of better.

Goku — Dragon Ball

A training arc never really ends.

Personal cosplay artwork. Same face. A different world.Enter the six-form training arc ↗
A little energy from the whole world.
1 of 8