Task — engineering-spec@1

"Calibrated LLM judge + golden-set growth + internal LongMemEval-style benchmark"

draftTASK-MEMORY-245
module memory · class product · priority p0 · created 2026-07-08 · shipped null
depends on TASK-MEMORY-209 · blocks none

TASK-MEMORY-245: Calibrated LLM judge + golden-set growth + internal LongMemEval-style benchmark

1. Description

judge agreement 75-90% vs human labels; 100+ golden cases; five ability buckets tracked per release

Migrated 2026-07-08 from the memory improvement backlog, folded into the task system as class: improvement. Source report refs: R46, R56.

Acceptance criteria