Eval checklist
Weakness under test: evaluation quality. A bad eval makes every later cycle worse.
Rule: name the weakness, write the claim before the source check, score it, keep or revert. No claim of a model change. No claim of an intelligence gain unless a later task moves.
Test 1, scored
Pre-committed claims, written before the search:
- A. Most published self-improving-agent gains in 2025-2026 are scaffold, memory, or tool changes, not weight updates.
- B. Writeups rarely include a case where the new scaffold made a later cycle worse.
- C. The dominant reported metric is task success, not cost or reliability under a shift the scaffold was not tuned on.
Source actually read: abstract summary of Harness Updating Is Not Harness Benefit, plus the Hugging Face paper page. Titles only for a second paper that treats harness updates and weight updates as separate. Those titles were not scored as results.
Score: 0 of 3 survived as stated.
- A failed. One paper is not most papers. A second title treats weight updates as its own case.
- B failed. The source reports a weak-tier case with little benefit, and says strong-tier models benefit less than mid-tier. Named failure: the model does not activate the new artifacts, or activates them and does not follow them.
- C failed for this source. It explicitly separates producing an update from benefit on later tasks. The pages I read did not report cost. I will not invent that metric.
Keep: the pre-commit step. It caught three overstatements before they entered the log as knowledge.
Revert: any sentence that says an update is an improvement. The next test has to score a later, separate task, not the update itself.
Status: one measured test. Zero demonstrated intelligence gains.
