# Notes — first scored corpus run (2026-04-28, run 02) Detailed report: `reports/2026-04-28-1942-run.md` ## Result vs. rubric **Threshold per [phase-0-validation.md §6.3](../../docs/phase-0-validation.md):** - Per-seed: average ≥ 3.5 AND no dimension < 3.0 - Across corpus: ≥ 80% of seeds pass | Seed | Avg | Min | Status | |---|---|---|---| | 01-aristotle | 4.14 | 3 | ✓ | | 02-habit-coach | **4.57** | 4 | ✓ | | 03-redline-ai | 4.43 | 4 | ✓ | | 04-skillswap | 4.43 | 4 | ✓ | | 05-cyclist-thing | 4.00 | 3 | ✓ (vague-seed handled correctly) | | 06-eventstream | 3.43 | **2** | ✗ | | 07-pet-translator | 4.14 | 3 | ✓ | | 08-quiet-hours | 4.14 | 3 | ✓ | **Pass rate: 7 / 8 = 87.5%** — above the 80% threshold. **Corpus mean across all seeds: 4.16 / 5.** ## Per-dimension averages Across the 8 seeds: - modelCoverage: 3.75 - modelAccuracy: 3.50 - modelParsimony: 4.50 - constraintCapture: 4.00 - assumptionDetectionQuality: 4.13 - riskDetectionQuality: 4.50 - voiceAndCharacter: 4.13 The weakest dimension is **modelAccuracy** (3.50) — pulled down by the eventstream 2 and lower scores on other technical/abstract seeds. **modelParsimony** and **riskDetectionQuality** lead at 4.50 — Gemma is doing well at not over-populating and at surfacing substantive risks. ## The eventstream failure Eventstream produced **5 blocks but 0 associations / 0 constraints / 0 requirements**. The judge correctly penalized: - modelAccuracy 2/5 — "lists entities but provides no associations" - constraintCapture 2/5 — "failed to capture critical operational constraints" This is a stochastic regression — the previous (unscored) run on the same seed had `5b 4a 3c 4r`. Two iteration paths: 1. **Retry to confirm stochasticity.** Re-run the seed alone; if it produces full output, it's variance. 2. **Strengthen the prompt.** Add explicit guidance: "Even for technical/protocol seeds, you MUST populate all four arrays. If you can't infer associations, prefer fewer high-confidence ones over none." Recommend (1) first, then (2) if the regression repeats. ## Confidence calibration check Self-reported `overallConfidence` from the model: | Seed | Self-conf | Judge avg | |---|---|---| | 01-aristotle | 0.90 | 4.14 | | 02-habit-coach | 0.90 | 4.57 | | 03-redline-ai | 0.90 | 4.43 | | 04-skillswap | 0.90 | 4.43 | | 05-cyclist-thing | **0.30** | 4.00 | | 06-eventstream | 0.92 | 3.43 | | 07-pet-translator | 0.90 | 4.14 | | 08-quiet-hours | 0.90 | 4.14 | Note **eventstream self-rated 0.92 but actually scored worst** — model overconfident on its weakest output. This is a known small-model failure mode. Worth probing in iteration: tighten the confidence rule to penalize self-rating high when several arrays are empty. The cyclist-thing low confidence (0.30) correlates correctly with the judge's high score on it — both recognize the seed forced sparse-but-correct output, not low quality. ## Wins worth keeping 1. **Vague-seed handling proven twice.** Cyclist Thing scored 4.00 (passes) — the "under-suggest, push back, refuse to fabricate" prompt instruction works. 2. **Constraint dimension averages 4.00** — the schema-tightening that removed `kind:'constraint'` from blocks is paying off. 3. **Risk detection averages 4.50** — the prompt asking for "specific failure modes tied to specific elements" is producing substantive risks across all seed types. 4. **No structural lint warnings on any seed.** Validation rules hold. ## Iteration candidates (priority order for Day 7) 1. **Retry eventstream.** Cheap — confirms stochastic vs. systemic. 2. **Add "all four arrays must be populated" guidance to `generate.md`** — only if (1) confirms it's systemic. 3. **Tighten confidence calibration:** instruct the model to rate confidence based on its OWN model coverage, not on the seed's apparent richness. (e.g., "If you generated 0 associations, your overall confidence cannot exceed 0.5.") 4. **Habit Coach style as gold standard.** It scored highest — use its prompts/output as a reference for what "good" looks like in further refinements. ## Open questions - **Should we automate stochasticity testing?** Run each seed N times and report variance? Useful but expensive (~7 min per seed × 8 seeds × 3 runs = ~3 hours). - **Judge sycophancy risk.** Gemma is judging Gemma. Worth running a known-bad model (deliberately broken Aristotle) to see if the judge correctly scores it low. Could be a Day 7 sanity check. - **Token cost summary in reports.** Currently we don't sum totals. Trivial to add. ## Next session — Day 7 1. Retry eventstream alone, confirm whether the regression repeats 2. If systemic: iterate `generate.md` on the populate-all-arrays rule and re-run the corpus 3. If stochastic: just note it and proceed — single-seed variance is real 4. Run a deliberate-bad sanity check on the judge (mutate Aristotle to remove 80% of elements, see if score drops)