# Phase 0 — Results _Closed 2026-04-28._ ## Verdict **GO for MVP.** The 10-seed corpus (8 working + 2 holdouts under final-round protocol) all pass the rubric defined in [phase-0-validation.md §6.3](phase-0-validation.md). Pass rate **100%**, corpus mean **4.32 / 5**, no individual dimension below 3 on any seed. The brain works on a 4-billion-parameter local model — the assumption Phase 0 was designed to test. ## Final round numbers Run report: [phase-0/reports/2026-04-28-FINAL-run.md](../phase-0/reports/2026-04-28-FINAL-run.md). Model: `google/gemma-4-e4b` via LM Studio. | # | Seed | Difficulty | Avg | Min | | |---|---|---|---|---|---| | 1 | 01-aristotle | easy | 4.43 | 4 | ✓ | | 2 | 02-habit-coach | medium | 4.29 | 3 | ✓ | | 3 | 03-redline-ai | medium | 4.43 | 4 | ✓ | | 4 | 04-skillswap | medium | 4.43 | 4 | ✓ | | 5 | 05-cyclist-thing | failure-prone | 4.29 | 3 | ✓ | | 6 | 06-eventstream | hard | 4.14 | 3 | ✓ | | 7 | 07-pet-translator | failure-prone | 4.14 | 3 | ✓ | | 8 | 08-quiet-hours | medium | 4.43 | 4 | ✓ | | 9 | **09-carbon-coach** _(HOLDOUT)_ | medium | **4.43** | 4 | ✓ | | 10 | **10-telemetry-lite** _(HOLDOUT)_ | hard | **4.14** | 3 | ✓ | **Holdouts both passed.** They were never used during prompt iteration — both `expectedDifficulty` levels (medium and hard) cleared the rubric on first contact, ruling out corpus-overfitting as a confound. ## What worked (shipping to MVP unchanged) 1. **`character.md`** — the Socrates voice prompt. Voice scored 4 across every seed. Anti-patterns ("never open with Great question", "no bullets in prose") held even on a 4B model. Promotes to `apps/web/lib/llm/prompts/socrates/character.md` verbatim. 2. **`generate.md`** — the seed → SysML model prompt, with two iterations applied during Phase 0: - **Run-03 fix:** explicit "constraint↔requirement separation" — items in `seed.constraints` go ONLY to the constraints array; requirements derive from `seed.desiredOutcome`. Eliminated content-duplication across all 8 seeds. - **Run-05 fix:** "vague desiredOutcome rule" — if the desired outcome is too vague to derive specific requirements, leave the requirements array empty rather than invent placeholders. 3. **Detection prompts** — `detect-assumptions.md`, `detect-risks.md`, `detect-inconsistencies.md`. Risk-detection averaged 4.50 across the corpus (highest of all dimensions). Promotes unchanged. 4. **`review.md`** — the conversational mode prompt. Numbered-options affordance works. Promotes unchanged. 5. **`propose.md` + the `oneOf` JSON schema** — the original "open `kind: string`" schema let Gemma cram all op fields into the kind name. Switching to `oneOf` per-op-kind made constrained decoding produce real ops. **Critical lesson:** small models need their JSON schema to enforce variant grammar, not just shapes. 6. **`judge.md`** — with the run-05 extension to invert coverage / accuracy / constraint-capture for failure-prone seeds (not just parsimony and risk). Survived a sanity check (deliberately broken Aristotle scored 1.29, real Aristotle 4.43+ — the judge isn't sycophantic). ## What didn't, and what changed - **Constraint duplication** (run-02 → run-03 fix). Every seed had REQ-NNN entries that literally restated `seed.constraints` text. Fixed with explicit prompt boundary. - **Vague-seed over-population** (run-03 → run-05 fix). The "derive requirements from desiredOutcome" instruction nudged Gemma to invent placeholder requirements on Cyclist Thing's vague seed. Fixed with explicit empty-array rule. - **Hallucinated element refs in detection** (existed throughout, mitigated by code). 4–5% of element ids in findings don't resolve to real model elements. Mitigated by `lib/socrates/post-validate.ts` — strips bad refs, drops findings with no remaining refs. - **Confidence > 1 emission** (run-04 schema rejection). Gemma occasionally emits `confidence: 1.05` etc despite the JSON Schema's `maximum: 1`. Fixed at the Zod layer with `.transform(n => clamp(n, 0, 1))`. - **Stochasticity** (eventstream regression in run-02). At temperature 0.3 the same seed produces different fullness across runs. Variance can be ±1 on the average score. Mitigated by accepting that one bad single-run isn't dispositive. ## What changed in the metamodel **Nothing.** [sysml-modeling.md](sysml-modeling.md) types remained valid across the corpus. Gemma successfully generated all 10 models against the metamodel without forcing us to add or change types. Specifically: - `system | actor | block | constraint` block kinds: sufficient for all 10 seeds - `association | composition | aggregation | generalization | constraintApplies` association kinds: sufficient - 4 property types (`string | number | boolean | enum`): sufficient - The Constraint-Block separation (Section 4.5 in sysml-modeling.md) was the right call — once the prompt enforced it, every seed produced clean separation **One open question for MVP** (deferred): EventStream and Telemetry Lite are technical/protocol seeds. They strain the metamodel because their natural vocabulary (`Stream`, `Backpressure Signal`, `Producer client library`) doesn't map cleanly to the PM-shaped Block/Actor distinction. Phase 0's metamodel handled them adequately (both passed) but a future "technical-systems profile" might add `Service`, `Capability`, or `Protocol` as block sub-kinds. **Not needed for MVP.** ## What gets promoted to MVP (file-by-file) Per the [phase-0-plan.md §12 hand-off contract](phase-0-plan.md): | Phase 0 artifact | MVP destination | |---|---| | `phase-0/src/prompts/character.md` | `apps/web/lib/llm/prompts/socrates/character.md` (verbatim) | | `phase-0/src/prompts/generate.md` | `apps/web/lib/llm/prompts/socrates/generate.md` (verbatim — used in seed-screen handoff M6) | | `phase-0/src/prompts/detect-*.md` | `apps/web/lib/llm/prompts/socrates/detect-*.md` (verbatim, M8) | | `phase-0/src/prompts/review.md` | `apps/web/lib/llm/prompts/socrates/review.md` (verbatim, M6 dock) | | `phase-0/src/prompts/propose.md` | `apps/web/lib/llm/prompts/socrates/propose.md` (M7) | | `phase-0/src/prompts/judge.md` | `apps/web/test/eval/judge.md` (CI eval) | | `phase-0/src/types.ts` | `apps/web/lib/sysml/model.ts` (canonical types) | | `phase-0/src/generate/json-schema.ts` | `apps/web/lib/llm/schemas/sysml-model.ts` (LLM JSON Schema for generation) | | `phase-0/src/socrates/post-validate.ts` | `apps/web/lib/socrates/post-validate.ts` (the [docs/socrates.md §6.3](socrates.md) reference-safety guard) | | `phase-0/src/socrates/propose.ts` (oneOf schema) | `apps/web/lib/socrates/propose.ts` (op schema lessons preserved) | | `phase-0/seeds/` | `apps/web/test/eval/seeds/` (regression eval suite) | | `phase-0/src/render/to-plantuml.ts` | `apps/web/lib/export/plantuml.ts` ("export model as PlantUML" feature) | ## Iteration history 5 corpus runs across 2 sessions. Reports committed to `phase-0/reports/`: - `2026-04-28-1735-run.md` — first corpus pass, no judge (visual-only review) - `2026-04-28-1942-run.md` — first scored run (run-02), 7/8 = 87.5% pass - `2026-04-28-2115-run.md` — run-03 with constraint↔requirement fix, 7/8, mean -0.05 but structure cleaner - `2026-04-28-2129-run.md` — run-04 added vague-outcome guard, 6/7 (one schema-validation failure) - `2026-04-28-2144-run.md` — run-05 with confidence-clamp + extended judge inversion, **8/8 = 100%**, mean 4.39 - `2026-04-28-FINAL-run.md` — **10/10 with holdouts, mean 4.32** Iteration notes per run in `phase-0/reports/notes-run-0N.md`. ## Cost LLM cost: **$0.** All inference was local via LM Studio (gemma-4-e4b on macOS). Wall-clock per scored corpus run: ~10–15 minutes for 8 seeds; ~17 minutes for 10 seeds (final round). ## Constraints / things to know going into MVP 1. **Hallucination floor.** Even with the post-validate guard, expect ~4–5% of LLM-emitted element refs to need stripping. The MVP `lib/socrates/post-validate.ts` should preserve the guard from day one. 2. **Variance.** At temperature 0.3 a single seed can swing ±1 on average score across runs. MVP's Sonnet path may be more stable but should not be assumed to be deterministic. 3. **Small-model JSON-schema lessons.** Use `oneOf` for tagged unions, clamp ranges instead of rejecting, default-fill optional arrays. These survive into MVP for the local-LM-Studio user-config path; can be relaxed for the Sonnet path but cost nothing to keep. 4. **Holdout discipline.** Whenever MVP iterates a prompt, hold out at least 20% of the eval corpus from iteration. Prevents overfitting that the corpus alone can't detect. 5. **The 4B-vs-Sonnet portability check (M6 task).** Phase 0 validated against gemma-4-e4b. MVP M6 must re-run the corpus through the Anthropic provider to confirm scores hold. Prompts validated locally are not guaranteed to transfer 1:1. ## Recommendation Begin **MVP M1** (visual port) per [implementation-plan.md](implementation-plan.md). The brain works. The risk that motivated Phase 0 is retired.