Sets up the Socrata project repo with: docs/ — strategy and design documents - idea.md: full product vision - implementation-plan.md: Phase 0 + Phase 1 MVP plan - phase-0-validation.md: 2-week validation experiment strategy - phase-0-plan.md: concrete Phase 0 build plan - phase-0-results.md: Phase 0 gate outcome — GO for MVP - sysml-modeling.md: metamodel + SE discipline + validation rules - socrates.md: agent character, surfaces, modes, prompts, lifecycle - sync.md: bidirectional text↔diagram sync engineering - design-source/: HTML/CSS/JS handoff bundle from Claude Design phase-0/ — validated harness (CLI, no UI, no DB) - LM Studio (local OpenAI-compatible) generation + detection + judge - PlantUML rendering for SysML model visualization - 10-seed corpus (8 working + 2 holdouts) - 5 corpus runs with iteration history in reports/ - Final gate: 10/10 pass, mean 4.32/5, holdouts validated Phase 1 MVP scope and milestones documented in implementation-plan.md.
6.0 KiB
Notes — second scored corpus run, post generate.md iteration (2026-04-28, run 03)
Detailed report: reports/2026-04-28-2115-run.md
Change applied
Tightened src/prompts/generate.md with explicit guidance separating constraints from requirements:
Each item from
seed.constraintsbelongs in EXACTLY ONE place — theconstraintsarray. Do NOT also output it as a requirement. (...) Therequirementsarray should contain things derived fromseed.desiredOutcome— NOT a re-encoding ofseed.constraints.
Headline result
| Run 02 (baseline) | Run 03 (this) | Δ | |
|---|---|---|---|
| Pass rate | 7/8 = 87.5% | 7/8 = 87.5% | flat |
| Corpus mean | 4.16 | 4.11 | -0.05 |
| Constraint↔requirement duplication | present every seed | eliminated every seed | ✓ fix |
Score did not move materially. Underlying output structure is meaningfully cleaner.
Per-seed deltas
| Seed | Run 02 | Run 03 | Δ |
|---|---|---|---|
| 01-aristotle | 4.14 / 3 | 4.14 / 3 | flat |
| 02-habit-coach | 4.57 / 4 | 4.43 / 4 | -0.14 |
| 03-redline-ai | 4.43 / 4 | 4.14 / 3 | -0.29 |
| 04-skillswap | 4.43 / 4 | 4.43 / 4 | flat |
| 05-cyclist-thing | 4.00 / 3 | 3.00 / 2 | -1.00 ✗ |
| 06-eventstream | 3.43 / 2 | 4.14 / 3 | +0.71 ✓ |
| 07-pet-translator | 4.14 / 3 | 4.43 / 4 | +0.29 |
| 08-quiet-hours | 4.14 / 3 | 4.14 / 3 | flat |
4 seeds flat, 2 up, 2 down. Eventstream's improvement was a re-run earlier in the session before this corpus pass — already known to be stochastic. So in pure-iteration terms: 4 flat, 1 up (pet-translator), 3 down (habit-coach, redline-ai, cyclist-thing).
Where the fix shows in the artifacts
Cross-checked four seeds with explicit seed.constraints. All show clean separation:
01-aristotle
C[3]: Must never output a complete solution... ; Response latency under 1.2s... ; FERPA tenancy...
R[2]: facilitate low-stakes re-engagement ; guide student without providing final solutions
↑ both derived from desiredOutcome, not from constraints
03-redline-ai
C[4]: Word integration ; no cross-customer training ; audit trail ; SOC 2
R[2]: first-pass redlines complete in minutes ; GC reviews AI suggestions
↑ both from desiredOutcome
06-eventstream
C[3]: Kafka wire-compat ; single binary 100k/s ; Apache 2.0
R[3]: backpressure first-class ; consumer-driven flow control default ; LLM-pipelined workloads
↑ from desiredOutcome and the protocol-design vision
08-quiet-hours
C[3]: GCal+Outlook RW ; timezone-respect ; no auto-blocking first 30 days
R[3]: 2 contiguous focus blocks/week ; invisible negotiation ; reduce social cost
↑ from desiredOutcome
Compare to run-02 where the same seeds had REQ-NNN entries that literally restated seed.constraints text.
The Cyclist Thing regression
Run 02: 2 blocks, 1 association, 0 constraints, 0 requirements, conf 0.30 → score 4.00 (vague handled correctly)
Run 03: 4 blocks, 3 associations, 0 constraints, 1 requirement, conf 0.30 → score 3.00 (judge penalizes the inventing)
Hypothesis: the new prompt's "derive requirements from seed.desiredOutcome" instruction nudged the model to invent a requirement from "Something useful for them" rather than leave the array empty. The vague-seed escape hatch ("if too vague, return sparse and set confidence < 0.3") is still in the prompt but didn't fire as cleanly.
Confidence stayed correctly low (0.30) — the model knows it's overreaching. So the lower judge score reflects real over-generation that the model itself flagged.
Cheap fix candidate: add to generate.md — "If seed.desiredOutcome is too vague to derive specific requirements (e.g., generic 'something useful'), leave the requirements array EMPTY rather than inventing a placeholder."
Judge sanity check (separate, this session)
Deliberately broken Aristotle (1 block, 0 of everything else, no findings) → judge gave 1.29 avg, min 1. Real Aristotle scores 4.14. The judge isn't rubber-stamping.
Decision matrix for next move
We've now validated:
- ✓ Generation works on a 4B local model
- ✓ Detection works (assumptions / risks / inconsistencies, with post-validation guard)
- ✓ Conversation loop works
- ✓ Propose-and-apply works (oneOf JSON-schema necessary for small-model op generation)
- ✓ LLM-as-judge works (sanity-checked against deliberately broken input)
- ✓ Vague-seed handling works (Cyclist Thing scored well in run-02; regressed in run-03 from a known cause)
- ✓ Feasibility-suspect handling works (Pet Translator)
- ✓ Constraint↔requirement separation now clean across all seeds
- ✓ Pass rate ≥ 80% (currently 87.5%)
- ✓ Corpus mean ~4.1 (well above the 3.5 per-seed threshold)
Phase 0 has met its quantitative pass criteria as defined in phase-0-validation.md §6.3.
Three plausible next moves:
A. One more iteration. Add the vague-desired-outcome guard, re-run. Cheap (~15 min). Likely restores cyclist-thing pass.
B. Day 9: human review. The user (you) reads the diagrams and findings and assigns scores. LLM-judge has been the sanity check; final scoring per the plan is human.
C. Day 10: final round with holdouts. Run Carbon Coach + Telemetry Lite under --final to verify the prompts generalize beyond what they were iterated against. Per the plan this is the gate decision.
Recommendation: A → C. Skip B until after C, since holdouts are the real test of whether iteration overfit.
Token usage roll-up (this run)
Approximately:
- Generation: 8 × ~3k in / ~1k out = 24k in / 8k out (Sonnet/local)
- Detection: 8 × ~5k in / ~1k out = 40k in / 8k out
- Judge: 8 × ~2.5k in / ~0.3k out = 20k in / 2.4k out
- Total ~84k in / 18.4k out for full scored corpus
- LLM cost: $0 (local)
- Wallclock: ~10 minutes
Updated open questions
- Iteration efficiency: 3 corpus runs done (~30 min compute, ~30 min author). Next iterations should be smaller, more targeted (single-seed re-runs) before triggering full corpus.
- Variance reporting: would help to run each seed N=3 times and report (mean, stdev) per dimension. Defer.