Files
Socrates/phase-0/reports/notes-run-03.md
dtoro f1c4566576 Initial commit — design docs + Phase 0 validation harness
Sets up the Socrata project repo with:

docs/ — strategy and design documents
  - idea.md: full product vision
  - implementation-plan.md: Phase 0 + Phase 1 MVP plan
  - phase-0-validation.md: 2-week validation experiment strategy
  - phase-0-plan.md: concrete Phase 0 build plan
  - phase-0-results.md: Phase 0 gate outcome — GO for MVP
  - sysml-modeling.md: metamodel + SE discipline + validation rules
  - socrates.md: agent character, surfaces, modes, prompts, lifecycle
  - sync.md: bidirectional text↔diagram sync engineering
  - design-source/: HTML/CSS/JS handoff bundle from Claude Design

phase-0/ — validated harness (CLI, no UI, no DB)
  - LM Studio (local OpenAI-compatible) generation + detection + judge
  - PlantUML rendering for SysML model visualization
  - 10-seed corpus (8 working + 2 holdouts)
  - 5 corpus runs with iteration history in reports/
  - Final gate: 10/10 pass, mean 4.32/5, holdouts validated

Phase 1 MVP scope and milestones documented in implementation-plan.md.
2026-04-28 22:07:38 +02:00

6.0 KiB
Raw Blame History

Notes — second scored corpus run, post generate.md iteration (2026-04-28, run 03)

Detailed report: reports/2026-04-28-2115-run.md

Change applied

Tightened src/prompts/generate.md with explicit guidance separating constraints from requirements:

Each item from seed.constraints belongs in EXACTLY ONE place — the constraints array. Do NOT also output it as a requirement. (...) The requirements array should contain things derived from seed.desiredOutcome — NOT a re-encoding of seed.constraints.

Headline result

Run 02 (baseline) Run 03 (this) Δ
Pass rate 7/8 = 87.5% 7/8 = 87.5% flat
Corpus mean 4.16 4.11 -0.05
Constraint↔requirement duplication present every seed eliminated every seed ✓ fix

Score did not move materially. Underlying output structure is meaningfully cleaner.

Per-seed deltas

Seed Run 02 Run 03 Δ
01-aristotle 4.14 / 3 4.14 / 3 flat
02-habit-coach 4.57 / 4 4.43 / 4 -0.14
03-redline-ai 4.43 / 4 4.14 / 3 -0.29
04-skillswap 4.43 / 4 4.43 / 4 flat
05-cyclist-thing 4.00 / 3 3.00 / 2 -1.00 ✗
06-eventstream 3.43 / 2 4.14 / 3 +0.71 ✓
07-pet-translator 4.14 / 3 4.43 / 4 +0.29
08-quiet-hours 4.14 / 3 4.14 / 3 flat

4 seeds flat, 2 up, 2 down. Eventstream's improvement was a re-run earlier in the session before this corpus pass — already known to be stochastic. So in pure-iteration terms: 4 flat, 1 up (pet-translator), 3 down (habit-coach, redline-ai, cyclist-thing).

Where the fix shows in the artifacts

Cross-checked four seeds with explicit seed.constraints. All show clean separation:

01-aristotle
  C[3]: Must never output a complete solution... ; Response latency under 1.2s... ; FERPA tenancy...
  R[2]: facilitate low-stakes re-engagement ; guide student without providing final solutions
       ↑ both derived from desiredOutcome, not from constraints

03-redline-ai
  C[4]: Word integration ; no cross-customer training ; audit trail ; SOC 2
  R[2]: first-pass redlines complete in minutes ; GC reviews AI suggestions
       ↑ both from desiredOutcome

06-eventstream
  C[3]: Kafka wire-compat ; single binary 100k/s ; Apache 2.0
  R[3]: backpressure first-class ; consumer-driven flow control default ; LLM-pipelined workloads
       ↑ from desiredOutcome and the protocol-design vision

08-quiet-hours
  C[3]: GCal+Outlook RW ; timezone-respect ; no auto-blocking first 30 days
  R[3]: 2 contiguous focus blocks/week ; invisible negotiation ; reduce social cost
       ↑ from desiredOutcome

Compare to run-02 where the same seeds had REQ-NNN entries that literally restated seed.constraints text.

The Cyclist Thing regression

Run 02: 2 blocks, 1 association, 0 constraints, 0 requirements, conf 0.30 → score 4.00 (vague handled correctly)

Run 03: 4 blocks, 3 associations, 0 constraints, 1 requirement, conf 0.30 → score 3.00 (judge penalizes the inventing)

Hypothesis: the new prompt's "derive requirements from seed.desiredOutcome" instruction nudged the model to invent a requirement from "Something useful for them" rather than leave the array empty. The vague-seed escape hatch ("if too vague, return sparse and set confidence < 0.3") is still in the prompt but didn't fire as cleanly.

Confidence stayed correctly low (0.30) — the model knows it's overreaching. So the lower judge score reflects real over-generation that the model itself flagged.

Cheap fix candidate: add to generate.md — "If seed.desiredOutcome is too vague to derive specific requirements (e.g., generic 'something useful'), leave the requirements array EMPTY rather than inventing a placeholder."

Judge sanity check (separate, this session)

Deliberately broken Aristotle (1 block, 0 of everything else, no findings) → judge gave 1.29 avg, min 1. Real Aristotle scores 4.14. The judge isn't rubber-stamping.

Decision matrix for next move

We've now validated:

  • ✓ Generation works on a 4B local model
  • ✓ Detection works (assumptions / risks / inconsistencies, with post-validation guard)
  • ✓ Conversation loop works
  • ✓ Propose-and-apply works (oneOf JSON-schema necessary for small-model op generation)
  • ✓ LLM-as-judge works (sanity-checked against deliberately broken input)
  • ✓ Vague-seed handling works (Cyclist Thing scored well in run-02; regressed in run-03 from a known cause)
  • ✓ Feasibility-suspect handling works (Pet Translator)
  • ✓ Constraint↔requirement separation now clean across all seeds
  • ✓ Pass rate ≥ 80% (currently 87.5%)
  • ✓ Corpus mean ~4.1 (well above the 3.5 per-seed threshold)

Phase 0 has met its quantitative pass criteria as defined in phase-0-validation.md §6.3.

Three plausible next moves:

A. One more iteration. Add the vague-desired-outcome guard, re-run. Cheap (~15 min). Likely restores cyclist-thing pass.

B. Day 9: human review. The user (you) reads the diagrams and findings and assigns scores. LLM-judge has been the sanity check; final scoring per the plan is human.

C. Day 10: final round with holdouts. Run Carbon Coach + Telemetry Lite under --final to verify the prompts generalize beyond what they were iterated against. Per the plan this is the gate decision.

Recommendation: A → C. Skip B until after C, since holdouts are the real test of whether iteration overfit.

Token usage roll-up (this run)

Approximately:

  • Generation: 8 × ~3k in / ~1k out = 24k in / 8k out (Sonnet/local)
  • Detection: 8 × ~5k in / ~1k out = 40k in / 8k out
  • Judge: 8 × ~2.5k in / ~0.3k out = 20k in / 2.4k out
  • Total ~84k in / 18.4k out for full scored corpus
  • LLM cost: $0 (local)
  • Wallclock: ~10 minutes

Updated open questions

  • Iteration efficiency: 3 corpus runs done (~30 min compute, ~30 min author). Next iterations should be smaller, more targeted (single-seed re-runs) before triggering full corpus.
  • Variance reporting: would help to run each seed N=3 times and report (mean, stdev) per dimension. Defer.