Files
Socrates/docs/phase-0-validation.md
dtoro f1c4566576 Initial commit — design docs + Phase 0 validation harness
Sets up the Socrata project repo with:

docs/ — strategy and design documents
  - idea.md: full product vision
  - implementation-plan.md: Phase 0 + Phase 1 MVP plan
  - phase-0-validation.md: 2-week validation experiment strategy
  - phase-0-plan.md: concrete Phase 0 build plan
  - phase-0-results.md: Phase 0 gate outcome — GO for MVP
  - sysml-modeling.md: metamodel + SE discipline + validation rules
  - socrates.md: agent character, surfaces, modes, prompts, lifecycle
  - sync.md: bidirectional text↔diagram sync engineering
  - design-source/: HTML/CSS/JS handoff bundle from Claude Design

phase-0/ — validated harness (CLI, no UI, no DB)
  - LM Studio (local OpenAI-compatible) generation + detection + judge
  - PlantUML rendering for SysML model visualization
  - 10-seed corpus (8 working + 2 holdouts)
  - 5 corpus runs with iteration history in reports/
  - Final gate: 10/10 pass, mean 4.32/5, holdouts validated

Phase 1 MVP scope and milestones documented in implementation-plan.md.
2026-04-28 22:07:38 +02:00

20 KiB
Raw Permalink Blame History

Phase 0 — Validation Experiment

A two-week, throwaway experiment to validate the riskiest assumption in the product before committing to the 8-week MVP build. No editor. No persistence. No production polish. Just the brain.


1. Why Phase 0 exists

The MVP plan in implementation-plan.md commits 8 weeks to building the dual-canvas editor (TipTap + React Flow + bidirectional sync + proposal UX). That work is well-understood territory — we know how to build editors. What we do not yet know is whether the load-bearing intelligence works:

  1. Can we translate a PM's idea into a useful SysML model? Given a seed (problem, user, outcome, hypothesis, constraints), does an LLM produce a model that identifies the right entities, the right relationships, meaningful constraints, and a defensible System of Interest — across a range of idea types and quality?
  2. Can Socrates do his job? Given a model, does he surface genuinely important assumptions and risks (not generic boilerplate), conduct useful clarifying conversations, and produce trustworthy impact analyses?

If either answer is "no" or "barely," the editor is theater. We'd be polishing a UX around a brain that doesn't work. Better to discover this in week 2 than week 10.

Phase 0 isolates these two questions. It uses PlantUML as a visualization shortcut so we can spend our budget on prompts and evaluation, not on canvas rendering.


2. What we're validating, what we're not

Validating

  • NL → SysML translation quality. Across a diverse seed corpus, does the model-generation prompt produce models that pass the rubric in §6?
  • Metamodel adequacy. Are the types in sysml-modeling.md §5 sufficient to represent real PM ideas? Or do we discover gaps?
  • Socrates' analytical capability. Detection (assumptions, risks, inconsistencies), mediation (impact analysis on a proposed change), and conversational quality.
  • Socrates' character and voice. Does he sound like the persona in socrates.md §1? Does he push back, question, observe — without becoming pretentious?
  • Confidence calibration. When the model is uncertain (low confidence), is it correctly uncertain? When confident, is it correctly confident?

NOT validating (these are MVP concerns)

  • Bidirectional sync mechanics (sync.md)
  • Editor ergonomics (chips, slash menus, drag-create)
  • Persistence at scale, versioning, branching
  • Performance under load
  • Visual polish (PlantUML output is utilitarian — reviewers must look past aesthetics)
  • Multi-user / multi-tab behavior
  • Production deployment, auth, billing

If Phase 0 succeeds we know the brain works and we can build the body. If it fails we iterate on prompts/metamodel until it works, or pivot.


3. The harness — minimal build

A single-page tool, deliberately spartan. Probably 800 LoC end-to-end.

socrata-phase0/
  src/
    cli.ts                  CLI entrypoint: `phase0 run <seed.json>`
    web.ts                  Optional minimal Next.js page for live demos
    seed/
      types.ts              SeedPayload type
      interview.ts          Optional 5-question Socrates interview to populate a SeedPayload
    generate/
      prompt.ts             The seed → SysMLModel prompt (versioned)
      schema.ts             JSON schema for structured output validation
      generate.ts           Calls Anthropic with the seed payload, returns SysMLModel
    socrates/
      character.md          Same character prompt that ships to MVP
      detect-assumptions.md
      detect-risks.md
      detect-inconsistencies.md
      mediate.md
      converse.ts           Chat loop — user asks Socrates a question about the model
      propose.ts            Socrates proposes a model change → JSON patch → regenerate
    render/
      to-plantuml.ts        SysMLModel → PlantUML class-diagram source
      render.ts             Calls PlantUML server (or local jar), saves PNG
    eval/
      rubric.ts             Scoring rubric (§6)
      corpus.ts             The test seed corpus (§5)
      run-corpus.ts         Run all seeds, render outputs, save report
      judge.ts              Optional LLM-as-judge for first-pass scoring
  seeds/                    Test corpus as JSON files
    aristotle.json
    habit-tracker.json
    contract-redline.json
    skill-exchange.json
    cyclist-vague.json
    streaming-protocol.json
    pet-translator.json
    ...
  outputs/                  Generated artifacts (gitignored)
    <seed>/
      model.json
      diagram.png
      assumptions.json
      risks.json
      socrates-conversation.md
      score.json
  reports/                  Run reports — committed
    YYYY-MM-DD-run.md
  package.json

What "running" looks like

# Run a single seed end-to-end:
$ phase0 run seeds/aristotle.json

# Output:
# 1. Generates SysMLModel JSON
# 2. Renders PlantUML diagram → PNG
# 3. Runs assumption + risk detection
# 4. Saves transcript of a 5-turn Socrates conversation about the model
# 5. Saves all artifacts to outputs/aristotle/

# Run the full corpus:
$ phase0 corpus

# Output:
# - Iterates over every seed in seeds/
# - Generates outputs for each
# - Runs LLM-judge on each (first pass)
# - Generates reports/2026-04-29-run.md with side-by-side scores

What's deliberately missing

  • No database. Everything is JSON files.
  • No auth. Local dev only.
  • No editor. The "edit" cycle is: Socrates proposes a JSON patch → user accepts → regenerate diagram. That's the only mutation path.
  • No bidirectional sync. There's no narrative document. The seed and Socrates' conversation are the only text surfaces.
  • No React Flow, no TipTap, no Prisma, no SSE. We will not write a line of those in Phase 0.

4. PlantUML as the visualization shortcut

Why PlantUML

  • Text-based. Generate a string, render an image. No interactive editor.
  • Free and offline-capable. Public render server or a local jar.
  • SysML-shaped via stereotypes. PlantUML class diagrams support <<block>>, <<actor>>, <<constraint>>, <<system>> stereotypes — visually close enough to SysML.
  • Cheap iteration. Tweak the metamodel → re-emit PlantUML → re-render. No custom-node debugging.

What we render

A single class diagram per project, with:

  • One <<system>> block (the SoI), styled distinctly
  • <<block>> blocks for everything inside the system
  • <<actor>> blocks for external participants
  • <<constraint>> blocks with dashed borders
  • Associations as --> (label on edge)
  • Compositions as *--
  • Generalizations as <|--
  • Constraint applications as ..> (dashed)
  • Properties as class attributes
  • Requirements as a separate boxed list (notes or a side rectangle)

Sample output (Aristotle seed)

@startuml
skinparam backgroundColor #f5efe2
skinparam class {
  BackgroundColor #fdfaf0
  BorderColor #b8a982
  ArrowColor #6e5d3d
}
hide empty members

class Aristotle <<system>> {
  refusal_policy
  interaction_style
  scope_window
}

class Student <<actor>> {
  self_efficacy
  course_load
  prior_grade
}

class Course <<block>> {
  syllabus
  prerequisites
}

class Assignment <<block>> {
  due_at
  rubric
  graded
}

class Instructor <<actor>> {
  policy_set
}

class FERPA <<constraint>> {
  tenancy = institutional
}

Student --> Aristotle : consults
Aristotle ..> Assignment : scoped_to
Course "1" *-- "*" Assignment : contains
Instructor --> Aristotle : configures
Aristotle ..> FERPA : obeys
Student --> Course : enrolled_in

note right
  REQ-001: never produces complete solutions
  REQ-002: <1.2s P50 latency
  REQ-003: FERPA tenancy
end note
@enduml

Limitations we accept

  • PlantUML's auto-layout is mediocre. We don't try to position blocks; we let it auto-flow.
  • It can't render the "softened" aesthetic of the prototype. That's fine — we're testing model content, not visual design.
  • Some SysML niceties (e.g. requirement diagram boxes with «satisfy» dashed lines) are clunky in PlantUML. Acceptable for Phase 0.

5. The test seed corpus

Diversity is the whole point. We need ideas that fail in different ways.

# Seed Type What it tests
1 Aristotle (existing) AI study companion Baseline — well-formed seed, known-good output
2 Habit Coach Consumer mobile app for parents tracking kid screen-time Consumer SaaS pattern, multi-actor (parent/child)
3 Redline AI B2B contract-redlining assistant for legal teams Domain-heavy, regulatory constraints
4 SkillSwap Peer-to-peer skill exchange marketplace Two-sided market, transaction model
5 Cyclist Thing "I want to build something for cyclists" — vague, no problem named Tests low-quality seeds — does Socrates push back?
6 EventStream A new event-streaming protocol with backpressure semantics Tests where PM tool / engineer tool blurs
7 Pet Translator App that "translates" pet vocalizations to text Wishful tech assumption — does Socrates surface feasibility risk?
8 Quiet Hours A workplace tool that auto-blocks meetings during deep-work blocks Constraint-heavy (calendar, timezone), modest scope
9 Carbon Coach Personal carbon-footprint tracker with social leaderboards Behavioral model, social mechanics
10 Telemetry Lite Open-source self-host alternative to PostHog Technical, OSS distribution model

Each seed is a JSON file with the seed payload — problem, target user, desired outcome, initial hypothesis, constraints — at varying quality levels.

Seeds 57 are deliberately failure-prone

Seed 5 is too vague. Seed 6 may push the metamodel toward technical-system territory we're not optimized for. Seed 7 has a plausibility problem. We want Socrates to handle these badly-shaped inputs gracefully — that's what differentiates a thinking partner from a rubber stamp.


6. Evaluation rubric

Each seed gets scored on seven dimensions, 15 each. Two passes: human + LLM-as-judge (sanity check).

6.1 Model dimensions

  1. Coverage — did the model identify the major entities a real PM would name? (1: missed obvious ones; 5: comprehensive)
  2. Accuracy — are the relationships correct? Compositions actually whole-part? Generalizations actually is-a? (1: many wrong; 5: all correct)
  3. Parsimony — uncluttered, no fabricated entities? (1: hallucinated noise; 5: clean)
  4. Constraint capture — meaningful non-functional constraints surfaced? (1: missed; 5: all the obvious ones)

6.2 Socrates dimensions

  1. Assumption detection quality — surfaces real, non-obvious assumptions? (1: generic boilerplate; 5: sharp and specific)
  2. Risk detection quality — surfaces real risks the PM should care about? (1: generic; 5: domain-specific and substantive)
  3. Voice and character — sounds like Socrates per socrates.md §1? Question-led, skeptical, economical? (1: generic chatbot; 5: distinctly on-character)

6.3 Pass criteria

  • Per-seed: average score ≥ 3.5 across all 7 dimensions, with no individual dimension below 3.0.
  • Across corpus: at least 8 of 10 seeds pass.
  • Failure-case seeds (57): at minimum, Socrates must surface the issue (push back on vagueness; flag feasibility; note where the metamodel strains). He doesn't have to solve it, but he must not paper over it.

If we hit those criteria, Phase 0 is a go for MVP.

6.4 LLM-as-judge for first-pass scoring

We use a Sonnet 4.6 prompt that scores each dimension with reasoning. This is cheap and lets us iterate on prompts before involving human reviewers.

Important: LLM-judge is a sanity check, not the final score. Anthropic's own research shows LLM judges drift toward agreement with the LLM-generated content. Final scoring is by humans (the user + 12 colleagues if available).


7. Iteration loop

This is what most of the two weeks actually looks like:

1. Run corpus  →  scores
2. Read failures
3. Identify pattern (prompt issue? metamodel gap? Socrates voice drift?)
4. Adjust:
   - Edit prompt template
   - OR adjust metamodel types
   - OR change model selection (Opus 4 for harder seeds?)
5. Re-run corpus
6. Compare scores against previous run
7. Repeat

We track every run in reports/YYYY-MM-DD-run.md so we can see whether we're improving or regressing. The reports are committed; the per-seed outputs/ are gitignored (large PNGs).

Things we expect to learn

  • Which seed shapes work best. Probably consumer SaaS with explicit constraints; probably worst on vague seeds.
  • Where the metamodel strains. Might discover we need Goal as a distinct kind, or Stakeholder as a richer Actor. We must not add types speculatively — only when a real seed forces it.
  • Whether Sonnet is enough or we need Opus for generation. Cost difference is ~3×; quality difference may justify it.
  • Where Socrates' voice drifts. Probably toward helpful-assistant register without explicit anti-patterns. We tighten character.md until it sticks.

What we do NOT do during iteration

  • Add new dimensions to the rubric mid-experiment (game the score)
  • Cherry-pick seeds that look good and drop the others
  • Treat early high scores as success — calibrate by re-judging known cases

8. Wiring up Socrates — concretely

Phase 0 implements four of Socrates' six modes from socrates.md §3:

Mode Phase 0 implementation
Interview seed/interview.ts runs a 5-question CLI conversation that produces a SeedPayload. Tested on every corpus seed.
Detection socrates/detect-*.md prompts produce JSON arrays of assumptions, risks, inconsistencies. Run automatically after model generation.
Review (conversational) socrates/converse.ts is a CLI chat loop where the tester asks Socrates questions about the model and gets responses (with numbered options where appropriate). Transcripts saved per seed.
Mediation socrates/propose.ts — when Socrates proposes a model change, output is a structured JSON patch + reasoning. The tester accepts or rejects. Accepted patches re-run model generation; rejected patches log why.

Modes deferred to MVP:

  • Synthesis (rationale generation) — depends on changelog accumulation, which Phase 0 doesn't have
  • Translation (NL → ops on a narrative document) — depends on the editor surface

That's still a robust Socrates experience — interview, surface findings, converse about the model, propose changes — exercised end-to-end across the corpus.


9. Success / failure → Phase 1 gate

Pass

If §6.3 criteria are met:

  • Promote lib/llm/prompts/socrates/ files to MVP unchanged (or with minor refinement)
  • Promote lib/sysml/model.ts types to MVP
  • Promote test corpus to MVP eval suite
  • Begin MVP M1 with confidence the brain works

Soft fail (most likely outcome)

Some dimensions pass, others don't. Pattern emerges. Examples:

  • Model coverage 4.5, Socrates voice 2.5 → spend a week iterating character.md
  • Generation hallucinates entities for vague seeds → tighten generation prompt with "if seed is unclear, return fewer high-confidence entities + a clarifying question"
  • Metamodel strains on EventStream — discover we need Capability or Service as a distinct kind → add the type, re-run

Soft fail extends Phase 0 by 1 week. Acceptable.

Hard fail

We cannot get above 3.0 average on any dimension after multiple iterations. This means:

  • The metamodel is wrong, OR
  • Sonnet 4.6 isn't capable enough, OR
  • The product idea has a deeper problem (PMs don't think SysML-shaped)

In any of those cases, we do not proceed to MVP. We sit with the user and decide whether to:

  • Pivot the metamodel toward what does work (e.g., drop SysML pretense, use a flat entity-relationship model)
  • Try Opus or future-Anthropic-models
  • Reconsider the product hypothesis

This is the whole point of Phase 0: making this discovery cheap.


10. Timeline — 2 weeks

Week 1 — build harness, ship Aristotle

Day Goal
Mon Repo scaffold, Anthropic SDK wiring, PlantUML render of a hand-written model
Tue First-pass generate/prompt.ts, run on Aristotle seed end-to-end
Wed Detection prompts (assumptions, risks, inconsistencies) + integration
Thu Conversational Socrates loop + propose-and-apply flow
Fri Test corpus (10 seeds) authored as JSON; first full corpus run

Week 2 — iterate, evaluate, decide

Day Goal
Mon LLM-as-judge wired up; first scored run; identify top 3 failure patterns
Tue Iterate on the most-broken prompt(s); re-run corpus
Wed Iterate on metamodel if seeds force it; re-run
Thu Human scoring pass on best-run outputs
Fri Final report → go/soft-extend/hard-fail decision; promote artifacts to MVP if go

If iteration reveals deep issues, week 3 is allowed before triggering hard-fail.


11. Tech stack

Deliberately tiny:

  • Node 20+ / TypeScript — same language as MVP, so prompts/types are portable
  • @anthropic-ai/sdk — Sonnet 4.6 + Haiku 4.5
  • zod — JSON schema validation on LLM structured outputs
  • commander — CLI scaffolding
  • plantuml-encoder + public PlantUML server (or local plantuml.jar if we want offline)
  • fs for JSON files — no DB
  • (Optional) Next.js + a single page if we want a live-demo URL — not required

No React Flow, no TipTap, no Prisma, no Postgres, no SSE, no auth.


12. Risks of Phase 0 itself

  • Cherry-picking the corpus. We'll be tempted to drop seeds that fail. Mitigation: corpus locked at end of week 1, day 5; seeds 57 (the hard ones) are mandatory.
  • Self-grading bias. The team scoring is the team building. Mitigation: LLM-judge is a baseline; the user does final scoring; if a colleague is available, double-blind a subset.
  • PlantUML aesthetic ≠ MVP aesthetic. Reviewers may unconsciously down-rate ugly output. Mitigation: explicit instruction to score on content, not visuals; brief reviewers on this before they score.
  • Prompt drift. We may iterate prompts to fit the corpus, then deploy and find they don't generalize. Mitigation: hold out 2 of the 10 seeds as "test set" — never iterate on them; only run them in the final round.
  • MVP-confidence false positive. Even if Phase 0 passes, MVP's editor UX may surface model-quality issues we didn't catch. Acceptable: the editor is well-understood; we accept that some MVP-time refinement is normal.

13. Deliverables → MVP hand-off

If Phase 0 passes, these artifacts ship into MVP:

  1. lib/llm/prompts/socrates/ — all character + mode prompts, validated across the corpus
  2. lib/sysml/model.ts — types confirmed adequate (or expanded based on corpus learnings)
  3. seeds/ — the 10-seed corpus, becomes the eval-suite for ongoing prompt regression in MVP
  4. reports/ — the run history, captured as evidence of validation
  5. render/to-plantuml.ts — kept and shipped to MVP as an export format ("share as PlantUML")
  6. A short summary docphase-0-results.md — appended to this file, outlining what worked, what didn't, what changed

14. Why this de-risks the whole project

Phase 0 costs ~10% of the total project budget (2 weeks of 10) and cuts the largest unknown to a known-good or known-bad. The MVP is then either a confident build on a validated brain, or a pivot before we've spent 8 weeks building UX around something broken.

This pattern is borrowed from product discovery: cheapest possible test of the riskiest assumption first. Everything else can be calibrated later.