Kit review: value and model-agnosticism

Status: refreshed 2026-09-01 (supersedes the Aug 30 draft on PR #22).
Open actions: kit-review-backlog.md

Verdict: Still valuable for teams that want hexagonal + DDD + TDD/XFN discipline and primarily run agents in Cursor (or another host that loads AGENTS.md + progressive skills). Still not model-agnostic in the strong sense: content is mostly portable markdown; discovery, MCP install, skill format, and the live EDD driver remain Cursor- and OpenAI-compatible-shaped. Public copy now names that split; skill-trigger evals still do not invoke a model.


What this kit is

Waykit is a governance + procedure pack for coding agents, plus a TypeScript CLI (kit) for bootstrap, MCP compose, security scan, and eval harnesses. Core ideas:

  1. Thin always-on bootstrap (AGENTS.md) with on-demand skills/SOPs/philosophy.
  2. Lifecycle roles (agent-*) and stack profiles (lang-* / framework-*).
  3. Architecture invariants (hexagonal, DDD, vertical slices, clean code) with applicability / opt-out (CODING_PHILOSOPHY.md, docs/ADRs).
  4. Eval-Driven Development for agent tool routing/schemas (kit eval).
  5. MCP profile catalog and multi-IDE pointer files via kit export-rules.

Rough inventory (as of this refresh): ~45 skills, SOPs, dual eval layers (evals/suites + evals/edd), MCP profiles, install/init, nightly live-model workflow (when keyed).


What still holds (strengths)

  1. Context budget as design: thin AGENTS.md, on-demand loads, one MCP profile, kit measure-context / kit check.
  2. Lifecycle taxonomy: phase → skill routing, handovers, tests-as-catalog, orchestrator + specialists.
  3. EDD shape: YAML metrics + JSONL cases, scripted merge gate, optional live path, shadow/prod→JSONL story (docs/edd.md).
  4. Architecture floor: explicit defaults with documented applicability/opt-out (landed after the Aug 30 draft).
  5. Supply-chain hygiene: layout verify, lockfile, audit, secrets-out-of-repo MCP fragments.
  6. Portable content layer: philosophy/SOPs/role prose travel; runtime discovery does not.

What changed since the Aug 30 draft

Finding (Aug 30)Status now
Architecture invariants feel non-optionalImproved: applicability and opt-out + seed ADRs (#36)
Closed-loop / live EDD aspirationalImproved: deeper EDD suites, shadow path, nightly edd-live.yml (still skips without KIT_EVAL_API_KEY)
Default CI is scripted keyword driverStill true: intentional merge gate; docs now say so clearly
Skill-trigger evals are theaterStill true: kit/src/edd/run_evals.ts does not invoke a model or assert required_patterns / required_output_sections
Multi-IDE peer-depth oversoldImproved: README, FAQ, and llms.txt name Cursor as the reference host; stubs vs discovery is explicit
Skill length budget slippingStill true: agent-prune / agent-orchestrator / agent-debug / agent-copy over ~150 lines
Thin stack profilesStill true: several framework-* / lang-* skills ~38–48 lines
Process weightStill true: shortcuts exist; default narrative is multi-phase

Model-agnosticism (focused answer)

LayerAgnostic?Notes
Philosophy / SOPs / role proseMostly yesMarkdown procedures; model-neutral instructions
Host discovery & MCPNoCursor skill format + .cursor/mcp.json dominate
Multi-IDE exportCosmeticPointer files ≠ equal capability
Skill routing evals (run_evals)N/ANo model in the loop
EDD live runnerWeaklyOpenAI-compatible /chat/completions; Anthropic key only via compatible gateway
Default CI proofNoScripted driver; live is nightly + secret

Bottom line: Model-agnostic at the documentation layer; provider-shaped (OpenAI-compatible) + Cursor-centric at the runtime layer.


Value judgment (unchanged in substance)

Worth adopting when: Cursor (or manual skill mapping), shared hexagonal/TDD norms, MCP/tool agents that need a routing harness, thin always-on rules.

Weak fit when: Equal first-class Claude Code / Gemini / Copilot / Windsurf depth, out-of-the-box multi-model proof CI, architecture-flexible guidance, or a minimal prompt pack without lifecycle ceremony.

Net: High concept and structure value; medium–rising execution value (EDD and honesty improved); still not “any model / any host, same guarantees.”


Highest-leverage remaining work

Tracked in kit-review-backlog.md. Top three:

  1. Make skill-trigger evals assert required_patterns / required_output_sections (copy already says they are registration / prompt hygiene, not a live model run).
  2. Enforce role skill line budget; cut or deepen thin stack profiles.
  3. Surface a shorter day-to-day path (debug + light XFN) without dropping the full lifecycle when the job is a product feature.

Review stance: product/architecture assessment of the kit as shipped on main, not a PR diff review. Evidence from AGENTS.md, README.md, docs/edd.md, kit/src/edd/run_evals.ts, MCP/install paths, skill line counts, and post–Aug 30 merges (#23–#43).

Markdown source