Files
sow-tools/.agents/skills/data-driven-design/SKILL.md
T
archvillainetteandClaude Fable 5 b5af42b53b
build-binaries / build-binaries (pull_request) Has been cancelled
test / test (pull_request) Has been cancelled
skills: add data-driven-design
Pattern skill for keeping specific knowledge (dataset names, column
encodings, layout) in data and only general mechanisms in code — the
direction sow-tools must move to be open-sourceable. Grounded in the
value_encodings mode zoo and the skills-dataset rename breakage;
baseline-tested against a subagent that reproduced both failures.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-19 09:35:42 +02:00

56 lines
3.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: data-driven-design
description: Use when adding a mode, switch case, config entry, special-cased name, or per-column/per-dataset handler to a tool that processes authored data — or when a tool "knows" a project's layout, dataset names, key prefixes, or format quirks and could not be open sourced as-is.
---
# Data-Driven Design
The tool implements **general mechanisms**; the data declares **all specific knowledge**. Every time code names a thing that lives in data — a dataset, a column, a format quirk — the tool stops being a tool and becomes a private extension of one project's layout.
**The test:** could this repo be open sourced and build a *different* project's data without code changes? Every hardcoded name, mode, and special case is a "no".
## The two smells
**1. The mode zoo.** A new data shape arrives and you add a named mode plus a switch case:
```yaml
- {column: AvailableHeadsMale, mode: external_hex_list, hex_width: 3}
- {column: PreferredAlignments, mode: alignment_set_mask, hex_width: 3}
- {column: ProficiencyFeats, mode: id_hex_list, hex_width: 4}
```
Three "modes" are one concept — *a set of ids with an engine-side encoding* — wearing three names. Each mode is code the data could have been. The fix is a smaller vocabulary of orthogonal primitives declared where the column is defined:
```json
"EquipSlots": {"shape": "id_set", "encode": "bitmask", "max": 31}
```
One parser for `id_set`, one encoder per encoding. Adding the next column is a data edit, zero code. The mode zoo *shrinks*: N special cases collapse into shape × encoding.
**2. Names in code.** `datasetRows("skills")`, `case "skills":`, `TrimPrefix(key, "skills:")` — the tool assumes the project's layout. When data moves (`skills``skills/core`), the tool breaks even though the data is self-consistent. Names must flow *from* the data: the spec that references a dataset names it; a key's namespace is whatever precedes its own colon. If code needs a name, some piece of data already knows it — read it from there.
## Deciding
Ask of every constant, case, and config knob: **whose knowledge is this?**
- *The mechanism's* (how to parse JSON, allocate rows, write a table) → code.
- *The data's* (which columns exist, how the engine encodes them, what things are called) → data, even if code is faster to write today.
Generalize only when it deletes: a primitive that replaces N modes is lazy; a framework "for future shapes" is not. One new shape may take its case *if* the case is spelled as a reusable primitive the next shape can declare.
## Rationalizations
| Excuse | Reality |
|---|---|
| "Just follow the existing pattern, it's proven" | The pattern *is* the defect. Each repetition raises the cost of the real fix. |
| "N cases isn't a crisis yet" | The crisis is per-case: every one couples the tool to one project forever. |
| "Release is tonight, generalize later" | Later never comes; tonight's mode is tomorrow's baseline. Declaring shape in data is usually the *same* hour of work. |
| "This name never changes" | Today's rename broke exactly such a name. Data moves; mechanisms shouldn't care. |
## Red flags — stop and re-shape
- A config entry that names both a dataset and a column to select behavior
- A new value in a `mode`/`kind`/`type` string enum with its own switch arm
- A literal in code that also appears in the data tree (`"skills"`, `"skills:"`)
- Per-project constants (row ranges, hex widths, id ceilings) living in Go instead of the dataset that owns them