Building Aruo: the project decision journal
Updated
Template selection: kind before ecosystem
Problem
aruo create's catalog grew to 8 templates (3 app frameworks, 5
libraries) across 5 languages. The interactive picker showed all 8 in one
flat, grouped-by-kind list; a user had to scan every ecosystem to find
the one relevant to them, even though "am I building an app or a
library" is almost always known before "which language."
Decision
Add a screen before the template list: What are you building? with two
options, Application and Library. Picking one filters the next screen
down to that kind only (3 apps or 5 libraries today). The screen is
skipped automatically when --kind, --template, or a non-interactive
session already answers the question, or if the catalog ever shrinks to
one kind.
Catalog-derived labels
Each kind option's helper text (Next.js application, Nuxt application, React application / Go library, JavaScript library, ...) is generated
from the catalog's actual entry names at prompt-build time
(kindEntryNames), not hand-written prose. A hand-written description
("frontend frameworks like React and Vue") would silently go stale the
moment a template is added or removed. Deriving it from the same catalog
data the picker itself reads means it's structurally impossible for the
description to list a template that doesn't exist, or omit one that does.
Correctness fix folded in during the later back-navigation rework
When this screen became reactive (part of Prompter.Guide, see the
companion architecture note), an edge case surfaced: if --language
narrows the catalog and a kind is chosen interactively, it's possible
for a kind to have zero matching templates for that language (no
python-app exists, for instance). Originally the kind list was built from
the full catalog, so a user could pick a kind that led to an empty
template screen. Fixed by computing the kind list from the
--language-filtered entry set, not the full catalog; a kind only
appears as a choice if it has at least one template under the
active language filter.
Limitations / open questions
- Only two kinds exist today (
app,library); the picker's copy ("What are you building?") and two-column layout weren't tested against a third kind.resolveKind's generic n-kind handling should still work, but the wording was chosen for a binary choice specifically.
Real defaults instead of TODO placeholders
Problem
aruo create's optional --description/--author fields defaulted to
static placeholder strings; Copyright (c) TODO: set an author landed
literally in generated LICENSE files. This reads as broken output, not a
convenience: "automatic" defaulting that still requires a manual
find-and-replace afterward isn't automatic; it relocates
the manual step from the CLI prompt into the generated repository, where
it's easier to miss.
Decision
Derive values from known data:
--authorleft blank → best-effortgit config --get user.name(500ms timeout, falls back to empty on any failure; missing git, unset config, timeout).--descriptionleft blank → the catalog entry's ownDescriptionfield ("A production-ready Go library with zero external dependencies, tests, CI, governance, security, and documentation"), since every template already carries an accurate one-line description for its own catalog listing.
Failure behavior
--module explicitly does not get this treatment; it becomes a literal
Go import path / npm package name / PyPI name, and a wrong guess produces
a broken go.mod/package.json/pyproject.toml, while an omitted
description only leaves incomplete metadata. Only auto-default a field when a wrong (or
empty) default is merely incomplete, never when it's broken.
For the fields that do get a default, when detectGitAuthor fails (no
git, no config, timeout), the fallback is an empty string; Copyright (c) with nothing after, not a second-tier placeholder. An incomplete
field is preferable to fabricated content because the omission is visible.
The contract is a derived value or an empty value, never a placeholder in
the prompt or generated output. This became the rule
documented in the CLI's copy style guide.
Limitations / open questions
detectGitAuthor's 500ms timeout was chosen without measuringgit configlatency across slow/networked filesystems. The value is an unmeasured timeout chosen to avoid stalling the prompt flow.- No equivalent auto-derivation exists yet for
--license(still a flag with a per-template default, never a prompt); not evaluated whether the same real-value-or-empty rule should extend there.
Defaults should not occupy editable input
The bug
internal/tux/charm/prompter.go's Input() seeded the Huh field's bound
value directly from the request's default:
value := ""
if request.Default != nil {
value = *request.Default
}
field := huh.NewInput().Value(&value)...
For an optional field like --author, this meant the prompt showed
Author or organization (Optional): with the input box already containing
Aruodore (from git config user.name), cursor at the end. To
leave it blank, the user had to backspace all four characters out; not
what "optional, leave blank to skip" implies.
Root cause of the asymmetry
The plain/accessible adapter (internal/tux/plain) never had this
problem, because it was never structured this way in the first place: its
Label [Default] convention is purely informational text printed in the
prompt line, never something sitting in an editable buffer the user has
to clear. The two adapters had quietly drifted to implement the same
InputRequest.Default contract two different ways.
Fix
Match the plain adapter's contract exactly: the field starts
empty; the default is shown only via Placeholder(...) (grayed-out hint
text that a keystroke replaces, never counted as real input); after the
form returns, if the submitted value is empty and a default exists, the
default is substituted then; outside the widget, after the fact.
value := ""
placeholder := request.Placeholder
if request.Default != nil && *request.Default != "" {
placeholder = *request.Default
}
field := huh.NewInput().Placeholder(placeholder).Value(&value)...
// ... after p.run(ctx, field):
if value == "" && request.Default != nil {
value = *request.Default
}
This same contract carried forward into Prompter.Guide's reactive input
steps later (see the back-navigation architecture note); the collected
Answers assembly applies the identical "substitute only after empty
submission" logic once, after the whole multi-group form completes.
Verified
Live pty check in --accessible mode confirmed the prompt shows
Author or organization (Optional) [Aruodore]: with an empty, untouched
input box, and pressing Enter alone still correctly produces Copyright (c) Aruodore in the generated LICENSE.
Limitations / open questions
- Only checked Huh v2.0.3's
Input/Select/Confirmfields for this same class of bug; didn't audit whetherMultiSelect's default-selection handling has an analogous "pre-selected vs. hint" distinction worth double-checking, since Aruo doesn't useMultiSelectin any shipped command yet.
Dogfooding exposed doctor detection gaps
The bug
aruo doctor scores a repository's test coverage signal two ways:
whether CI runs a recognized test command, and whether test files exist
under a recognized naming convention. checkTests' native-test-runner
allowlist had go test, pytest, cargo test, and
npm/pnpm/bun/deno test; but not node --test or unittest.
TestFiles() only matched Python's _test.py suffix convention, not
the more common test_*.py prefix convention, which is what
python-library's generated tests use.
Net effect: a freshly aruo create-generated js-library project lost 5
of 15 points on the tests category despite its CI running the real node --test command; a generated python-library project lost 8 points on
top of that for using test_*.py, the more idiomatic Python convention,
instead of the suffix form doctor was looking for.
How it was found
Not by auditing checkTests/TestFiles and reasoning about what
conventions they cover. Found by running aruo doctor against a
project generated by aruo create --template js-library (and separately
python-library) and noticing the reported score was lower than the
equivalent, equally-complete go-library project's score, for no
substantive difference in project quality. The undercounting was only
visible by producing an artifact and pointing the tool at itself ;
reading the detection code by itself wouldn't have flagged "this list is
incomplete" without something to compare it against.
Fix
Extended checkTests' containsAny list with "unittest" and "node --test"; extended TestFiles() to also recognize the test_*.py prefix
alongside the existing suffix/substring checks. Added
TestGeneratedJSLibraryScoresA and TestGeneratedPythonLibraryScoresA
alongside the existing Go equivalent specifically so a future template
addition can't silently regress this same class of gap again without a
test catching it.
Limitations / open questions
- This was found opportunistically while building unrelated template
catalog entries (adding
js-library,python-library), not from a systematic audit of every ecosystem's test-runner naming conventions doctor should recognize. Other gaps of the same shape likely exist for ecosystems Aruo doesn't have a template for yet (Rust'scargo nextest, Java'smvn test/gradle test, etc.) and would need the same dogfooding process; generate a real project, run doctor against it, compare the score to expectation; to surface.
Verify templates with real current tools
Decision
For every new ecosystem template added to aruo create's catalog
(TypeScript library, React app, Nuxt app, Vue library, Next.js app,
Python library), the process was: run the current official
scaffolding tool first (npm create vite@latest, npx nuxi@latest init,
npm create vue@latest, npx create-next-app@latest) to see today's
actual output; not the shape remembered from training data; then build
a lean, Aruo-specific version by hand, then verify the whole generated
project with npm install, its test runner, and its build command.
Aruo's own Go test suite still stays hermetic; it only checks the file
plan exists (TestXHasRequiredFiles), never runs npm install itself,
so go test ./... has no network dependency. The real install/build/test
verification happened by hand, once, while writing each template, on the
actual aruo create-generated output.
What this caught that memory wouldn't have
vite.config.tsneedsdefineConfigimported from"vitest/config", not"vite", or thetestoption silently doesn't type-check under strict TypeScript. Not a runtime failure; a type error that only shows up runningtscfor real.- jsdom 30 fails to load under Node 20 after emitting an engine warning. This session's sandbox defaulted to Node 20.20.2; jsdom 30 requires Node ≥22.22.2/24.15.0/26.0.0. Confirmed by running it and reading the failure, then downloading a Node 26.7.0 tarball to run the verification instead of assuming compatibility from the package's stated engines field.
@nuxt/test-utils'smountSuspendedneeds@vue/test-utilsas an undocumented peer dependency. Fails with an unhelpful resolve error without it; not listed as a required dependency anywhere obvious. Also needsenvironment: "nuxt"set explicitly invitest.config.ts, which in turn needshappy-dominstalled specifically (notjsdom) or it fails with "Could not resolve happy-dom." None of this is discoverable without running it.- A hand-written Vue template lost its literal "Hello, "/"!" text when
the
.vue.tmplfile was rendered and run throughnpm test; a bug in the template's Go-template escaping for embedding literal Vue mustache syntax ({{ "{{ name }}" }}), caught only because the generated project's own test failed, not because Aruo's own file-existence-only unit test could have caught a content bug like this. next-app'svitest.config.tsproduced a real warning without"type": "module"inpackage.json; again, only visible by running it.
Ongoing verification rule
Every one of these is the kind of detail that "I know how Vite/Nuxt/Next scaffolding generally works" gets subtly wrong, because ecosystem tooling changes versions constantly and these specific failure modes are undocumented implementation details, not documented API contracts. None of them would have been caught by writing the template from memory and only checking it compiles; they only surface when the generated project's own real toolchain runs against it.
Limitations / open questions
- This verification was done once, by hand, at the time each template was
written; there's no CI job that re-runs a real
npm installagainst the generated templates on a schedule, so a future breaking change in any of these tools (a new Nuxt major, a Vite config format change) wouldn't be caught automatically; it would need to be caught the same way, by hand, if and when someone re-verifies. - Real installs against the live npm registry are inherently non-reproducible in the exact-version sense; the specific versions verified (jsdom 30, typescript 7.0.2, @types/node 26.1.2, etc.) will drift out of date; the verification method is the lasting part, not the specific version numbers recorded in commit messages at the time.
Back-navigation uses two mechanisms
The ask
"Why can't I go back in this tool?"; aruo create walks through up to 7
screens (name, kind, template, module, description, author, confirm),
each issued as its own independent Prompter.Input/Select/Confirm
call. Wanted: backward navigation across the whole flow.
Root cause of "can't"
Every screen built its own isolated huh.NewForm(huh.NewGroup(field)).
Huh's Prev (shift+tab) key is tested and wired by default; but
it only does anything within one continuous multi-group form. A lone
single-field form is structurally always "the first field of the first
group," so Prev is permanently disabled by Huh's own FieldPosition.IsFirst()
check. The plain/accessible adapter is a hand-rolled forward-only
bufio.Scanner loop with no concept of "back" at all.
Options considered
A. One continuous multi-group Huh form, using WithHideFunc for
steps a flag already answered and OptionsFunc/TitleFunc for content
that depends on an earlier answer (the template list narrowing to the
chosen kind). Gets native shift+tab for free, including automatic help-bar
advertising of the key. Cost: glue code to make Huh's reactive
*Func mechanism work across steps, and a genuine gotcha found along the
way (see the companion bug note on OptionsFunc/empty bindings).
B. Keep each screen as its own isolated form, and instead make each
individual prompt call able to signal "the user wants to go back" (a
sentinel error), with a shared orchestrator loop managing an index and
re-invoking the previous step's prompt call. Simpler reactivity (just
ordinary Go closures re-run each time, no OptionsFunc needed); but
there's no native "back" signal to hook into for a structurally-isolated
single-field form (per the root cause above), so this would mean either
fighting Huh's keymap-recomputation lifecycle to force-enable a key it
actively disables, or writing a custom Bubble Tea program from scratch ;
arguably more complexity than option A, just moved to a different spot.
Chose A for the rich adapter: it builds on Huh's own tested, native,
documented mechanism (TestPrevGroup, TestHideGroup exist in Huh's own
suite) rather than fighting the library.
What the plain adapter does instead
No Huh involved at all; it's Aruo's own code. Added an index-based loop
over []tux.Step that recognizes the literal word back (trimmed,
case-insensitive) as a navigation command before running each field's own
validate/required checks. Chose a bare reserved word over a sigil-prefixed
one (:back, !back) for maximum discoverability, accepting the
collision cost: someone who wants the literal text "back" as a
project name/description/author has to use the corresponding flag instead
of the interactive prompt. Documented, not hidden.
The shared interface
Both adapters implement one new Prompter.Guide(ctx, []Step) (Answers, error) method. Step.Skip ended up typed as func() bool rather than
func(Answers) bool; every skip condition in this catalog turned out to
be a static flag decided before the guide even starts (--kind,
--template, --yes, single-kind catalog), never something that changes
based on an answer gathered mid-flow. Keeping it non-reactive is a real
simplification (no need to seed bound variables from flags before a
hidden group's field ever runs) that would need revisiting the day a
step's skip condition depends on an earlier interactive answer.
Known, accepted quirk
If a user goes back and changes kind after already picking a template,
Huh's Select.selectOption() tries to preserve the old selection by
value first; if the old template ID isn't in the new (re-filtered)
option list, it falls back to clamping the cursor index into the new
list rather than resetting to the first option. Not silent; the user
sees whichever option that lands on highlighted before they press Enter ;
but not necessarily a sensible default either. Documented in the terminal
UX spec rather than worked around, since fixing it would mean depending
on undocumented Huh internals for a low-severity cosmetic issue.
Limitations / open questions
SecretandMultiSelectstep kinds don't exist yet; nothing increate's flow needs them today. Would needStepSecret/StepMultiSelectadded totux.StepKindif a future command's guided flow needs either.- No other command uses
Prompterinteractively yet, so the blast radius of this interface change was contained to one command's wiring; unclear how well theSkip func() bool(non-reactive) simplification holds up once a second interactive command exists.
Huh empty-bindings hash collision
Context
Building multi-screen back-navigation for Aruo's aruo create command
(charm.land/huh v2.0.3, bubble-tea v2.0.2). The rich prompter now builds
one continuous multi-group huh.Form instead of one isolated form per
screen, so Huh's native shift+tab works across the whole flow. Steps
whose content depends on an earlier answer (the template list narrowing
to the chosen project kind) use Select.OptionsFunc(f func() []Option[T], bindings any); the field re-evaluates f when hash(bindings) changes.
Every step, including the very first one with no preceding steps to
depend on, was wired the same way for uniformity: build a []any of the
preceding steps' bound pointers as bindings. For step 0 that's an empty
slice, []any{}.
The bug
The very first step's OptionsFunc/TitleFunc never fired. Down/Up
arrow keys on the first Select had no visible effect; the field silently
kept its zero-value state as if Options had never been set.
Root cause
Huh's Eval[T] (eval.go) decides whether to recompute via:
func (e *Eval[T]) shouldUpdate() (bool, uint64) {
if e.fn == nil {
return false, 0
}
newHash := hash(e.bindings)
return e.bindingsHash != newHash, newHash
}
bindingsHash is a uint64 that starts at its Go zero value, 0.
hash() wraps hashstructure.Hash(val, hashstructure.FormatV2, nil).
Confirmed directly:
hashstructure.Hash([]any{}, hashstructure.FormatV2, nil) // → 0, nil
An empty slice hashes to exactly 0; the same as the field's untouched
bindingsHash. shouldUpdate() compares 0 != 0, gets false, and the
recompute never happens. Every subsequent step in the form was fine,
because their bindings included at least one prior step's pointer,
producing a non-zero hash on first evaluation.
Huh's public API and documentation do not mention this behavior. It shows
up by noticing that a dynamically-configured field (OptionsFunc,
TitleFunc) with a data-free binding silently behaves as if it were
never configured at all, while the identical field wired with .Options()
(the static setter) works.
Fix
Seed bindings with a nonempty value; the step's
own ID is enough, since it's guaranteed present and stable per step:
bindings := make([]any, 0, index+1)
bindings = append(bindings, step.ID)
for i := 0; i < index; i++ {
bindings = append(bindings, fields[i].pointer)
}
How it was found
Reading the Huh source didn't surface it; the mechanism looks
correct on paper. Found by writing a throwaway test that drives a
huh.NewSelect[...]().OptionsFunc(...) field directly through Form.Init()
- simulated keypresses (see the companion note on testing Huh forms
without a PTY) and printing whether the closure itself ever ran. It
didn't, for the field with an empty
bindingsslice; it did, immediately, for an otherwise-identical field with a non-empty one. Isolating to that single variable took three progressively smaller repro tests.
Limitations / open questions
- Only verified against Huh v2.0.3 and
hashstructure/v2v2.0.2; not confirmed whether this is fixed in a later Huh release, or whether it's considered a bug upstream at all. - Didn't check whether other zero-hashing
bindingsvalues exist besides an empty slice (e.g. a nil pointer, a zero-value struct); the fix here sidesteps the whole class by never passing an empty/trivial binding, but a more general fix would need to know exactly which values hash to 0.
Testing Huh forms without a PTY
Problem
Aruo's rich terminal prompter is built on Huh v2 / Bubble Tea v2. Earlier
work in the same project had already established that this sandbox can't
run a real PTY through Bubble Tea's Program loop end to end; the rich
TUI sends terminal capability queries (sync-output mode probe, Kitty
keyboard protocol probe) that a basic pty.fork()-based test harness
can't answer, so it just hangs. That's why the project's existing charm
prompter tests exercise Huh's own WithAccessible(true) mode instead; a
built-in scriptable-stdin fallback.
That fallback stopped being usable for the specific feature being built:
multi-group forms with WithHideFunc (conditionally hidden groups) and
OptionsFunc/TitleFunc (reactive content). Reading Huh's source
directly confirmed Form.runAccessible is a plain nested loop over every
group and field, unconditionally; it does not check group.hide at all,
and each field's RunAccessible reads only the field's static .val,
never invoking .fn. So testing the real navigation and reactivity
mechanics needed something else.
Direct model updates
huh.Form satisfies tea.Model (via a package-level alias,
type Model = compat.Model, itself ultimately tea.Model). Update
is just an ordinary Go method: func (f *Form) Update(msg tea.Msg) (Model, tea.Cmd). Nothing about calling it requires a running tea.Program,
a terminal, or any I/O at all; it's pure state transition logic that
happens to normally be driven by Bubble Tea's runtime loop.
So: build the form, call Init(), then call Update(tea.KeyPressMsg{...})
directly with real key values (tea.KeyDown, tea.KeyEnter,
tea.Key{Code: tea.KeyTab, Mod: tea.ModShift} for shift+tab), reading
Key.String() output to confirm each constructs the exact string Huh's
key.Matches compares against ("down", "enter", "shift+tab").
Draining commands
A single Update call frequently returns a non-nil tea.Cmd; a
func() tea.Msg that would normally be handed back to Bubble Tea's
runtime, executed asynchronously, and its result fed back into another
Update call. Some of these carry real state (the message produced by an
OptionsFunc recompute); some are internal batching wrappers
(tea.BatchMsg, and Huh's own unexported sequenceMsg, both literally
[]tea.Cmd under a different name); some are time-based
(spinner.Tick, which wraps time.After).
A test has to replicate enough of that loop by hand:
func drainGuideForm(t *testing.T, model huh.Model, cmd tea.Cmd) huh.Model {
pending := []tea.Cmd{cmd}
for len(pending) > 0 {
next := pending[0]
pending = pending[1:]
if next == nil {
continue
}
msg, ok := callWithTimeout(next) // see below
if !ok || msg == nil {
continue
}
if more, ok := asCmdSlice(msg); ok { // reflection, see below
pending = append(pending, more...)
continue
}
var newCmd tea.Cmd
model, newCmd = model.Update(msg)
if newCmd != nil {
pending = append(pending, newCmd)
}
}
return model
}
Two details were required:
- Unwrapping batch/sequence messages via reflection, not type
assertion.
sequenceMsgis unexported; a test outside thehuhpackage can't writemsg.(huh.sequenceMsg). But bothtea.BatchMsgandsequenceMsgare just named[]tea.Cmd.reflect.ValueOf(msg).Kind() == reflect.Slice, iterate.Index(i).Interface().(tea.Cmd), and both shapes unwrap identically without ever needing the unexported type name. - Bounding each
cmd()call with a short timeout instead of calling it inline. Callingspinner.Tick(or anything wrappingtea.Tick) directly blocks on a real timer. Running each cmd in its own goroutine with aselectagainsttime.After(20ms)lets real, fast, synchronous cmds (the actual data-carrying ones) resolve normally, while anything that blocks longer than that gets abandoned; assumed to be a cosmetic animation command irrelevant to the state under test. The abandoned goroutine can leak past the deadline; acceptable for a short-lived test process.
With both of those in place, a test can type-simulate an entire
interaction; select an option, submit, shift+tab back, change the
earlier answer, move forward again; and assert on the form's bound Go
values directly, with zero PTY involvement, fully deterministic, and fast
(each of the three tests built this way ran in under 0.2s).
Where this technique came from
Not discovered by reading Huh's docs; its README doesn't mention testing
at all. Found by locating Huh's own test suite (huh_test.go) and
noticing it drives navigation via direct message injection into Form.Update
too, just using Huh's unexported message constructors internally
(f.Update(prevGroup()) etc., since the test lives inside package huh).
The adaptation for external test code was working out the reflection-based
generic unwrap and the timeout-bounded cmd execution, neither of which
Huh's own tests needed since they never send commands that produce further
async cmds requiring drainage across a whole form assembled from scratch
outside the package.
Limitations / open questions
- This proves the logic is correct (real keystrokes, real Huh navigation/reactivity code paths) but not the actual terminal rendering or the capability-query handshake that made PTY testing impractical in the first place; those remain unverified in this sandbox.
- The 20ms cmd timeout is a heuristic tuned against this specific form's
cmd shapes (title/description/options recompute, focus, spinner tick).
A form with a slow synchronous
OptionsFunc(e.g. a real network call) would need a longer timeout or a different signal to distinguish "real but slow" from "cosmetic and endless."