Building Aruo: the project decision journal

By

Updated


Template selection: kind before ecosystem

Problem

aruo create's catalog grew to 8 templates (3 app frameworks, 5 libraries) across 5 languages. The interactive picker showed all 8 in one flat, grouped-by-kind list; a user had to scan every ecosystem to find the one relevant to them, even though "am I building an app or a library" is almost always known before "which language."

Decision

Add a screen before the template list: What are you building? with two options, Application and Library. Picking one filters the next screen down to that kind only (3 apps or 5 libraries today). The screen is skipped automatically when --kind, --template, or a non-interactive session already answers the question, or if the catalog ever shrinks to one kind.

Catalog-derived labels

Each kind option's helper text (Next.js application, Nuxt application, React application / Go library, JavaScript library, ...) is generated from the catalog's actual entry names at prompt-build time (kindEntryNames), not hand-written prose. A hand-written description ("frontend frameworks like React and Vue") would silently go stale the moment a template is added or removed. Deriving it from the same catalog data the picker itself reads means it's structurally impossible for the description to list a template that doesn't exist, or omit one that does.

Correctness fix folded in during the later back-navigation rework

When this screen became reactive (part of Prompter.Guide, see the companion architecture note), an edge case surfaced: if --language narrows the catalog and a kind is chosen interactively, it's possible for a kind to have zero matching templates for that language (no python-app exists, for instance). Originally the kind list was built from the full catalog, so a user could pick a kind that led to an empty template screen. Fixed by computing the kind list from the --language-filtered entry set, not the full catalog; a kind only appears as a choice if it has at least one template under the active language filter.

Limitations / open questions

  • Only two kinds exist today (app, library); the picker's copy ("What are you building?") and two-column layout weren't tested against a third kind. resolveKind's generic n-kind handling should still work, but the wording was chosen for a binary choice specifically.

Real defaults instead of TODO placeholders

Problem

aruo create's optional --description/--author fields defaulted to static placeholder strings; Copyright (c) TODO: set an author landed literally in generated LICENSE files. This reads as broken output, not a convenience: "automatic" defaulting that still requires a manual find-and-replace afterward isn't automatic; it relocates the manual step from the CLI prompt into the generated repository, where it's easier to miss.

Decision

Derive values from known data:

  • --author left blank → best-effort git config --get user.name (500ms timeout, falls back to empty on any failure; missing git, unset config, timeout).
  • --description left blank → the catalog entry's own Description field ("A production-ready Go library with zero external dependencies, tests, CI, governance, security, and documentation"), since every template already carries an accurate one-line description for its own catalog listing.

Failure behavior

--module explicitly does not get this treatment; it becomes a literal Go import path / npm package name / PyPI name, and a wrong guess produces a broken go.mod/package.json/pyproject.toml, while an omitted description only leaves incomplete metadata. Only auto-default a field when a wrong (or empty) default is merely incomplete, never when it's broken.

For the fields that do get a default, when detectGitAuthor fails (no git, no config, timeout), the fallback is an empty string; Copyright (c) with nothing after, not a second-tier placeholder. An incomplete field is preferable to fabricated content because the omission is visible. The contract is a derived value or an empty value, never a placeholder in the prompt or generated output. This became the rule documented in the CLI's copy style guide.

Limitations / open questions

  • detectGitAuthor's 500ms timeout was chosen without measuring git config latency across slow/networked filesystems. The value is an unmeasured timeout chosen to avoid stalling the prompt flow.
  • No equivalent auto-derivation exists yet for --license (still a flag with a per-template default, never a prompt); not evaluated whether the same real-value-or-empty rule should extend there.

Defaults should not occupy editable input

The bug

internal/tux/charm/prompter.go's Input() seeded the Huh field's bound value directly from the request's default:

value := ""
if request.Default != nil {
    value = *request.Default
}
field := huh.NewInput().Value(&value)...

For an optional field like --author, this meant the prompt showed Author or organization (Optional): with the input box already containing Aruodore (from git config user.name), cursor at the end. To leave it blank, the user had to backspace all four characters out; not what "optional, leave blank to skip" implies.

Root cause of the asymmetry

The plain/accessible adapter (internal/tux/plain) never had this problem, because it was never structured this way in the first place: its Label [Default] convention is purely informational text printed in the prompt line, never something sitting in an editable buffer the user has to clear. The two adapters had quietly drifted to implement the same InputRequest.Default contract two different ways.

Fix

Match the plain adapter's contract exactly: the field starts empty; the default is shown only via Placeholder(...) (grayed-out hint text that a keystroke replaces, never counted as real input); after the form returns, if the submitted value is empty and a default exists, the default is substituted then; outside the widget, after the fact.

value := ""
placeholder := request.Placeholder
if request.Default != nil && *request.Default != "" {
    placeholder = *request.Default
}
field := huh.NewInput().Placeholder(placeholder).Value(&value)...
// ... after p.run(ctx, field):
if value == "" && request.Default != nil {
    value = *request.Default
}

This same contract carried forward into Prompter.Guide's reactive input steps later (see the back-navigation architecture note); the collected Answers assembly applies the identical "substitute only after empty submission" logic once, after the whole multi-group form completes.

Verified

Live pty check in --accessible mode confirmed the prompt shows Author or organization (Optional) [Aruodore]: with an empty, untouched input box, and pressing Enter alone still correctly produces Copyright (c) Aruodore in the generated LICENSE.

Limitations / open questions

  • Only checked Huh v2.0.3's Input/Select/Confirm fields for this same class of bug; didn't audit whether MultiSelect's default-selection handling has an analogous "pre-selected vs. hint" distinction worth double-checking, since Aruo doesn't use MultiSelect in any shipped command yet.

Dogfooding exposed doctor detection gaps

The bug

aruo doctor scores a repository's test coverage signal two ways: whether CI runs a recognized test command, and whether test files exist under a recognized naming convention. checkTests' native-test-runner allowlist had go test, pytest, cargo test, and npm/pnpm/bun/deno test; but not node --test or unittest. TestFiles() only matched Python's _test.py suffix convention, not the more common test_*.py prefix convention, which is what python-library's generated tests use.

Net effect: a freshly aruo create-generated js-library project lost 5 of 15 points on the tests category despite its CI running the real node --test command; a generated python-library project lost 8 points on top of that for using test_*.py, the more idiomatic Python convention, instead of the suffix form doctor was looking for.

How it was found

Not by auditing checkTests/TestFiles and reasoning about what conventions they cover. Found by running aruo doctor against a project generated by aruo create --template js-library (and separately python-library) and noticing the reported score was lower than the equivalent, equally-complete go-library project's score, for no substantive difference in project quality. The undercounting was only visible by producing an artifact and pointing the tool at itself ; reading the detection code by itself wouldn't have flagged "this list is incomplete" without something to compare it against.

Fix

Extended checkTests' containsAny list with "unittest" and "node --test"; extended TestFiles() to also recognize the test_*.py prefix alongside the existing suffix/substring checks. Added TestGeneratedJSLibraryScoresA and TestGeneratedPythonLibraryScoresA alongside the existing Go equivalent specifically so a future template addition can't silently regress this same class of gap again without a test catching it.

Limitations / open questions

  • This was found opportunistically while building unrelated template catalog entries (adding js-library, python-library), not from a systematic audit of every ecosystem's test-runner naming conventions doctor should recognize. Other gaps of the same shape likely exist for ecosystems Aruo doesn't have a template for yet (Rust's cargo nextest, Java's mvn test/gradle test, etc.) and would need the same dogfooding process; generate a real project, run doctor against it, compare the score to expectation; to surface.

Verify templates with real current tools

Decision

For every new ecosystem template added to aruo create's catalog (TypeScript library, React app, Nuxt app, Vue library, Next.js app, Python library), the process was: run the current official scaffolding tool first (npm create vite@latest, npx nuxi@latest init, npm create vue@latest, npx create-next-app@latest) to see today's actual output; not the shape remembered from training data; then build a lean, Aruo-specific version by hand, then verify the whole generated project with npm install, its test runner, and its build command.

Aruo's own Go test suite still stays hermetic; it only checks the file plan exists (TestXHasRequiredFiles), never runs npm install itself, so go test ./... has no network dependency. The real install/build/test verification happened by hand, once, while writing each template, on the actual aruo create-generated output.

What this caught that memory wouldn't have

  • vite.config.ts needs defineConfig imported from "vitest/config", not "vite", or the test option silently doesn't type-check under strict TypeScript. Not a runtime failure; a type error that only shows up running tsc for real.
  • jsdom 30 fails to load under Node 20 after emitting an engine warning. This session's sandbox defaulted to Node 20.20.2; jsdom 30 requires Node ≥22.22.2/24.15.0/26.0.0. Confirmed by running it and reading the failure, then downloading a Node 26.7.0 tarball to run the verification instead of assuming compatibility from the package's stated engines field.
  • @nuxt/test-utils's mountSuspended needs @vue/test-utils as an undocumented peer dependency. Fails with an unhelpful resolve error without it; not listed as a required dependency anywhere obvious. Also needs environment: "nuxt" set explicitly in vitest.config.ts, which in turn needs happy-dom installed specifically (not jsdom) or it fails with "Could not resolve happy-dom." None of this is discoverable without running it.
  • A hand-written Vue template lost its literal "Hello, "/"!" text when the .vue.tmpl file was rendered and run through npm test; a bug in the template's Go-template escaping for embedding literal Vue mustache syntax ({{ "{{ name }}" }}), caught only because the generated project's own test failed, not because Aruo's own file-existence-only unit test could have caught a content bug like this.
  • next-app's vitest.config.ts produced a real warning without "type": "module" in package.json; again, only visible by running it.

Ongoing verification rule

Every one of these is the kind of detail that "I know how Vite/Nuxt/Next scaffolding generally works" gets subtly wrong, because ecosystem tooling changes versions constantly and these specific failure modes are undocumented implementation details, not documented API contracts. None of them would have been caught by writing the template from memory and only checking it compiles; they only surface when the generated project's own real toolchain runs against it.

Limitations / open questions

  • This verification was done once, by hand, at the time each template was written; there's no CI job that re-runs a real npm install against the generated templates on a schedule, so a future breaking change in any of these tools (a new Nuxt major, a Vite config format change) wouldn't be caught automatically; it would need to be caught the same way, by hand, if and when someone re-verifies.
  • Real installs against the live npm registry are inherently non-reproducible in the exact-version sense; the specific versions verified (jsdom 30, typescript 7.0.2, @types/node 26.1.2, etc.) will drift out of date; the verification method is the lasting part, not the specific version numbers recorded in commit messages at the time.

Back-navigation uses two mechanisms

The ask

"Why can't I go back in this tool?"; aruo create walks through up to 7 screens (name, kind, template, module, description, author, confirm), each issued as its own independent Prompter.Input/Select/Confirm call. Wanted: backward navigation across the whole flow.

Root cause of "can't"

Every screen built its own isolated huh.NewForm(huh.NewGroup(field)). Huh's Prev (shift+tab) key is tested and wired by default; but it only does anything within one continuous multi-group form. A lone single-field form is structurally always "the first field of the first group," so Prev is permanently disabled by Huh's own FieldPosition.IsFirst() check. The plain/accessible adapter is a hand-rolled forward-only bufio.Scanner loop with no concept of "back" at all.

Options considered

A. One continuous multi-group Huh form, using WithHideFunc for steps a flag already answered and OptionsFunc/TitleFunc for content that depends on an earlier answer (the template list narrowing to the chosen kind). Gets native shift+tab for free, including automatic help-bar advertising of the key. Cost: glue code to make Huh's reactive *Func mechanism work across steps, and a genuine gotcha found along the way (see the companion bug note on OptionsFunc/empty bindings).

B. Keep each screen as its own isolated form, and instead make each individual prompt call able to signal "the user wants to go back" (a sentinel error), with a shared orchestrator loop managing an index and re-invoking the previous step's prompt call. Simpler reactivity (just ordinary Go closures re-run each time, no OptionsFunc needed); but there's no native "back" signal to hook into for a structurally-isolated single-field form (per the root cause above), so this would mean either fighting Huh's keymap-recomputation lifecycle to force-enable a key it actively disables, or writing a custom Bubble Tea program from scratch ; arguably more complexity than option A, just moved to a different spot.

Chose A for the rich adapter: it builds on Huh's own tested, native, documented mechanism (TestPrevGroup, TestHideGroup exist in Huh's own suite) rather than fighting the library.

What the plain adapter does instead

No Huh involved at all; it's Aruo's own code. Added an index-based loop over []tux.Step that recognizes the literal word back (trimmed, case-insensitive) as a navigation command before running each field's own validate/required checks. Chose a bare reserved word over a sigil-prefixed one (:back, !back) for maximum discoverability, accepting the collision cost: someone who wants the literal text "back" as a project name/description/author has to use the corresponding flag instead of the interactive prompt. Documented, not hidden.

The shared interface

Both adapters implement one new Prompter.Guide(ctx, []Step) (Answers, error) method. Step.Skip ended up typed as func() bool rather than func(Answers) bool; every skip condition in this catalog turned out to be a static flag decided before the guide even starts (--kind, --template, --yes, single-kind catalog), never something that changes based on an answer gathered mid-flow. Keeping it non-reactive is a real simplification (no need to seed bound variables from flags before a hidden group's field ever runs) that would need revisiting the day a step's skip condition depends on an earlier interactive answer.

Known, accepted quirk

If a user goes back and changes kind after already picking a template, Huh's Select.selectOption() tries to preserve the old selection by value first; if the old template ID isn't in the new (re-filtered) option list, it falls back to clamping the cursor index into the new list rather than resetting to the first option. Not silent; the user sees whichever option that lands on highlighted before they press Enter ; but not necessarily a sensible default either. Documented in the terminal UX spec rather than worked around, since fixing it would mean depending on undocumented Huh internals for a low-severity cosmetic issue.

Limitations / open questions

  • Secret and MultiSelect step kinds don't exist yet; nothing in create's flow needs them today. Would need StepSecret/StepMultiSelect added to tux.StepKind if a future command's guided flow needs either.
  • No other command uses Prompter interactively yet, so the blast radius of this interface change was contained to one command's wiring; unclear how well the Skip func() bool (non-reactive) simplification holds up once a second interactive command exists.

Huh empty-bindings hash collision

Context

Building multi-screen back-navigation for Aruo's aruo create command (charm.land/huh v2.0.3, bubble-tea v2.0.2). The rich prompter now builds one continuous multi-group huh.Form instead of one isolated form per screen, so Huh's native shift+tab works across the whole flow. Steps whose content depends on an earlier answer (the template list narrowing to the chosen project kind) use Select.OptionsFunc(f func() []Option[T], bindings any); the field re-evaluates f when hash(bindings) changes.

Every step, including the very first one with no preceding steps to depend on, was wired the same way for uniformity: build a []any of the preceding steps' bound pointers as bindings. For step 0 that's an empty slice, []any{}.

The bug

The very first step's OptionsFunc/TitleFunc never fired. Down/Up arrow keys on the first Select had no visible effect; the field silently kept its zero-value state as if Options had never been set.

Root cause

Huh's Eval[T] (eval.go) decides whether to recompute via:

func (e *Eval[T]) shouldUpdate() (bool, uint64) {
    if e.fn == nil {
        return false, 0
    }
    newHash := hash(e.bindings)
    return e.bindingsHash != newHash, newHash
}

bindingsHash is a uint64 that starts at its Go zero value, 0. hash() wraps hashstructure.Hash(val, hashstructure.FormatV2, nil). Confirmed directly:

hashstructure.Hash([]any{}, hashstructure.FormatV2, nil) // → 0, nil

An empty slice hashes to exactly 0; the same as the field's untouched bindingsHash. shouldUpdate() compares 0 != 0, gets false, and the recompute never happens. Every subsequent step in the form was fine, because their bindings included at least one prior step's pointer, producing a non-zero hash on first evaluation.

Huh's public API and documentation do not mention this behavior. It shows up by noticing that a dynamically-configured field (OptionsFunc, TitleFunc) with a data-free binding silently behaves as if it were never configured at all, while the identical field wired with .Options() (the static setter) works.

Fix

Seed bindings with a nonempty value; the step's own ID is enough, since it's guaranteed present and stable per step:

bindings := make([]any, 0, index+1)
bindings = append(bindings, step.ID)
for i := 0; i < index; i++ {
    bindings = append(bindings, fields[i].pointer)
}

How it was found

Reading the Huh source didn't surface it; the mechanism looks correct on paper. Found by writing a throwaway test that drives a huh.NewSelect[...]().OptionsFunc(...) field directly through Form.Init()

  • simulated keypresses (see the companion note on testing Huh forms without a PTY) and printing whether the closure itself ever ran. It didn't, for the field with an empty bindings slice; it did, immediately, for an otherwise-identical field with a non-empty one. Isolating to that single variable took three progressively smaller repro tests.

Limitations / open questions

  • Only verified against Huh v2.0.3 and hashstructure/v2 v2.0.2; not confirmed whether this is fixed in a later Huh release, or whether it's considered a bug upstream at all.
  • Didn't check whether other zero-hashing bindings values exist besides an empty slice (e.g. a nil pointer, a zero-value struct); the fix here sidesteps the whole class by never passing an empty/trivial binding, but a more general fix would need to know exactly which values hash to 0.

Testing Huh forms without a PTY

Problem

Aruo's rich terminal prompter is built on Huh v2 / Bubble Tea v2. Earlier work in the same project had already established that this sandbox can't run a real PTY through Bubble Tea's Program loop end to end; the rich TUI sends terminal capability queries (sync-output mode probe, Kitty keyboard protocol probe) that a basic pty.fork()-based test harness can't answer, so it just hangs. That's why the project's existing charm prompter tests exercise Huh's own WithAccessible(true) mode instead; a built-in scriptable-stdin fallback.

That fallback stopped being usable for the specific feature being built: multi-group forms with WithHideFunc (conditionally hidden groups) and OptionsFunc/TitleFunc (reactive content). Reading Huh's source directly confirmed Form.runAccessible is a plain nested loop over every group and field, unconditionally; it does not check group.hide at all, and each field's RunAccessible reads only the field's static .val, never invoking .fn. So testing the real navigation and reactivity mechanics needed something else.

Direct model updates

huh.Form satisfies tea.Model (via a package-level alias, type Model = compat.Model, itself ultimately tea.Model). Update is just an ordinary Go method: func (f *Form) Update(msg tea.Msg) (Model, tea.Cmd). Nothing about calling it requires a running tea.Program, a terminal, or any I/O at all; it's pure state transition logic that happens to normally be driven by Bubble Tea's runtime loop.

So: build the form, call Init(), then call Update(tea.KeyPressMsg{...}) directly with real key values (tea.KeyDown, tea.KeyEnter, tea.Key{Code: tea.KeyTab, Mod: tea.ModShift} for shift+tab), reading Key.String() output to confirm each constructs the exact string Huh's key.Matches compares against ("down", "enter", "shift+tab").

Draining commands

A single Update call frequently returns a non-nil tea.Cmd; a func() tea.Msg that would normally be handed back to Bubble Tea's runtime, executed asynchronously, and its result fed back into another Update call. Some of these carry real state (the message produced by an OptionsFunc recompute); some are internal batching wrappers (tea.BatchMsg, and Huh's own unexported sequenceMsg, both literally []tea.Cmd under a different name); some are time-based (spinner.Tick, which wraps time.After).

A test has to replicate enough of that loop by hand:

func drainGuideForm(t *testing.T, model huh.Model, cmd tea.Cmd) huh.Model {
    pending := []tea.Cmd{cmd}
    for len(pending) > 0 {
        next := pending[0]
        pending = pending[1:]
        if next == nil {
            continue
        }
        msg, ok := callWithTimeout(next) // see below
        if !ok || msg == nil {
            continue
        }
        if more, ok := asCmdSlice(msg); ok { // reflection, see below
            pending = append(pending, more...)
            continue
        }
        var newCmd tea.Cmd
        model, newCmd = model.Update(msg)
        if newCmd != nil {
            pending = append(pending, newCmd)
        }
    }
    return model
}

Two details were required:

  1. Unwrapping batch/sequence messages via reflection, not type assertion. sequenceMsg is unexported; a test outside the huh package can't write msg.(huh.sequenceMsg). But both tea.BatchMsg and sequenceMsg are just named []tea.Cmd. reflect.ValueOf(msg).Kind() == reflect.Slice, iterate .Index(i).Interface().(tea.Cmd), and both shapes unwrap identically without ever needing the unexported type name.
  2. Bounding each cmd() call with a short timeout instead of calling it inline. Calling spinner.Tick (or anything wrapping tea.Tick) directly blocks on a real timer. Running each cmd in its own goroutine with a select against time.After(20ms) lets real, fast, synchronous cmds (the actual data-carrying ones) resolve normally, while anything that blocks longer than that gets abandoned; assumed to be a cosmetic animation command irrelevant to the state under test. The abandoned goroutine can leak past the deadline; acceptable for a short-lived test process.

With both of those in place, a test can type-simulate an entire interaction; select an option, submit, shift+tab back, change the earlier answer, move forward again; and assert on the form's bound Go values directly, with zero PTY involvement, fully deterministic, and fast (each of the three tests built this way ran in under 0.2s).

Where this technique came from

Not discovered by reading Huh's docs; its README doesn't mention testing at all. Found by locating Huh's own test suite (huh_test.go) and noticing it drives navigation via direct message injection into Form.Update too, just using Huh's unexported message constructors internally (f.Update(prevGroup()) etc., since the test lives inside package huh). The adaptation for external test code was working out the reflection-based generic unwrap and the timeout-bounded cmd execution, neither of which Huh's own tests needed since they never send commands that produce further async cmds requiring drainage across a whole form assembled from scratch outside the package.

Limitations / open questions

  • This proves the logic is correct (real keystrokes, real Huh navigation/reactivity code paths) but not the actual terminal rendering or the capability-query handshake that made PTY testing impractical in the first place; those remain unverified in this sandbox.
  • The 20ms cmd timeout is a heuristic tuned against this specific form's cmd shapes (title/description/options recompute, focus, spinner tick). A form with a slow synchronous OptionsFunc (e.g. a real network call) would need a longer timeout or a different signal to distinguish "real but slow" from "cosmetic and endless."