Testing¶
Four tiers, matching spec §12.
1. Unit tests¶
swift test
Runs SpacialShellKitTests, SpacialShellPlatformTests, and AXAppLivenessTests (the one
member of PlatformIntegrationTests that needs no Accessibility grant — it pins that every AXApp
call answers and that destroy() is final): the model and its invariants, every
command on hand-built worlds, layout frame math, the reconciler's write planning, config parsing,
hotkey parsing/tables, and property tests that throw random world/command/event sequences at the
model and assert spec §4.2's five invariants still hold plus that Layout.frames never returns
overlapping frames. No permissions needed, and it stays well under 5 seconds — if it's taking
longer, something regressed.
Two property-style sweeps are worth knowing about specifically, since a failure names a seed you can reproduce directly:
PropertyTests.randomSequencesPreserveInvariants(Tests/SpacialShellKitTests/PropertyTests.swift) drivesWorlddirectly through 200 random seeds.WorldStoreTests.randomSnapshotsPreserveInvariants(Tests/SpacialShellKitTests/WorldStoreTests.swift) drives theWorldStoreactor through aFakeBackend— 50 seeds, each replaying 30 randomapply(.snapshot(...))/run(command)/.windowMovedevents — so it also exercises adoption-from-snapshot, native focus, and the parking side tables that the direct-Worldsweep above never touches.
Both use Tests/SpacialShellKitTests/Support/TestRNG.swift, a deterministic SplitMix64 generator
seeded per test case, so swift test --filter randomSnapshotsPreserveInvariants reproduces a
specific seed's failure exactly by reading the seed number out of the failure message.
Declare every module a test imports (#156)¶
A test target must list every package module it imports in Package.swift, even one that
another dependency already brings in. Otherwise incremental swift test can segfault after a
change to a Kit struct:
- SwiftPM re-runs a target's compile only when a declared dependency's
.swiftmodulechanges. A module reached transitively (ShellStoryTests→SpacialShellUI→SpacialShellKit) can be imported, but it is not a build input. - Adding a stored property to
ConfigorWorldrewritesSpacialShellKit.swiftmodule. Platform and UI recompile but keep the same interface, and the compiler leaves those.swiftmodulefiles untouched. So a test target that lists only Platform or UI is never rebuilt, and its objects keep the struct's old layout. - Value-copy helpers such as "outlined init with copy of Config" are weak symbols, emitted into
every object that copies the struct. The linker keeps one of them. If it keeps the stale copy,
it breaks every caller, recompiled ones included. The crash is signal 11 in
swift_retain←outlined init with copy of Config←Config.init(), in whichever test first copies the struct (SettingsTests,AXWindowBackendTests,KeyboardGrammarTests, …).
Scripts/check-target-deps.py fails when a target imports a sibling target it doesn't declare.
The pre-commit hook and CI both run it. SwiftPM's own
--explicit-target-dependency-import-check accepts transitive modules, so it doesn't catch this.
If a local build crashes this way anyway, swift package clean recovers.
2. Classifier corpus regression¶
swift test --filter ClassifierCorpusTests
Tests/SpacialShellPlatformTests/ClassifierCorpusTests.swift replays every fixture under
axDumps/*.json5 — real AX-tree dumps captured from live apps — through the window classifier and
checks the result against the expected tile/float/ignore verdict recorded alongside each fixture.
This is what pins classifier behaviour against real-world AX quirks (Zoom, Chrome, dialogs, popups)
without needing those apps installed or a live Accessibility grant. Also no permissions needed.
3. Platform integration test (opt-in, needs the Accessibility grant)¶
Tests/PlatformIntegrationTests/TextEditTests.swift is the one tier that talks to real windows.
It is gated behind an environment variable so it never runs by accident (in CI or otherwise) and
never blocks swift test. (Its target-mate AXAppLivenessTests is not gated — it needs no grant
and runs in tier 1 above.)
swift test # TextEditTests is skipped, suite stays green
SPACIAL_INTEGRATION=1 swift test --filter TextEditTests # actually runs it
What it does: opens two plain .txt files in TextEdit (via
NSWorkspace.open(_:withApplicationAt:configuration:) — deliberately not AppleScript/Apple Events,
which would need a separate Automation TCC grant on top of Accessibility), drives a real
AXWindowBackend + WorldStore against them, and asserts that .maximize parks one window while
showing the other, and that cycling to .split shows both. It terminates TextEdit and deletes its
scratch files when done.
Before it can pass, the test host needs the Accessibility grant. swift test runs the suite
inside an .xctest bundle (currently
.build/arm64-apple-macosx/debug/SpacialShellPackageTests.xctest, architecture/build-config
folder aside) — like the debug app binary described in Scripts/README, TCC identifies this loose
binary by its path, not a bundle id, so:
- Run
SPACIAL_INTEGRATION=1 swift test --filter TextEditTestsonce. It will fail (AXError/ timeouts) — that first run is what makes the process request the Accessibility permission and register in TCC. - Open System Settings → Privacy & Security → Accessibility and tick the checkbox for the test
runner that just appeared (it may be listed as the
.xctestbundle,xctest, orSpacialShellPackageTests, depending on Xcode/toolchain version). - Re-run the same command. It should pass.
- Rebuilding changes the ad-hoc signature of the test binary, which can silently invalidate the
grant (checkbox stays ticked,
AXIsProcessTrusted()returns false anyway) — the fix is the same untick/re-tick or remove-and-re-grant dance described inScripts/README.
The suite creates exactly one AXWindowBackend and calls its terminal stop() exactly once, at
the end of the one test — AXApp's registry is process-global, so a second backend in the same
process would be unsound.
4. Manual checklist¶
Story snapshots (Tests/ShellStoryTests)¶
Storybook for the shell: every interesting state of the rail, tab bar, overview and cheat sheet is
a fixture "story" (Stories.swift), rendered at true panel geometry in light and dark. Each story
is image-snapshotted (references under __Snapshots__/, committed; SNAPSHOT_RECORD=1 swift test
--filter ShellStoryTests re-records) and run through LayoutLint, which fails on overlapping
text, content escaping its container, or content that wants more space than the panel has —
the wrap/truncate/line-height class of regressions a pixel diff alone can hide. Stories flagged
knownOverflow model conditions the design hasn't built handling for yet (tab "+N" badge, rail
overflow) and run under withKnownIssue. Before its snapshot, each story is spun on the run loop
until five passes in a row render identically (#159), not for a fixed time. So a scroll into view
is in the picture however loaded the machine is. A story whose view focuses a field as it appears
(the overview) is marked awaitsFocus, and it also waits until that field holds focus. A story that
never stops changing fails as "never settled", so give it fixed inputs (reduceMotion: true, a
fixed elapsed, fixture images at a fixed pixel size). macOS-only; new UI states get a story with their PR,
and the recorded stills feed the PR-media rule in AGENTS.md.
The items below need a human, a real display arrangement, and things a test suite cannot simulate (unplugging a monitor, locking the screen, a real password field). None of this is automated; record what you observe.
| # | Check | Procedure | What to look for |
|---|---|---|---|
| 1 | 3-display unplug/replug | Run SpacialShell with three displays attached. Unplug one, wait, replug it. | Windows on the surviving displays keep their positions throughout; the reconnected display's workspace stack comes back instead of being silently dropped; no crash or stuck "parked forever" window. |
| 2 | Lock/unlock | ⌃⌘Q to lock the screen, wait a few seconds, unlock. |
No writes happen while locked (spec §7.7); a fresh reconcile runs on unlock and the model matches reality afterward. |
| 3 | Hung app | Launch a managed app, then kill -STOP <pid> it from a terminal. Try to focus/move its window. Then kill -CONT <pid>. |
The hung app's AX calls time out (bounded by ax-timeout-ms) without blocking commands aimed at other windows; after kill -CONT, the app catches back up on the next reconcile without needing a restart. |
| 4 | Native fullscreen round-trip | Put a managed window into native fullscreen (green-button double-click or the fullscreen shortcut), then take it back out. | While fullscreen the window keeps its tab, is never moved or parked by the shell, and the rail and tab bar do not draw over its fullscreen Space. Focus another app, then click its tab (or Fn+A/D to it): macOS switches back to the fullscreen Space. When it leaves fullscreen it retiles in its own slot — not duplicated, not lost. |
| 5 | Zoom parking | Get a Zoom window parked (switch away from its workspace, or to a layout that doesn't show it). | Zoom is one of the zeroSliverBundleIDs apps — confirm it still parks cleanly (no visible sliver, no window left onscreen) rather than clamping to something visible like some AX-strict apps do. |
| 6 | Secure Input in a password field | Click into a real password field (a browser login form, login in Terminal, etc.) so the OS enables Secure Input, then press a bound chord. |
Secure Input drops keyDown events for every tap system-wide (spec §13.2) — bound chords do nothing while focus is in the field, then work again as soon as focus leaves it. No crash, no stuck modifier state. Within about 2 s the rail cog lists "App is waiting for a password ('window title', on workspace N)…" (#193), and Show window brings that window up; it clears within about 2 s of leaving the field. A password prompt on another workspace (a wallet's unlock window, a Keychain prompt) appears centred on the current one within about 2 s, in front but not focused, and is parked again once its field loses focus or it closes (#195). |
| 7 | Quit restores windows | With several windows parked (on background workspaces or other screens), quit SpacialShell: Ctrl-C on the dev binary, or kill <pid> (SIGTERM) otherwise. There is no menu bar, so no ⌘Q, and Fn+Q with nothing focused is a no-op, not a quit shortcut. |
Either path runs the termination gate (state save → restore parked windows → stop): every parked window ends up centred on its screen (spec §7.4) — nothing is left sitting in a parking corner after the app is gone. Windows that were tiled and visible stay exactly where they were: they are already reachable, so the restore leaves them alone rather than piling the whole layout into the middle of the screen. |
Related empirical checks — the hotkey-tap-vs-symbolic-hotkey race, Fn+arrow keycodes, parking
sliver geometry, stacked-display parking, and tap survival across sleep/lock — are tracked
separately in docs/platform-notes.md, with a pending grant result until someone with the
Accessibility grant runs through them. That page also explains why those specific checks can't be
scripted (they need a human to watch for a system UI popping up, or a real keyboard chord).