* fix(server): harden Codex/Windows startup and provider resolution
Addresses #452, #443, #353, #418, #403, #221, #307, and locks in #284.
- Replace custom where.exe/which output parsing with npm `which@^5` plus a
spawn probe. findExecutable now enumerates all PATH+PATHEXT candidates
and returns the first invokable one. A WindowsApps-ACL'd codex.exe no
longer wins over a working codex.cmd (#452).
- Make default provider isAvailable() check the binary instead of always
returning true. Codex/Claude/OpenCode default isAvailable() now defer
to isCommandAvailable() so missing CLIs surface as unavailable instead
of throwing later from spawn (#221, #443).
- Gate AgentManager.resumeAgentFromPersistence on isAvailable() so a
persisted agent record with a missing binary cannot reach provider
spawn during rehydration. The daemon stays up and the agent reports
unavailable (#443, #353, #418, #403).
- Drop --path-format=absolute from rev-parse callers and validate stdout
through a shared parser that rejects multi-line output and unknown-flag
echoes. --show-toplevel is absolute by default in modern Git;
--git-common-dir is resolved against the command cwd. Fixes workspace
registration on pre-2.31 Git that echoed the unknown flag and produced
a two-line "path" (#307).
Tests:
- Real-FS executable.test.ts using temp PATH fixtures, covering the .cmd
fallback after .exe pre-spawn failure (Windows-only) plus the
null-on-no-invokable-candidate case.
- provider-availability.test.ts builds real provider clients against a
temp-dir-only PATH for Codex/Claude/OpenCode.
- bootstrap-provider-availability.test.ts builds the daemon and triggers
ensureAgentLoaded so it actually exercises resumeAgentFromPersistence.
- claude-agent.spawn.test.ts asserts shell: false reaches spawnProcess
from the Claude SDK spawn override, locking in 39b56af4 for #284.
- checkout-git-rev-parse.test.ts covers nested-checkout resolution and
the old-Git multi-line stdout case via a tightly-scoped runGitCommand
fake.
CI:
- server-tests-windows now also runs the new and modified test files so
the Windows behaviors are exercised on windows-latest.
* fix: update lockfile signatures and Nix hash
* fix(server): handle synchronous spawn UNKNOWN on Windows
On Windows, child_process.spawn() throws synchronously when invoked on
a corrupt or invalid .exe (e.g., a WindowsApps stub the current user
cannot execute, or a zero-byte file). The executable probe did not
guard the spawn call, so the synchronous throw rejected the probe
promise instead of resolving false, preventing findExecutable() from
trying the next candidate. This is the root cause of the daemon-crash
pattern in #452: codex.exe from WindowsApps would hard-fail before the
.cmd shim ever got a chance.
Wrap spawn() in try/catch and settle false on sync throw. The existing
error-event and exit-event handlers already cover async failure modes;
sync throw just needed one more guard.
Also adjust three Windows-only test comparisons that were asserting
platform-dependent string equality:
- executable.test: compare .cmd paths case-insensitively (which@5
preserves PATHEXT casing, which is uppercase in production).
- workspace-registry-model.test: expect normalizeWorkspaceId(path),
not the hardcoded POSIX form.
- checkout-git-rev-parse.test: normalize separators/case when
comparing git's Windows forward-slash output against realpathSync.
* test(server): canonicalize repo root via git on Windows to fix short-name mismatch
realpathSync on Windows preserves 8.3 short names (e.g. RUNNER~1) while git's
rev-parse --show-toplevel always returns the long-name form (runneradmin).
Use git as the canonicalizer on both sides of the assertion so the comparison
holds regardless of how Windows exposes the temp directory path.
* fix(server): Claude is always available in default mode; SDK bundles cli.js
The Phase 2 change wrongly tied Claude's default-mode isAvailable() to
isCommandAvailable("claude"). Claude's default runtime does not use an
external `claude` binary — @anthropic-ai/claude-agent-sdk ships its own
cli.js and spawnClaudeCodeProcess runs it via process.execPath. The
previous `return true` was correct; restore it.
Update provider-availability.test.ts to assert the truthful behavior:
Claude reports available even when no `claude` binary is on PATH.
Codex and OpenCode genuinely require their binaries on PATH, so their
availability checks remain unchanged.
---------
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
The full server test suite has deep Unix assumptions (hardcoded /tmp
paths, mkdir -p, Unix sockets, bash variable expansion). Rather than
porting the entire suite, run only the tests that exercise
Windows-specific code paths: executable resolution, spawn/exec,
git command handling, provider registry, and config loading.
Provider isAvailable() was using executableExists() which only checks
filesystem paths, not PATH. Commands like ["claude", "--flag"] would
show as unavailable even though they'd launch fine. Switch to
isCommandAvailable() which uses findExecutable() for proper PATH
resolution. Un-export executableExists from the public API.
Fix Windows cmd.exe metacharacter escaping — &, |, ^, <, >, (, ), !
are now properly escaped with ^, and % is doubled. Add shared
escapeWindowsCmdValue helper used by both quoteWindowsCommand and
quoteWindowsArgument. Replace local quoteForCmd in spawn.ts.
Add server-tests-windows CI job to catch Windows-specific regressions.
* ci: add CI status tracker for test fix iteration
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* cli: honor daemon connect timeout
* app: fix e2e helper server path
* e2e: fix helper imports and ws cleanup
* cli: align daemon status tests
* style: autoformat with biome to fix 195 formatting errors
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* server: fix 4 stale test expectations and setup bug
- logger.test.ts: update expected default level from trace to debug
matching intentional product change
- session.workspaces.test.ts: only opened worktree reconciles, not
siblings; add explicit reconcileWorkspaceRecord before owner-change
assertion
- worktree.test.ts: add explicit git checkout -B main origin/main
for deterministic CI branch state
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* server: align daemon-client test expectations with ignoreWhitespace field
normalizeCheckoutDiffCompare() now always emits ignoreWhitespace: false,
so update the two checkout-diff subscribe test assertions to match.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* style: format daemon-client test to satisfy biome
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* cli: update provider test for new providers and model lineup
Update 15-provider test expectations to match current product state:
- Provider count: 3 → 5 (added copilot, pi)
- Claude models: 3 → 4 (added claude-opus-4-6[1m])
- Codex models: replace retired gpt-5.1-* with gpt-5.4/gpt-5.4-mini
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* app: fix 18 failing test files with vitest setup and stale expectations
Add vitest.setup.ts to define __DEV__, shim Expo globals, mock
react-native-unistyles/expo-linking, and stub @xterm/addon-ligatures.
Update stale test expectations across combined-model-selector,
use-settings, tool-call-display, sidebar-project-row-model,
sidebar-shortcuts, keyboard-shortcuts, host-runtime,
use-agent-form-state, desktop-permissions, and voice-runtime
to match current source behavior.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* ci: add missing highlight build step to app-tests job
The app-tests CI job was missing the `npm run build --workspace=@getpaseo/highlight`
step that all other jobs have. This caused diff-highlighter.test.ts to fail with
"Failed to resolve entry for package @getpaseo/highlight" because the dist/ directory
did not exist.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* ci: install codex and opencode CLIs in cli-tests job
The 15-provider test expects `provider models codex` and
`provider models opencode` to succeed, which requires the
actual CLI binaries to be on PATH.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* app: fix Playwright e2e helpers using import.meta.url in CJS context
Revert dynamic import path resolution from `new URL(..., import.meta.url)`
to `path.resolve(__dirname, ...)` + `pathToFileURL(...)` in three e2e
helpers. Playwright's TS loader emits CommonJS, where import.meta.url
is undefined, causing "exports is not defined in ES module scope" and
blocking all e2e test discovery.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* ci: add missing server build step to playwright job
The Playwright e2e helpers dynamically import from
packages/server/dist/server/server/exports.js, but the CI job
only built highlight and relay dependencies. Add the server build
step after relay so the dist artifacts exist when tests run.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* test(cli): make codex model assertions resilient to catalog changes
The 15-provider test hard-coded exact model IDs (gpt-5.3-codex-spark,
etc.) which depend on the external codex CLI's model/list endpoint.
Replace with structural checks: all IDs in gpt- family, at least one
codex-optimized model, and all models have required fields.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* test(cli): make opencode model assertions resilient to catalog changes
The 15-provider test hard-coded exact model IDs (opencode/gpt-5-nano,
openrouter/openai/gpt-5.3-codex) which break when the external opencode
CLI updates its model catalog. Replace with structural checks: at least
one first-party opencode model, at least one OpenAI-backed model, at
least one codex-optimized model, and all models have required fields.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* app: fix Playwright E2E WebSocket errors on Node 20 + fix format
E2E helpers used DaemonClient without a webSocketFactory, which relies
on globalThis.WebSocket — unavailable in Node 20 (CI). Add a shared
node-ws-factory helper using the ws package (matching the CLI pattern)
and inject it in all three E2E connection helpers. Also fix a biome
formatting issue in the cli provider test.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: update lockfile signatures and Nix hash
* test(cli): replace OpenAI-specific assertions with generic third-party check
The opencode provider test asserted models with "openai/" or
"openrouter/openai/" prefixes, but these are environment-dependent —
opencode returns whatever providers are connected, and CI may not have
OpenAI configured. Replace with a generic check that at least one
non-opencode/ namespaced model exists.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: advance stale schedule nextRunAt on daemon restart
On restart, persisted nextRunAt could be in the past, showing stale
dates in `schedule ls`. Now recoverInterruptedRuns() advances any
past-due nextRunAt forward to the next future tick.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* ci: install agent CLIs in playwright job + remove env-dependent opencode test assertions
The Playwright E2E tests (archive-tab, terminal-performance) fail because
the CI job doesn't install Codex/OpenCode binaries that the tests need to
create agents. Add the same install step already used in cli-tests.
The 15-provider CLI test still asserts third-party providers are connected
in OpenCode, which is environment-dependent. Remove that assertion and
lower the minimum model count to 1, keeping only deterministic structural
checks (namespacing, required fields, first-party models).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: apply biome formatting to schedule service files
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(cli): parse localDaemon field in daemon supervisor test assertions
The daemon status JSON uses `localDaemon` not `status`, so the polling
condition was always null and timed out after 120s on CI.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: apply biome formatting to daemon supervisor test files
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(cli): spawn supervisor directly via node --import tsx instead of npx wrapper
The daemon supervisor tests (22, 23, 25) spawned the supervisor through
`npx tsx` which creates a wrapper process. When SIGINT was sent, the npx
wrapper died with signal=SIGINT instead of forwarding it to the actual
supervisor which handles graceful shutdown. Now matches production startup
pattern using process.execPath with --import tsx.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(server): acquire pid lock for unsupervised daemon workers
The direct worker path (non-supervisor) was not writing paseo.pid,
so `paseo daemon status` could not detect the running daemon. This
caused test 26-daemon-restart-unsupervised to time out in CI waiting
for the status to become "running".
Acquire the pid lock before daemon creation, update it with the
listen address after start, and release on shutdown/error. Supervision
detection now requires both PASEO_SUPERVISED=1 and an active IPC
channel to avoid misclassification when the env var is inherited.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(e2e): correct terminal tab testid and navigation URL in Playwright helper
navigateToTerminal() used the bare workspace route instead of the URL
with ?open=terminal:<id> intent, and looked for testid
"workspace-tab-terminal:<id>" (colon) when the real tab key is
"terminal_<id>" (underscore). Both bugs prevented the terminal surface
from ever appearing, causing terminal-performance tests to timeout.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(e2e): wait for open-intent redirect before interacting with terminal surface
The navigateToTerminal helper was racing the workspace layout's ?open= redirect,
which returns null during the useEffect cycle. Now waits for the clean workspace
URL before looking for the terminal tab, and removes the fallback that created
a different terminal via the "New terminal tab" button.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: apply biome formatting to terminal-perf.ts
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(app): ungate host routes from storeReady to preserve deep links
Host routes (workspace, agent, sessions, etc.) were inside
Stack.Protected guard={storeReady}, causing deep links to be rejected
during initial storage hydration and redirected to /welcome. Move them
outside the guard so they can render their own loading state while
stores hydrate.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(app): gate open-intent consumption on navigation readiness
After moving host routes outside Stack.Protected, the workspace layout
could mount before the root navigator was ready. router.replace() would
silently fail, but consumedIntentRef was already set, preventing retries.
Gate the effect on rootNavigationState.key so the intent is only consumed
once Expo Router can actually process the replace.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(app): use history.replaceState to strip ?open since Expo Router skips query-only changes
Expo Router's findDivergentState ignores search params, so
router.replace with the same pathname but without ?open is a no-op.
Use window.history.replaceState on web to directly strip the query
param, and track intent consumption in component state so
WorkspaceScreen renders once the tab is prepared.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: update stale codex model from gpt-5.1-codex-mini to gpt-5.4-mini
The Codex CLI no longer supports gpt-5.1-codex-mini. Update all
references to gpt-5.4-mini, which is available in the current CLI.
This fixes archive-tab Playwright tests that create codex agents
which were erroring due to the unsupported model.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(ci): pin codex CLI to 0.105.0 and improve turn_failed diagnostics
Playwright archive-tab tests fail because CI installs @openai/codex@latest
(0.120.0) which has breaking protocol changes vs the known-working 0.105.0.
Pin the version and add diagnostics for future debugging: elevate turn_failed
log from TRACE to WARN, and include error details in test assertions.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(e2e): switch archive-tab tests from codex to opencode provider
Codex CLI requires OAuth auth (~/.codex/auth.json) that CI lacks —
OPENAI_API_KEY alone gets 401. These tests verify archive tab behavior,
not any specific provider. Switch to opencode/gpt-5-nano which
authenticates via standard OpenAI API that the CI key supports.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* ci: retrigger CI for PR #236
* ci: trigger CI for opencode provider switch
Previous commits (3cbd7976, bdf768d6) were pushed with a token
that did not trigger GitHub Actions workflows. This empty commit
forces a fresh pull_request event.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* ci: add workflow_dispatch trigger to CI workflow
Enable manual triggering of the CI workflow. Previous pushes to
ci/fix-tests-green failed to trigger pull_request events for the
CI workflow, so this allows manual dispatch as a fallback.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(ci): unblock remaining test lanes
* test(server): allow slower CI Claude model listing
* fix(cli): stabilize unsupervised restart regression
* test(ci): harden restart and archive e2e timing
* test(server): align workspace reconciliation expectation
* test(server): tolerate platform-specific workspace ids
* fix(app): gate host route logic on bootstrap readiness
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Add a create-release job that runs first and creates the GitHub release
before any platform build jobs start. This prevents the race condition
where parallel electron-builder jobs each create their own release.
Platform jobs now depend on create-release and just upload assets to
the existing release. Single-platform retry tags (desktop-macos-v*,
desktop-linux-v*, desktop-windows-v*) skip the create-release job
since the release already exists from the original build.
electron-builder overwrites latest-mac.yml when parallel arm64/x64
builds publish independently — whichever finishes last wins, leaving
the other architecture's users downloading the wrong binary.
Add a finalize-mac-manifest job that runs after both macOS builds
complete, merges their per-arch manifest artifacts into a single
latest-mac.yml containing all files, and uploads it to the release.
The website fetches the latest release version at build time, so it
needs to redeploy when a new release is published to pick up updated
download links.
The fallback was "release", so pushing a tag without pre-creating a
draft release on GitHub caused electron-builder to create a published
release instead of a draft.
Hash mismatch was hidden by continue-on-error. Now the Nix Build
workflow fails visibly, and fix-nix-hash runs on every push/PR that
touches the lockfile (not just Dependabot).
* fix: add missing resolved/integrity fields to package-lock.json
npm omits resolved URLs and integrity hashes for workspace-local
node_modules overrides. This breaks offline installers like Nix's
npm ci. Add the missing fields for 25 workspace-hoisted packages.
* feat: add Nix flake with package and NixOS module
Add a Nix flake that builds the Paseo daemon (server + CLI) and
provides a NixOS module for declarative deployment.
Package (nix/package.nix):
- Builds relay, server, and CLI workspaces
- Skips onnxruntime-node install script (sandbox-incompatible)
- Rebuilds only node-pty for native terminal support
- Source filter excludes app/website/desktop workspaces
NixOS module (nix/module.nix):
- Systemd service with configurable user, port, listen address
- allowedHosts for DNS rebinding protection
- relay.enable to toggle remote access via app.paseo.sh
- inheritUserEnvironment to expose user tools (git, ssh) to agents
- openFirewall and extra environment variables
ci: add Nix hash maintenance scripts and workflows
scripts/fix-lockfile.mjs:
Adds missing resolved/integrity fields to package-lock.json for
workspace-local overrides. Idempotent, uses `npm view`.
scripts/update-nix.sh:
Runs fix-lockfile.mjs, prefetches deps, computes NAR hash, and
updates npmDepsHash in nix/package.nix. Supports --check for CI.
.github/workflows/nix-build.yml:
Builds the Nix package on push/PR and verifies the lockfile and
hash are up to date.
.github/workflows/fix-nix-hash.yml:
Auto-fixes lockfile signatures and Nix hash on dependabot PRs.
fix: update npmDepsHash after upstream sync
nix: allowlist workspace symlinks instead of blocklist
Prevents build failures when upstream adds new workspace packages.
* don't block PRs on nix failures
* better document npm workaround
* fix hash update script, and update hash
* integrate with npm run build:daemon
* ci: trigger nix build on highlight changes
* fix(nix): update npmDepsHash
---------
Co-authored-by: Mohamed Boudra <boudra.moha@gmail.com>
The NODE_OPTIONS --require patch for metro config was being applied
during build:workspace-deps, causing expo-module-build (a bash script)
to be loaded through Node's JS loader on Windows.
- Change electron-builder publish owner from anthropics to getpaseo
- Remove CSC_NAME (auto-discovered from cert, secret had rejected prefix)
- Remove CSC_IDENTITY_AUTO_DISCOVERY=false from build script (breaks Windows cmd.exe)
Fix desktop-release workflow to reference correct package path after
Tauri→Electron migration. Add macOS entitlements and notarization config
for electron-builder.
The linuxdeploy-plugin-gtk hook forces GDK_BACKEND=x11, which
prevents GTK initialization on Wayland-only systems. The bundled
libgdk-3.so already has Wayland support built in.
Add a post-build step that extracts the AppImage, comments out the
GDK_BACKEND=x11 line, and repackages with appimagetool.
Tauri notarizes the .app but not the .dmg container. When users
download the DMG from GitHub Releases, macOS quarantines everything
and Gatekeeper doesn't clear quarantine on embedded helper binaries
(like the bundled Node runtime), causing SIGTRAP on first launch.
Add a post-build step that signs, notarizes, and staples the DMG,
then re-uploads it to the release.
The linuxdeploy failure was caused by CUDA shared library references in
onnxruntime-node, not by linuxdeploy itself. The CUDA stripping step
added in the previous commit fixes the root cause, so AppImage bundling
should work now.
linuxdeploy-plugin-appimage's "continuous" release on GitHub is broken,
causing every AppImage build to fail. The downloaded binary is actually
an HTML error page. Switch to deb format which doesn't depend on
linuxdeploy at all.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
linuxdeploy is an AppImage itself and needs FUSE to run. GitHub Actions
runners don't always have working FUSE support. Setting this env var
tells AppImage tools to extract-and-run instead, avoiding the FUSE
dependency.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
linuxdeploy scans all ELF binaries in the AppDir and fails when it
can't find libcublasLt.so.12 (a CUDA library referenced by the
onnxruntime native module). Use patchelf to remove these optional
CUDA dependencies since we only need CPU inference.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The build script fix only applies to future tags. For v0.1.24 (and any
tag built before that fix), we need the workflow itself to remove the
CUDA .so files after building the managed runtime.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
linuxdeploy is distributed as an AppImage and may need libfuse2 to
execute even with APPIMAGE_EXTRACT_AND_RUN=1. ubuntu-22.04 runners
don't have it by default.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>