Files
paseo/packages/server
Mohamed Boudra 617cf8a7bf refactor(server): extract voice mode subsystem into VoiceSession (#1640)
* refactor(server): extract voice mode subsystem into VoiceSession

session.ts was a 10.5k-line god object and the most-churned file in the
repo. ~1,100 of those lines were an entire voice/audio subsystem — the
STT/TTS/dictation managers, the barge-in audio-buffering state machine,
voice-turn orchestration, and the MCP voice bridge — interleaved field by
field and method by method with workspace/git/agent/provider concerns.

Move the whole subsystem into a new VoiceSession deep module. Session now
holds one `voiceSession` field, constructs it once, and delegates the
voice/dictation/abort message types, the permission auto-allow gate, and
cleanup to it. VoiceSession owns all 18 voice fields and ~25 methods and
reaches back into the agent run only through a narrow VoiceSessionHost
seam (emit, loadAgent, reloadAgentSession, sendSpokenInput,
interruptAgentIfRunning, hasActiveAgentRun).

Behavior is preserved (verbatim method moves); deleting voice-session.ts
now removes the feature mechanically. The voice unit tests drive the
VoiceSession boundary instead of reaching into Session internals, so
future voice tests can construct a VoiceSession with a fake host.

session.ts: 10,468 -> 9,272 lines.

* refactor(server): move VoiceSession into voice/ and test it at the boundary

Address Greptile review on #1640:

- Move voice-session.ts into the existing voice/ subdirectory, alongside
  voice-turn-controller.ts, instead of adding another peer to the 30+ file
  server/ directory. The directory now carries the domain.
- Add voice/voice-session.test.ts beside the module: it constructs a
  VoiceSession with a fake VoiceSessionHost and drives it through the public
  API (handleSetVoiceMode, then turn-detection/STT events), asserting on the
  host seam (sendSpokenInput) and emitted messages. No Session, no private
  field/method access.
- Remove the voice tests from session.test.ts that reached through Session's
  private voiceSession field into VoiceSession internals.

Same three behaviors are covered (streaming final -> agent, finalization
timeout empty path, low-confidence filtering); driving handleSetVoiceMode
additionally exercises the enable path the old tests bypassed.
2026-06-21 07:14:23 +00:00
..
2026-06-18 15:02:42 +08:00
2026-05-28 01:58:18 +08:00
2026-06-21 00:30:25 +07:00
2026-04-23 21:44:48 +07:00

Voice Assistant

A voice-controlled terminal assistant that runs as a single local service.

Quick Start

# Install dependencies
npm install

# Copy environment variables
cp .env.example .env
# Edit .env and add your API keys (OpenAI, Deepgram)

# Run development servers
npm run dev

# Open browser to http://localhost:5173

Architecture

  • Express Server (port 3000) - Serves API and built UI in production
  • Vite Dev Server (port 5173) - Hot-reload React UI in development
  • WebSocket (/ws) - Real-time bidirectional communication
  • Agent - STT → LLM → TTS pipeline with terminal control
  • Daemon - tmux-based terminal management (in-process)

Development

# Run both servers (recommended)
npm run dev

# Or run separately:
npm run dev:server  # Express on port 3000
npm run dev:ui      # Vite on port 5173

# Type checking
npm run typecheck

# Build for production
npm run build

# Start production server
npm start

Project Status

Completed (Phases 1-2):

  • Package setup and configuration
  • Express server with WebSocket
  • React UI with Vite
  • WebSocket client with ping/pong testing

In Progress (Phase 3):

  • Terminal control (tmux integration)

📋 Planned (Phases 4-9):

  • LLM integration (OpenAI GPT-4)
  • Agent orchestrator
  • Speech-to-Text (Deepgram)
  • Text-to-Speech (OpenAI)
  • Audio streaming
  • UI polish

See IMPLEMENTATION_PLAN.md for complete details.

Environment Variables

OPENAI_API_KEY=your-openai-key-here      # GPT-4 and TTS
DEEPGRAM_API_KEY=your-deepgram-key-here  # Streaming STT
STT_MODEL=whisper-1        # Optional: override to gpt-4o-transcribe, etc.
STT_CONFIDENCE_THRESHOLD=-3.0  # Optional: reject low-confidence clips
STT_DEBUG_AUDIO_DIR=.stt-debug # Optional: persist raw dictation audio for debugging
PASEO_HOME=~/.paseo        # Runtime state directory (agents/, etc.)
PASEO_LISTEN=127.0.0.1:6767  # Listen address (host:port or /path/to/socket)

PASEO_HOME defaults to ~/.paseo and isolates runtime artifacts like agents/. PASEO_LISTEN controls the daemon listen address. For blue/green testing you can run a parallel server without touching production state:

PASEO_HOME=~/.paseo-blue PASEO_LISTEN=127.0.0.1:7777 npm run dev

Tech Stack

  • Server: Express, TypeScript, ws (WebSocket)
  • Client: React 18, Vite, TypeScript
  • Terminal: tmux (via child_process)
  • AI: OpenAI (LLM + TTS), Deepgram (STT)

Testing

Currently manual testing via:

  1. Start servers: npm run dev
  2. Open http://localhost:5173
  3. Test WebSocket connection (green status indicator)
  4. Click "Send Ping" button to test communication

More testing guidance as features are implemented.

License

MIT