mirror of
https://github.com/getpaseo/paseo.git
synced 2026-07-29 12:01:31 +00:00
* refactor(server): extract voice mode subsystem into VoiceSession session.ts was a 10.5k-line god object and the most-churned file in the repo. ~1,100 of those lines were an entire voice/audio subsystem — the STT/TTS/dictation managers, the barge-in audio-buffering state machine, voice-turn orchestration, and the MCP voice bridge — interleaved field by field and method by method with workspace/git/agent/provider concerns. Move the whole subsystem into a new VoiceSession deep module. Session now holds one `voiceSession` field, constructs it once, and delegates the voice/dictation/abort message types, the permission auto-allow gate, and cleanup to it. VoiceSession owns all 18 voice fields and ~25 methods and reaches back into the agent run only through a narrow VoiceSessionHost seam (emit, loadAgent, reloadAgentSession, sendSpokenInput, interruptAgentIfRunning, hasActiveAgentRun). Behavior is preserved (verbatim method moves); deleting voice-session.ts now removes the feature mechanically. The voice unit tests drive the VoiceSession boundary instead of reaching into Session internals, so future voice tests can construct a VoiceSession with a fake host. session.ts: 10,468 -> 9,272 lines. * refactor(server): move VoiceSession into voice/ and test it at the boundary Address Greptile review on #1640: - Move voice-session.ts into the existing voice/ subdirectory, alongside voice-turn-controller.ts, instead of adding another peer to the 30+ file server/ directory. The directory now carries the domain. - Add voice/voice-session.test.ts beside the module: it constructs a VoiceSession with a fake VoiceSessionHost and drives it through the public API (handleSetVoiceMode, then turn-detection/STT events), asserting on the host seam (sendSpokenInput) and emitted messages. No Session, no private field/method access. - Remove the voice tests from session.test.ts that reached through Session's private voiceSession field into VoiceSession internals. Same three behaviors are covered (streaming final -> agent, finalization timeout empty path, low-confidence filtering); driving handleSetVoiceMode additionally exercises the enable path the old tests bypassed.
Voice Assistant
A voice-controlled terminal assistant that runs as a single local service.
Quick Start
# Install dependencies
npm install
# Copy environment variables
cp .env.example .env
# Edit .env and add your API keys (OpenAI, Deepgram)
# Run development servers
npm run dev
# Open browser to http://localhost:5173
Architecture
- Express Server (port 3000) - Serves API and built UI in production
- Vite Dev Server (port 5173) - Hot-reload React UI in development
- WebSocket (
/ws) - Real-time bidirectional communication - Agent - STT → LLM → TTS pipeline with terminal control
- Daemon - tmux-based terminal management (in-process)
Development
# Run both servers (recommended)
npm run dev
# Or run separately:
npm run dev:server # Express on port 3000
npm run dev:ui # Vite on port 5173
# Type checking
npm run typecheck
# Build for production
npm run build
# Start production server
npm start
Project Status
✅ Completed (Phases 1-2):
- Package setup and configuration
- Express server with WebSocket
- React UI with Vite
- WebSocket client with ping/pong testing
⏳ In Progress (Phase 3):
- Terminal control (tmux integration)
📋 Planned (Phases 4-9):
- LLM integration (OpenAI GPT-4)
- Agent orchestrator
- Speech-to-Text (Deepgram)
- Text-to-Speech (OpenAI)
- Audio streaming
- UI polish
See IMPLEMENTATION_PLAN.md for complete details.
Environment Variables
OPENAI_API_KEY=your-openai-key-here # GPT-4 and TTS
DEEPGRAM_API_KEY=your-deepgram-key-here # Streaming STT
STT_MODEL=whisper-1 # Optional: override to gpt-4o-transcribe, etc.
STT_CONFIDENCE_THRESHOLD=-3.0 # Optional: reject low-confidence clips
STT_DEBUG_AUDIO_DIR=.stt-debug # Optional: persist raw dictation audio for debugging
PASEO_HOME=~/.paseo # Runtime state directory (agents/, etc.)
PASEO_LISTEN=127.0.0.1:6767 # Listen address (host:port or /path/to/socket)
PASEO_HOME defaults to ~/.paseo and isolates runtime artifacts like agents/. PASEO_LISTEN controls the daemon listen address. For blue/green testing you can run a parallel server without touching production state:
PASEO_HOME=~/.paseo-blue PASEO_LISTEN=127.0.0.1:7777 npm run dev
Tech Stack
- Server: Express, TypeScript, ws (WebSocket)
- Client: React 18, Vite, TypeScript
- Terminal: tmux (via child_process)
- AI: OpenAI (LLM + TTS), Deepgram (STT)
Testing
Currently manual testing via:
- Start servers:
npm run dev - Open http://localhost:5173
- Test WebSocket connection (green status indicator)
- Click "Send Ping" button to test communication
More testing guidance as features are implemented.
License
MIT