- Install react-native-vad and expo-audio-studio for voice detection
- Create use-realtime-audio hook for continuous recording + VAD segmentation
- Add Android FOREGROUND_SERVICE permissions for background mic access
- Update UI components with Lucide icons and markdown rendering
- Add npm overrides to resolve peer dependency conflicts with Expo SDK 54
- Integrate realtime mode toggle with pulse animation for speech detection
Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)
Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Allow setting session mode when sending prompts, and include current mode in activity output.
Changes:
- sendPrompt now accepts options object with maxWait and sessionMode
- If sessionMode specified, sets mode before sending prompt
- send_agent_prompt MCP tool exposes sessionMode parameter
- get_agent_activity now returns currentModeId in output
- Enables workflow: send_agent_prompt({ sessionMode: "plan", prompt: "..." })
🤖 Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)
Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Remove grouping and clever formatting. Just convert activity feed to chronological text.
Changes:
- Process updates in order (chronological)
- Buffer message/thought chunks and flush when needed
- Output tool calls and plans as they appear
- No grouping, no reorganizing, no sections
- Simple text format: messages, [Tool] status, [Plan] entries
🤖 Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)
Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Create activity curator that consolidates hundreds of tiny message chunks into clean, human-readable text. Saves tokens and improves readability dramatically.
Changes:
- Add activity-curator.ts with curateAgentActivity() function
- Consolidate agent_message_chunk fragments into complete messages
- Format tool calls with status, input/output in readable structure
- Display plans with status icons (✅ completed, 🔄 in_progress, ⏳ pending)
- Add priority indicators (🔴 high, 🟡 medium, 🟢 low)
- Update get_agent_activity to support format='curated' (default) or 'raw'
- Curated format returns clean markdown text instead of JSON array
Example: 363 raw updates consolidating to single clean message instead of 363 JSON objects with fragments like " it.", " fix", "ify what".
🤖 Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)
Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Change sendPrompt to return immediately by default instead of blocking, allowing the orchestrator to continue while agents process in the background. Add optional maxWait parameter for cases where waiting is needed.
Changes:
- sendPrompt now returns { didComplete, stopReason? } immediately if no maxWait
- Background promise handles completion and updates agent status via protocol
- Optional maxWait parameter races completion vs timeout, returns didComplete status
- Update send_agent_prompt MCP tool to expose maxWait parameter
- Agent continues processing in background even if maxWait expires
- Orchestrator can poll get_agent_status or get_agent_activity to check progress
This prevents orchestrator from blocking on long-running agent tasks while maintaining protocol-based completion tracking.
🤖 Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)
Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Replace unreliable 2-second timeout heuristic with proper protocol-based completion detection using PromptResponse.stopReason from ACP SDK.
Changes:
- Capture and handle PromptResponse.stopReason in sendPrompt()
- Set agent status based on stop reason: end_turn → completed, refusal → failed, cancelled → ready
- Remove 2-second timeout completion heuristic (was causing drift)
- Update SessionStateMessageSchema to match AgentInfo null types
This ensures agent completion is detected via the protocol spec, not manual polling/timing.
Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)
Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Replace all undefined values with null in agent MCP tool responses to prevent JSON serialization issues that break the MCP client.
Changes:
- Update AgentInfo interface to use `| null` instead of optional `?` syntax
- Convert undefined to null in listAgents(), getCurrentMode(), getAvailableModes()
- Update all Zod schemas from .optional() to .nullable()
- Fix tool responses in create_coding_agent and set_agent_mode
This ensures all MCP tool responses properly serialize optional fields as null instead of undefined, preventing client-side errors.
🤖 Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)
Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Replace client-side permission mode hack with proper ACP session modes:
- Add typed Client interface implementation to ACPClient
- Import proper types from ACP SDK (SessionNotification, RequestPermissionRequest, etc.)
- Fix critical sessionUpdate bug where params.update was undefined
- Replace PermissionMode with SessionMode from ACP spec
- Add getCurrentMode(), getAvailableModes(), setSessionMode() to AgentManager
- Update MCP tools: replace set_agent_permission_mode with set_agent_mode
- Add session_state message type to send live agents/commands to client
- Update UI to display session mode badges (ask/code/architect)
- Remove all permissionsMode references from tests
Agents now have ONE mode (session mode) that controls their behavior,
system prompts, and tool availability - not separate permission/session modes.
Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)
Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
- Created new Expo app with SDK 54+ and Expo Router
- Installed and configured NativeWind v5 for Tailwind styling
- Upgraded voice-assistant to React 19 for consistency
- Kept voice-mobile outside workspace to avoid Metro bundler conflicts
- Disabled React Compiler experimental feature for stability
Setup includes:
- postcss.config.mjs with Tailwind PostCSS plugin
- global.css with Tailwind and NativeWind imports
- metro.config.js with NativeWind wrapper
- lightningcss pinned to 1.30.1 for stability
Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)
Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Major enhancements to voice assistant with conversation persistence
and mobile app planning:
- Add conversation management with multiple active conversations
- Implement conversation persistence (save/load from disk)
- Add conversation selector UI for switching between conversations
- Enhance agent prompt with per-conversation tmux sessions
- Add mobile app technical specification (React Native + Expo)
- Improve UI with active processes view and agent stream
- Add ACP (Agent Control Protocol) server implementation
Mobile spec details:
- React Native with Expo SDK 54+
- NativeWind v5 for styling
- Share Zod schemas/types from server
- Android in-call experience (InCallManager)
- Direct WebSocket connection via Tailscale
- Same audio quality as web prototype
🤖 Generated with [Claude Code](https://claude.com/claude-code)
via [Happy](https://happy.engineering)
Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
This commit implements isolated tmux sessions for each conversation and
refines the terminal MCP tools to use consistent command terminology.
**Per-Conversation Tmux Sessions:**
- Each Session now creates its own Terminal MCP client with conversation ID as tmux session name
- Complete isolation: commands from different conversations don't interfere
- Automatic cleanup: tmux session killed when conversation ends
- Updated getAllTools() to accept custom terminal tools per session
- Session.cleanup() now async to properly await tmux session termination
**Command-Based MCP Tools:**
- Removed old terminal-based tools (list_terminals, create_terminal, etc.)
- Kept only command-based tools (execute_command, list_commands, etc.)
- Updated kill_command to accept multiple command IDs for batch cleanup
- Fixed bug where all commands showed same metadata (working directory, command)
- Added window-specific option storage using set-window-option/show-window-options
**Bug Fixes:**
- Fixed listCommands() showing wrong working directory and command for dead commands
- Stored command metadata as tmux window options for proper retrieval
- Enhanced test to verify multiple commands retain unique metadata
**Testing:**
- Added comprehensive test suite with 26 passing tests
- All tests verify command execution, capture, interaction, and cleanup
- TypeScript: no errors
- Backward compatibility maintained with global singleton fallback
Generated with [Claude Code](https://claude.com/claude-code)
via [Happy](https://happy.engineering)
Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
- Add basic HTTP authentication (express-basic-auth)
- Improve artifact tool with better source type handling (command_output, file, text)
- Update agent prompt with clearer artifact usage patterns
- Refactor tmux command execution to use execFile for safer argument handling
- Add UI cancellation capability for in-progress operations
- Switch session dumps to use util.inspect for better formatting
- Update artifact tool schema to require title parameter first
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
refactor: improve session management, audio handling, and MCP configuration
- Change default session name from 'voice-dev' to '__voice-dev'
- Change default server port from 3000 to 6767
- Add conversation history persistence to .conversations directory
- Defer stream abort until after transcription completes
- Return empty content arrays from MCP tool handlers (use structuredContent only)
- Add pause/resume support to audio player for handling false positive speech detection
- Add .conversations to .gitignore
```
- Update TTSManager to accept an optional abort signal, allowing for graceful cancellation of TTS playback.
- Implement abort handling in the generateAndWaitForPlayback method to clean up pending playbacks and reject promises when playback is aborted.
- Modify Session class to pass the abort signal during TTS generation, improving control over playback interruptions.
These changes improve the robustness of the voice assistant's TTS functionality and enhance user experience by allowing for better management of audio playback.
- Introduce a new `currentStreamPromise` to track ongoing LLM processing, allowing for better handling of stream interruptions and cleanup.
- Update message handling to utilize `ModelMessage` for improved type safety and structure.
- Refactor `processWithLLM` to integrate new streaming capabilities with tool execution and error handling.
- Enhance logging for better traceability during audio playback and message processing.
- Remove deprecated `streamLLM` function and replace it with a more modular approach using `getAllTools`.
These changes improve the robustness of the voice assistant's session management and streamline the integration of LLM processing.
- Introduce a new Session class to encapsulate conversation state and message processing, replacing the previous ClientSession structure.
- Implement STTManager and TTSManager for dedicated handling of speech-to-text and text-to-speech functionalities.
- Refactor WebSocket server to utilize the new session management, improving message routing and state handling.
- Remove deprecated orchestrator functions and streamline message schemas using Zod for validation.
- Enhance audio processing with improved buffering and interruption handling, ensuring seamless user interactions.
This refactor enhances maintainability and prepares the codebase for future feature expansions.
- Remove legacy terminal management code and integrate a new Terminal MCP server.
- Implement TerminalManager for handling terminal operations within a specific tmux session.
- Introduce tools for creating, listing, renaming, and killing terminals with improved error handling.
- Update references and documentation to reflect the new architecture.
- Ensure compatibility with existing features while enhancing maintainability and scalability.
This refactor streamlines terminal operations and prepares the codebase for future enhancements.
feat: add audio buffering and interruption handling for voice input
Implement processing phase tracking (idle/transcribing/llm) and audio segment buffering to handle speech interruptions gracefully. Adds abort_request message type and buffer timeout mechanism to process accumulated audio segments together.
```
chore: configure VAD assets and update Playwright MCP settings
- Add vite-plugin-static-copy to bundle VAD WASM files
- Configure VAD with adjusted speech detection thresholds
- Switch Playwright MCP to omit image responses
- Fix WebSocket abort controller handling for interruptions
- Add WebFetch permission for VAD documentation
```
fix: improve audio playback interruption and response persistence
- Remove TTS timeout to prevent premature playback cancellation
- Save assistant responses to history before waiting for TTS completion
- Stop audio playback when sending message or starting recording
- Properly reject current audio promise when stopping playback
- Ensure partial responses are saved even if stream is aborted
```
Adds error callbacks to LLM orchestrator for tool and stream failures, updates UI to show failed tool calls with error details, removes deprecated tool-definitions.ts file, and updates agent prompt to remove `--dangerously-skip-permissions` flag from Claude Code commands.
Extend capture_terminal to wait for terminal activity to settle before
capturing output. The wait parameter is renamed to maxWait and now triggers
intelligent polling that detects when terminal output has stabilized (1s of
no changes) rather than a simple sleep.
Changes:
- Export waitForPaneActivityToSettle from tmux.ts for reuse
- Update captureTerminal to use settling logic when maxWait provided
- Update capture_terminal tool schema with new maxWait parameter
- Add error handling callbacks to StreamLLMParams interface
- Improve send_text/send_keys tool descriptions for clarity
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
refactor: stream text segments for TTS preparation
Refactor LLM streaming to emit complete text segments before tool calls for future TTS integration. Moves from step-based callbacks to chunk-based streaming with text buffering and segment flushing at tool boundaries.
```
chore: upgrade Vite to v7 and update OpenAI services
- Upgrade vite from ^5.0.0 to ^7.1.10 with new Node.js requirements
- Add dotenv package for environment variable management
- Update STT to use gpt-4o-transcribe model
- Remove verbose_json format from STT response
- Format code with consistent style (quotes, spacing)
- Add HMR WebSocket configuration for development
- Improve UI layout with overflow handling
- Reorganize chat interface component order
- Add user message broadcasting to activity log
```
```
feat: add AI SDK dependencies and refine agent voice interaction config
- Add @ai-sdk packages for enhanced AI capabilities
- Install Deepgram SDK, livekit-client, and werift for voice/WebRTC
- Expand agent system prompt with multi-tool workflows and Claude Code interaction patterns
- Adjust VAD timing (1.2s silence threshold) and turn detection for better voice UX
- Switch TTS to inworld/inworld-tts-1 with Olivia voice
- Configure interruption timing (0.2s detection, no word requirement)
- Add comprehensive guidelines for conversational responses vs structured output
- Document plan mode triggers and Claude response type handling
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
```
feat: refactor MCP tools to use terminal names instead of IDs
- Replace terminalId parameter with terminalName across all tools
- Add window name uniqueness validation for create/rename operations
- Implement terminal name-to-window resolution in all MCP tools
- Add home directory tilde expansion in working directory paths
- Update terminal resources to use encoded names in URIs
- Remove terminal IDs from API responses (list-terminals, resources)
- Add type checking and error handling for Python agent
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
```
```
feat: improve terminal state tracking and creation reliability
- Switch to ElevenLabs TTS (eleven_turbo_v2_5) in voice agent
- Add current command tracking to terminal listings
- Use session prefix '__voice-dev' to avoid conflicts
- Create windows with working directory and run initial commands atomically
- Disable automatic window renaming to preserve user-set names
- Improve working directory detection with lsof fallback
- Enhance command detection to show full command with arguments
```
```
refactor: simplify MCP server to use terminals model instead of tmux hierarchy
Replace complex session/window/pane tools with simplified terminal-focused API:
- Add list-terminals, create-terminal, capture-terminal tools
- Remove list, create-session, create-window, split-pane, kill tools
- Terminals are tmux windows in a default 'voice-dev' session
- Each terminal has contextual working directory
- Extract agent system prompt to external markdown file
- Update CLAUDE.md documentation with new terminal model
This simplifies voice interactions from "create window in session X" to
natural "create a terminal for the web project".
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
```
Delete the existing README.md and mcp-server README.md files to streamline documentation. Introduce CLAUDE.md, which provides comprehensive guidance for the voice-dev project, including architecture, setup instructions, and component interactions.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
```
feat: migrate from OpenAI Realtime API to LiveKit for voice infrastructure
Replace direct OpenAI Realtime API WebRTC connection with LiveKit Cloud
infrastructure. Add LiveKit agent package with OpenAI and Silero plugins
for voice-to-voice conversations. Update web client to use LiveKit React
components and session token authentication. Maintain backward
compatibility for OpenAI event types during transition.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
```
This commit message:
- Uses `feat:` type since this is a new feature/major change
- Describes the main change (migration from OpenAI to LiveKit)
- Provides context in the body about what was added/changed
- Mentions backward compatibility consideration
- Includes the Claude Code attribution as requested
refactor: rename project references from tmux-mcp to voice-dev-mcp
Update internal references, package scripts, binary name, service identifiers, and tool descriptions to reflect the voice-dev branding. Changes default HTTP port from 3000 to 6767.
```