Update handleSystemMessage() in claude-agent.ts to read the actual model from the SDK's init message (message.model) instead of just echoing back the configured model. This gives us the real runtime model that the SDK selected. Also invalidates cached runtime info when the model is updated to ensure getRuntimeInfo() returns the new value. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
6.7 KiB
Plan
Context
Voice-controlled terminal assistant using OpenAI's Realtime API. Monorepo with Express backend and Expo cross-platform app.
Focus: Agent model info tracking. Two distinct concepts:
- Configured model: What the user requested (default or specific model)
- Runtime model: What the agent is actually using (from the agent process itself)
Hard requirement: We must get the actual runtime model, not just echo back the requested config.
Environment
- Expo app: Running in tmux session
moboudra:mobile- check logs withtmux capture-pane -t moboudra:mobile -p - Server: Running in tmux session
moboudra:server- check logs withtmux capture-pane -t moboudra:server -p - Web testing: Use Playwright MCP at
http://localhost:8081
Guiding Principles
- Keep changes minimal and focused
- Run typecheck after every change
- Don't break existing functionality
- Test in the running app when possible
Testing Requirements (CRITICAL)
- Nothing is done until tested. Every implementation task must be followed by a testing task using Playwright MCP.
- Test tasks verify specific behaviors. Not "test the feature" but "verify X does Y when Z".
- Test failures spawn fix tasks. If a test finds issues, add fix tasks immediately after.
- Fix tasks get their own test tasks. After a fix task, add a re-test task to verify the fix.
- Loop until it works. The cycle is: implement → test → fix → re-test → ... until verified working.
- Testers own the plan. Test tasks can and should add new tasks (fixes, re-tests) to keep the loop going.
- Never claim done without verification. If you can't test it with Playwright MCP, you can't mark it complete.
Tasks
-
Plan: Audit current model info implementation and expand this plan.
- Find where model config is set when creating agents.
- Find where runtime model info is fetched/displayed.
- Document how each provider (Claude SDK, Codex SDK) exposes actual model info.
- Check if we're currently showing requested model vs actual runtime model.
- Add implementation tasks based on findings.
- Add test tasks as needed.
- Add another Plan task at an appropriate position to re-audit after some progress.
- Done (2025-12-21 12:45): Completed full audit. See findings below.
Audit Findings: Model Info Implementation
Current Architecture (Working Well):
- Model config is set via
AgentSessionConfig.modelinagent-sdk-types.ts:148-164 - Runtime info uses
AgentRuntimeInfotype withmodelfield inagent-sdk-types.ts:126-132 - UI correctly uses
extractAgentModel()which readsruntimeInfo.model(not config) - Displayed in agent screen header at line 555 of
[agentId].tsx
Claude SDK:
claude-agent.ts:423-435-getRuntimeInfo()returnsthis.lastOptionsModel- Model comes from Claude SDK options passed during agent creation
- Currently echoes back the configured model (not truly "runtime detected")
Codex SDK:
codex-agent.ts:407-426-getRuntimeInfo()reads actual runtime modelcodex-agent.ts:1229-1261- Parses rollout file forturn_context.payload.model- This IS the actual runtime model (correctly implemented)
- Falls back to config model if rollout not available yet
Gap Identified:
- Claude agents: Model info is the configured model, not detected from Claude SDK
- Codex agents: Model info IS detected from runtime (via rollout file) ✓
- Claude SDK may expose actual model in response metadata - needs investigation
Data Flow:
- Config:
config.model→ agent creation - Runtime:
agent.session.getRuntimeInfo()called after stream/turn completion - Storage:
runtimeInfopersisted in agent registry - Client: Receives via WebSocket, displayed via
extractAgentModel()
-
Investigate: Check if Claude SDK exposes actual model in response/stream metadata.
- Look at
@anthropic-ai/claude-agent-sdktypes for model info in responses - Check if
query()returns model information - Document findings for implementation task
- Done (2025-12-21 13:15): YES, the Claude SDK exposes actual model info in TWO places:
SDKSystemMessage(subtype: "init") hasmodel: stringfield (sdk.d.ts:334)SDKResultMessagehasmodelUsage: { [modelName: string]: ModelUsage }(sdk.d.ts:299-300)
- Current implementation ignores both - only uses
config.model(claude-agent.ts:712) - Fix: Update
handleSystemMessage()to capturemessage.modelaslastOptionsModel - Alternatively: Extract model name from
modelUsagekeys inSDKResultMessage
- Look at
-
Implement: Capture actual model from Claude SDK init message.
- In
handleSystemMessage(), setthis.lastOptionsModel = message.model - This gives us the actual model the SDK selected (not just our config)
- Simpler than parsing modelUsage, same result
- Done (2025-12-21 13:25): Updated
handleSystemMessage()inclaude-agent.ts:868-882to capturemessage.modelfrom the SDK init message and invalidate cached runtime info. Typecheck passes.
- In
-
Test: Verify current Codex runtime model detection works.
- Create a Codex agent with default model
- Wait for first turn to complete
- Verify the model displayed matches actual runtime model (e.g.,
gpt-4.1) - Check that it's not just echoing configured model
-
Test: Verify Claude agent model display behavior.
- Create a Claude agent with default model
- Wait for first turn to complete
- Check what model is displayed
- Document whether it's configured or runtime model
-
Plan: Re-audit after investigation and initial tests complete.
- Review test results
- Determine if Claude SDK exposes runtime model info
- Add implementation tasks if improvements needed
- Add fix tasks if tests reveal issues
-
Plan: Design and implement agent parent/child hierarchy.
- Add
parentIdfield to agents. - Agents created via MCP should auto-set parentId to the calling agent.
- Homepage should only show top-level agents (no parentId).
- Agent screen three-dot menu should show sub-agents of that agent.
- Sub-agents should be navigable from the menu.
- Ensure back button works properly when navigating agent hierarchy.
- Add implementation tasks based on findings.
- Add another Plan task at an appropriate position to re-audit after some progress.
- Add
-
agent=codex Review: Code quality and types review.
- Review all changed files for code quality issues.
- Check TypeScript types are correct and complete.
- Look for any type errors or unsafe casts.
- Check for proper error handling.
- Add fix tasks for any issues found.
- Add another
agent=codex **Review**task after fix tasks to verify fixes.