- Added REPO_INVENTORY.md with all repos, branches, remotes, and staging info - Added .gitignore - Synced all existing docs from local workspace - Centralized documentation hub for GrowQR team
12 KiB
GrowQR Integration Problem Analysis
Current Pain Points
1. Hardcoded Routing Maps
The orchestrator has hardcoded PAGE_TO_AGENT and ACTION_TO_AGENT dictionaries in session.py. When new services are added or actions renamed, these maps must be updated, requiring:
- Code changes to orchestrator
- Redeployment of orchestrator
- Coordination between your team (service changes) and original team (orchestrator updates)
2. Redis Streams Complexity
Current event-driven architecture uses Redis Streams with specific naming conventions:
- Stream naming:
tasks:{agent_name} - Response channels:
responses:{task_id} - Message format with specific JSON schema
- Pub/sub completion sentinels (
__task_complete__)
Contributors must understand this protocol, implement consumers, handle backpressure, etc.
3. Dual Transport Maintenance
Services must implement both:
- Redis Stream worker (for streaming/async)
- HTTP A2A endpoints (for synchronous fallback)
This doubles the integration surface area.
4. Contract Breaking on Changes
When you modify microservices:
- Action names change
- Message formats evolve
- New parameters added
- Response structures modified
The original team must update all consumers (orchestrator, other services) to match.
5. No Runtime Tool Registration
Agent cards (/.well-known/agent-card.json) are static JSON files. To add/modify skills:
- Change JSON
- Redeploy service
- Orchestrator must rediscover
There's no dynamic "tool marketplace" where contributors can register capabilities at runtime.
Goals for Improvement
- Minimize breaking changes - Changes to services shouldn't require orchestrator updates
- Simplify contributor integration - New developers should add tools without deep Redis/protocol knowledge
- Dynamic tool management - Frontend-visible tools should be registrable at runtime
- Pluggable interface - SDK-like experience without building a full SDK
- Preserve architecture - Keep the 4-lane model, orchestrator pattern, Q-Score signals
Potential Solution Paths
Path A: Dynamic Tool Registry (Database-Driven)
Replace hardcoded maps with a database table that stores:
- Tool ID, name, description
- Target agent/service URL
- Input schema (JSON Schema)
- Output schema (JSON Schema)
- Routing rules (page patterns, action patterns)
- Version info
Pros:
- Runtime tool registration
- Versioning support
- Frontend can show available tools dynamically
- No orchestrator redeploy for new tools
Cons:
- Still need to define schemas
- Database migration required
- Some complexity in registry UI/management
Path B: HTTP-Only Standardization (Simplify Transport)
Standardize on HTTP A2A for ALL service communication. Remove Redis Streams complexity.
Changes:
- All services implement
/a2a/tasksendpoint - Standardized request/response format
- Streaming via SSE (Server-Sent Events) instead of Redis pub/sub
- Orchestrator routes all traffic via HTTP
Pros:
- Single transport to understand
- HTTP is familiar to all developers
- Easy to test/debug with curl/Postman
- Load balancing built-in
- No Redis consumer complexity
Cons:
- Loses Redis persistence/replay capability
- Slightly higher latency (acceptable?)
- WebSocket to HTTP bridging needed
Path C: Schema Registry + Code Generation
Maintain a central schema registry (like Confluent Schema Registry or simple JSON Schema store):
- Services register their message schemas
- Orchestrator validates messages against schemas
- Code generation for type-safe clients
- Schema evolution rules (backward/forward compatibility)
Pros:
- Strong typing
- Schema evolution managed
- Clear contracts
Cons:
- Complex to implement
- Still requires coordination
- Overhead for simple services
Path D: GraphQL Federation
Replace REST/A2A with GraphQL:
- Single GraphQL gateway (orchestrator becomes gateway)
- Services expose GraphQL schemas
- Frontend queries exactly what it needs
- Natural stitching of services
Pros:
- Single query for multi-service data
- Strong typing
- Introspection = automatic discovery
- Frontend-driven queries
Cons:
- Major architectural shift
- Learning curve for team
- Q-Score signals need separate path
- Breaking change to entire system
Path E: Plugin Architecture with Self-Registration
Services self-register at startup via a "Plugin Registry" API:
# In service startup
await plugin_registry.register({
"service_name": "interview-service",
"version": "2.0.0",
"tools": [
{
"tool_id": "configure_interview",
"name": "Configure Interview",
"endpoint": "/a2a/tasks",
"input_schema": {...},
"triggers": ["page:interview", "action:configure_interview"],
"outputs": ["interview_configured", "qscore_signals"]
}
]
})
Pros:
- Services describe themselves
- No hardcoded maps
- Runtime discovery
- Can include health checks
Cons:
- Services must implement registration logic
- Registration order matters
- Need deregistration/cleanup
Path F: OpenAPI-Based Dynamic Client Generation
Services expose OpenAPI specs at /.well-known/openapi.json:
- Orchestrator fetches specs
- Generates client code dynamically (or uses generic client)
- Routes based on operationIds
- Validates requests/responses
Pros:
- Industry standard
- Auto-generated docs
- Code generation tools exist
- Language agnostic
Cons:
- Still need routing logic
- OpenAPI can be verbose
- Doesn't solve the action/page routing problem
Path G: gRPC with Reflection
Replace HTTP A2A with gRPC:
- Services define .proto files
- gRPC reflection for discovery
- Strongly typed, efficient
- Streaming support built-in
Pros:
- Type safety
- Performance
- Streaming
- Reflection = discovery
Cons:
- Major architectural change
- Frontend needs gRPC-web or gateway
- Learning curve
- Overkill for simple CRUD
Recommended Hybrid Approach
Based on the constraints (minimal breaking changes, contributor-friendly, pluggable), I recommend a hybrid approach combining Path A + Path B + Path E:
Phase 1: Dynamic Tool Registry (Immediate Win)
Create a Tool Registry Service (or add to orchestrator):
# Tool Registry Schema
tool_registry = {
"tool_id": "resume_analyze",
"name": "AI Resume Analysis",
"description": "Analyzes resume for ATS compatibility",
"service_name": "resume-builder", # Maps to agent name
"service_url": "http://resume-builder:8000",
"triggers": {
"pages": ["resume-hub", "resume-editor"],
"actions": ["ai_analyze", "analyze_resume"],
"skills": ["resume.analysis"]
},
"input_schema": {
"type": "object",
"properties": {
"resume_id": {"type": "string"},
"deep_analysis": {"type": "boolean"}
},
"required": ["resume_id"]
},
"output_schema": {
"type": "object",
"properties": {
"score": {"type": "number"},
"suggestions": {"type": "array"}
}
},
"qscore_signals": ["resume.ats_score", "resume.completeness"],
"version": "2.1.0",
"enabled": True,
"created_by": "contributor-xyz",
"created_at": "2024-01-15"
}
Changes to Orchestrator:
- Replace hardcoded
PAGE_TO_AGENTandACTION_TO_AGENTwith registry lookups - Cache registry in memory with periodic refresh
- New admin endpoints for tool CRUD
- Frontend can query available tools for a page
Phase 2: HTTP-First with Optional Redis (Simplify Transport)
Standardize on HTTP A2A as primary transport:
- Make HTTP A2A the default (already implemented)
- Deprecate Redis Streams for new services (keep for existing)
- Add SSE streaming support for real-time updates
- Contributors only need to implement:
POST /a2a/tasks
Service Template for Contributors:
# Minimal service template
from fastapi import FastAPI
from pydantic import BaseModel
app = FastAPI()
class TaskRequest(BaseModel):
task_id: str
action: str
params: dict
user_id: str
user_context: dict | None = None
@app.post("/a2a/tasks")
async def handle_task(request: TaskRequest):
# Route to appropriate handler
if request.action == "my_tool_action":
result = await my_tool.execute(request.params, request.user_context)
return {
"task_id": request.task_id,
"status": "completed",
"messages": [{
"type": "agent_data",
"action": "my_tool_complete",
"data": result
}]
}
Phase 3: Self-Registration with Health Checks (Plugin Architecture)
Services self-register on startup:
# Service startup
async def register_tools():
async with httpx.AsyncClient() as client:
await client.post(
"http://orchestrator:8000/api/v1/tools/register",
json={
"service_name": "my-new-service",
"tools": [
{
"tool_id": "my_tool",
"triggers": {
"pages": ["my-page"],
"actions": ["my_action"]
},
"endpoint": "/a2a/tasks",
"health_check": "/health"
}
]
},
headers={"Authorization": f"Bearer {REGISTRATION_TOKEN}"}
)
Orchestrator enhancements:
- Accept dynamic registrations
- Health check polling
- Automatic deregistration on unhealthy services
- Tool versioning (multiple versions of same tool)
Phase 4: Frontend Tool Marketplace
Allow dynamic tool enabling from frontend:
// Admin interface
interface ToolConfig {
toolId: string;
enabled: boolean;
visibleToUsers: boolean;
requirePlan: 'free' | 'pro' | 'enterprise';
customParams: Record<string, any>;
}
// Users see tools based on registry + their plan
const availableTools = await fetch('/api/v1/tools/available?page=resume-hub');
Implementation Roadmap (Minimal Breaking Changes)
Week 1-2: Registry Foundation
- Create
ToolRegistryclass in orchestrator (in-memory first) - Add
/api/v1/toolsCRUD endpoints - Load hardcoded maps into registry at startup
- Switch routing to use registry lookups
Breaking change: None - just refactoring
Week 3-4: Service Self-Registration
- Add registration endpoint to orchestrator
- Create simple registration helper library
- Update one service (e.g., interview-service) to self-register
- Test end-to-end
Breaking change: None for existing services
Week 5-6: Frontend Integration
- Add tool discovery API:
/api/v1/tools?page={page} - Frontend fetches available tools per page
- Dynamic UI rendering based on tool registry
- Admin interface for tool management
Breaking change: None - additive feature
Week 7-8: HTTP-First Migration
- Ensure all services have HTTP A2A endpoints
- Switch orchestrator default to HTTP
- Add SSE streaming for real-time updates
- Document contributor template
Breaking change: Redis becomes optional (not default)
Week 9-10: Contributor SDK (Lightweight)
- Python decorator library:
from growqr_sdk import tool, register_service
@tool(
tool_id="analyze_resume",
triggers=["page:resume-hub", "action:analyze"]
)
async def analyze_resume(params, user_context):
# Implementation
return {"score": 85}
register_service(
name="my-service",
orchestrator_url="http://orchestrator:8000",
tools=[analyze_resume]
)
- JavaScript/TypeScript equivalent for Node services
Summary: Why This Works
- Minimal breaking changes: Existing services keep working, Redis optional
- Contributor-friendly: Single HTTP endpoint to implement, decorator SDK optional
- Pluggable: Runtime registration, dynamic UI, no orchestrator redeploy 4 Preserves architecture: Still 4-lane, still orchestrated, still Q-Score signals 5 Your team benefits: You change services independently, original team gets stability 6 Original team benefits: Less integration code, dynamic capabilities, easier debugging
The key insight: The orchestrator should ask "what tools are available?" not "is this action in my hardcoded map?"