feat: initial docs repo with project inventory and all documentation
- Added REPO_INVENTORY.md with all repos, branches, remotes, and staging info - Added .gitignore - Synced all existing docs from local workspace - Centralized documentation hub for GrowQR team
This commit is contained in:
403
integration-problem-analysis.md
Normal file
403
integration-problem-analysis.md
Normal file
@@ -0,0 +1,403 @@
|
||||
# GrowQR Integration Problem Analysis
|
||||
|
||||
## Current Pain Points
|
||||
|
||||
### 1. Hardcoded Routing Maps
|
||||
The orchestrator has hardcoded `PAGE_TO_AGENT` and `ACTION_TO_AGENT` dictionaries in `session.py`. When new services are added or actions renamed, these maps must be updated, requiring:
|
||||
- Code changes to orchestrator
|
||||
- Redeployment of orchestrator
|
||||
- Coordination between your team (service changes) and original team (orchestrator updates)
|
||||
|
||||
### 2. Redis Streams Complexity
|
||||
Current event-driven architecture uses Redis Streams with specific naming conventions:
|
||||
- Stream naming: `tasks:{agent_name}`
|
||||
- Response channels: `responses:{task_id}`
|
||||
- Message format with specific JSON schema
|
||||
- Pub/sub completion sentinels (`__task_complete__`)
|
||||
|
||||
Contributors must understand this protocol, implement consumers, handle backpressure, etc.
|
||||
|
||||
### 3. Dual Transport Maintenance
|
||||
Services must implement both:
|
||||
- Redis Stream worker (for streaming/async)
|
||||
- HTTP A2A endpoints (for synchronous fallback)
|
||||
|
||||
This doubles the integration surface area.
|
||||
|
||||
### 4. Contract Breaking on Changes
|
||||
When you modify microservices:
|
||||
- Action names change
|
||||
- Message formats evolve
|
||||
- New parameters added
|
||||
- Response structures modified
|
||||
|
||||
The original team must update all consumers (orchestrator, other services) to match.
|
||||
|
||||
### 5. No Runtime Tool Registration
|
||||
Agent cards (`/.well-known/agent-card.json`) are static JSON files. To add/modify skills:
|
||||
- Change JSON
|
||||
- Redeploy service
|
||||
- Orchestrator must rediscover
|
||||
|
||||
There's no dynamic "tool marketplace" where contributors can register capabilities at runtime.
|
||||
|
||||
## Goals for Improvement
|
||||
|
||||
1. **Minimize breaking changes** - Changes to services shouldn't require orchestrator updates
|
||||
2. **Simplify contributor integration** - New developers should add tools without deep Redis/protocol knowledge
|
||||
3. **Dynamic tool management** - Frontend-visible tools should be registrable at runtime
|
||||
4. **Pluggable interface** - SDK-like experience without building a full SDK
|
||||
5. **Preserve architecture** - Keep the 4-lane model, orchestrator pattern, Q-Score signals
|
||||
|
||||
## Potential Solution Paths
|
||||
|
||||
### Path A: Dynamic Tool Registry (Database-Driven)
|
||||
Replace hardcoded maps with a database table that stores:
|
||||
- Tool ID, name, description
|
||||
- Target agent/service URL
|
||||
- Input schema (JSON Schema)
|
||||
- Output schema (JSON Schema)
|
||||
- Routing rules (page patterns, action patterns)
|
||||
- Version info
|
||||
|
||||
**Pros:**
|
||||
- Runtime tool registration
|
||||
- Versioning support
|
||||
- Frontend can show available tools dynamically
|
||||
- No orchestrator redeploy for new tools
|
||||
|
||||
**Cons:**
|
||||
- Still need to define schemas
|
||||
- Database migration required
|
||||
- Some complexity in registry UI/management
|
||||
|
||||
### Path B: HTTP-Only Standardization (Simplify Transport)
|
||||
Standardize on HTTP A2A for ALL service communication. Remove Redis Streams complexity.
|
||||
|
||||
**Changes:**
|
||||
- All services implement `/a2a/tasks` endpoint
|
||||
- Standardized request/response format
|
||||
- Streaming via SSE (Server-Sent Events) instead of Redis pub/sub
|
||||
- Orchestrator routes all traffic via HTTP
|
||||
|
||||
**Pros:**
|
||||
- Single transport to understand
|
||||
- HTTP is familiar to all developers
|
||||
- Easy to test/debug with curl/Postman
|
||||
- Load balancing built-in
|
||||
- No Redis consumer complexity
|
||||
|
||||
**Cons:**
|
||||
- Loses Redis persistence/replay capability
|
||||
- Slightly higher latency (acceptable?)
|
||||
- WebSocket to HTTP bridging needed
|
||||
|
||||
### Path C: Schema Registry + Code Generation
|
||||
Maintain a central schema registry (like Confluent Schema Registry or simple JSON Schema store):
|
||||
- Services register their message schemas
|
||||
- Orchestrator validates messages against schemas
|
||||
- Code generation for type-safe clients
|
||||
- Schema evolution rules (backward/forward compatibility)
|
||||
|
||||
**Pros:**
|
||||
- Strong typing
|
||||
- Schema evolution managed
|
||||
- Clear contracts
|
||||
|
||||
**Cons:**
|
||||
- Complex to implement
|
||||
- Still requires coordination
|
||||
- Overhead for simple services
|
||||
|
||||
### Path D: GraphQL Federation
|
||||
Replace REST/A2A with GraphQL:
|
||||
- Single GraphQL gateway (orchestrator becomes gateway)
|
||||
- Services expose GraphQL schemas
|
||||
- Frontend queries exactly what it needs
|
||||
- Natural stitching of services
|
||||
|
||||
**Pros:**
|
||||
- Single query for multi-service data
|
||||
- Strong typing
|
||||
- Introspection = automatic discovery
|
||||
- Frontend-driven queries
|
||||
|
||||
**Cons:**
|
||||
- Major architectural shift
|
||||
- Learning curve for team
|
||||
- Q-Score signals need separate path
|
||||
- Breaking change to entire system
|
||||
|
||||
### Path E: Plugin Architecture with Self-Registration
|
||||
Services self-register at startup via a "Plugin Registry" API:
|
||||
```python
|
||||
# In service startup
|
||||
await plugin_registry.register({
|
||||
"service_name": "interview-service",
|
||||
"version": "2.0.0",
|
||||
"tools": [
|
||||
{
|
||||
"tool_id": "configure_interview",
|
||||
"name": "Configure Interview",
|
||||
"endpoint": "/a2a/tasks",
|
||||
"input_schema": {...},
|
||||
"triggers": ["page:interview", "action:configure_interview"],
|
||||
"outputs": ["interview_configured", "qscore_signals"]
|
||||
}
|
||||
]
|
||||
})
|
||||
```
|
||||
|
||||
**Pros:**
|
||||
- Services describe themselves
|
||||
- No hardcoded maps
|
||||
- Runtime discovery
|
||||
- Can include health checks
|
||||
|
||||
**Cons:**
|
||||
- Services must implement registration logic
|
||||
- Registration order matters
|
||||
- Need deregistration/cleanup
|
||||
|
||||
### Path F: OpenAPI-Based Dynamic Client Generation
|
||||
Services expose OpenAPI specs at `/.well-known/openapi.json`:
|
||||
- Orchestrator fetches specs
|
||||
- Generates client code dynamically (or uses generic client)
|
||||
- Routes based on operationIds
|
||||
- Validates requests/responses
|
||||
|
||||
**Pros:**
|
||||
- Industry standard
|
||||
- Auto-generated docs
|
||||
- Code generation tools exist
|
||||
- Language agnostic
|
||||
|
||||
**Cons:**
|
||||
- Still need routing logic
|
||||
- OpenAPI can be verbose
|
||||
- Doesn't solve the action/page routing problem
|
||||
|
||||
### Path G: gRPC with Reflection
|
||||
Replace HTTP A2A with gRPC:
|
||||
- Services define .proto files
|
||||
- gRPC reflection for discovery
|
||||
- Strongly typed, efficient
|
||||
- Streaming support built-in
|
||||
|
||||
**Pros:**
|
||||
- Type safety
|
||||
- Performance
|
||||
- Streaming
|
||||
- Reflection = discovery
|
||||
|
||||
**Cons:**
|
||||
- Major architectural change
|
||||
- Frontend needs gRPC-web or gateway
|
||||
- Learning curve
|
||||
- Overkill for simple CRUD
|
||||
|
||||
## Recommended Hybrid Approach
|
||||
|
||||
Based on the constraints (minimal breaking changes, contributor-friendly, pluggable), I recommend a **hybrid approach combining Path A + Path B + Path E**:
|
||||
|
||||
### Phase 1: Dynamic Tool Registry (Immediate Win)
|
||||
|
||||
**Create a Tool Registry Service (or add to orchestrator):**
|
||||
|
||||
```python
|
||||
# Tool Registry Schema
|
||||
tool_registry = {
|
||||
"tool_id": "resume_analyze",
|
||||
"name": "AI Resume Analysis",
|
||||
"description": "Analyzes resume for ATS compatibility",
|
||||
"service_name": "resume-builder", # Maps to agent name
|
||||
"service_url": "http://resume-builder:8000",
|
||||
"triggers": {
|
||||
"pages": ["resume-hub", "resume-editor"],
|
||||
"actions": ["ai_analyze", "analyze_resume"],
|
||||
"skills": ["resume.analysis"]
|
||||
},
|
||||
"input_schema": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"resume_id": {"type": "string"},
|
||||
"deep_analysis": {"type": "boolean"}
|
||||
},
|
||||
"required": ["resume_id"]
|
||||
},
|
||||
"output_schema": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"score": {"type": "number"},
|
||||
"suggestions": {"type": "array"}
|
||||
}
|
||||
},
|
||||
"qscore_signals": ["resume.ats_score", "resume.completeness"],
|
||||
"version": "2.1.0",
|
||||
"enabled": True,
|
||||
"created_by": "contributor-xyz",
|
||||
"created_at": "2024-01-15"
|
||||
}
|
||||
```
|
||||
|
||||
**Changes to Orchestrator:**
|
||||
1. Replace hardcoded `PAGE_TO_AGENT` and `ACTION_TO_AGENT` with registry lookups
|
||||
2. Cache registry in memory with periodic refresh
|
||||
3. New admin endpoints for tool CRUD
|
||||
4. Frontend can query available tools for a page
|
||||
|
||||
### Phase 2: HTTP-First with Optional Redis (Simplify Transport)
|
||||
|
||||
**Standardize on HTTP A2A as primary transport:**
|
||||
1. Make HTTP A2A the default (already implemented)
|
||||
2. Deprecate Redis Streams for new services (keep for existing)
|
||||
3. Add SSE streaming support for real-time updates
|
||||
4. Contributors only need to implement: `POST /a2a/tasks`
|
||||
|
||||
**Service Template for Contributors:**
|
||||
```python
|
||||
# Minimal service template
|
||||
from fastapi import FastAPI
|
||||
from pydantic import BaseModel
|
||||
|
||||
app = FastAPI()
|
||||
|
||||
class TaskRequest(BaseModel):
|
||||
task_id: str
|
||||
action: str
|
||||
params: dict
|
||||
user_id: str
|
||||
user_context: dict | None = None
|
||||
|
||||
@app.post("/a2a/tasks")
|
||||
async def handle_task(request: TaskRequest):
|
||||
# Route to appropriate handler
|
||||
if request.action == "my_tool_action":
|
||||
result = await my_tool.execute(request.params, request.user_context)
|
||||
return {
|
||||
"task_id": request.task_id,
|
||||
"status": "completed",
|
||||
"messages": [{
|
||||
"type": "agent_data",
|
||||
"action": "my_tool_complete",
|
||||
"data": result
|
||||
}]
|
||||
}
|
||||
```
|
||||
|
||||
### Phase 3: Self-Registration with Health Checks (Plugin Architecture)
|
||||
|
||||
**Services self-register on startup:**
|
||||
```python
|
||||
# Service startup
|
||||
async def register_tools():
|
||||
async with httpx.AsyncClient() as client:
|
||||
await client.post(
|
||||
"http://orchestrator:8000/api/v1/tools/register",
|
||||
json={
|
||||
"service_name": "my-new-service",
|
||||
"tools": [
|
||||
{
|
||||
"tool_id": "my_tool",
|
||||
"triggers": {
|
||||
"pages": ["my-page"],
|
||||
"actions": ["my_action"]
|
||||
},
|
||||
"endpoint": "/a2a/tasks",
|
||||
"health_check": "/health"
|
||||
}
|
||||
]
|
||||
},
|
||||
headers={"Authorization": f"Bearer {REGISTRATION_TOKEN}"}
|
||||
)
|
||||
```
|
||||
|
||||
**Orchestrator enhancements:**
|
||||
- Accept dynamic registrations
|
||||
- Health check polling
|
||||
- Automatic deregistration on unhealthy services
|
||||
- Tool versioning (multiple versions of same tool)
|
||||
|
||||
### Phase 4: Frontend Tool Marketplace
|
||||
|
||||
**Allow dynamic tool enabling from frontend:**
|
||||
```typescript
|
||||
// Admin interface
|
||||
interface ToolConfig {
|
||||
toolId: string;
|
||||
enabled: boolean;
|
||||
visibleToUsers: boolean;
|
||||
requirePlan: 'free' | 'pro' | 'enterprise';
|
||||
customParams: Record<string, any>;
|
||||
}
|
||||
|
||||
// Users see tools based on registry + their plan
|
||||
const availableTools = await fetch('/api/v1/tools/available?page=resume-hub');
|
||||
```
|
||||
|
||||
## Implementation Roadmap (Minimal Breaking Changes)
|
||||
|
||||
### Week 1-2: Registry Foundation
|
||||
1. Create `ToolRegistry` class in orchestrator (in-memory first)
|
||||
2. Add `/api/v1/tools` CRUD endpoints
|
||||
3. Load hardcoded maps into registry at startup
|
||||
4. Switch routing to use registry lookups
|
||||
|
||||
**Breaking change:** None - just refactoring
|
||||
|
||||
### Week 3-4: Service Self-Registration
|
||||
1. Add registration endpoint to orchestrator
|
||||
2. Create simple registration helper library
|
||||
3. Update one service (e.g., interview-service) to self-register
|
||||
4. Test end-to-end
|
||||
|
||||
**Breaking change:** None for existing services
|
||||
|
||||
### Week 5-6: Frontend Integration
|
||||
1. Add tool discovery API: `/api/v1/tools?page={page}`
|
||||
2. Frontend fetches available tools per page
|
||||
3. Dynamic UI rendering based on tool registry
|
||||
4. Admin interface for tool management
|
||||
|
||||
**Breaking change:** None - additive feature
|
||||
|
||||
### Week 7-8: HTTP-First Migration
|
||||
1. Ensure all services have HTTP A2A endpoints
|
||||
2. Switch orchestrator default to HTTP
|
||||
3. Add SSE streaming for real-time updates
|
||||
4. Document contributor template
|
||||
|
||||
**Breaking change:** Redis becomes optional (not default)
|
||||
|
||||
### Week 9-10: Contributor SDK (Lightweight)
|
||||
1. Python decorator library:
|
||||
```python
|
||||
from growqr_sdk import tool, register_service
|
||||
|
||||
@tool(
|
||||
tool_id="analyze_resume",
|
||||
triggers=["page:resume-hub", "action:analyze"]
|
||||
)
|
||||
async def analyze_resume(params, user_context):
|
||||
# Implementation
|
||||
return {"score": 85}
|
||||
|
||||
register_service(
|
||||
name="my-service",
|
||||
orchestrator_url="http://orchestrator:8000",
|
||||
tools=[analyze_resume]
|
||||
)
|
||||
```
|
||||
|
||||
2. JavaScript/TypeScript equivalent for Node services
|
||||
|
||||
## Summary: Why This Works
|
||||
|
||||
1. **Minimal breaking changes**: Existing services keep working, Redis optional
|
||||
2. **Contributor-friendly**: Single HTTP endpoint to implement, decorator SDK optional
|
||||
3. **Pluggable**: Runtime registration, dynamic UI, no orchestrator redeploy
|
||||
4 **Preserves architecture**: Still 4-lane, still orchestrated, still Q-Score signals
|
||||
5 **Your team benefits**: You change services independently, original team gets stability
|
||||
6 **Original team benefits**: Less integration code, dynamic capabilities, easier debugging
|
||||
|
||||
The key insight: **The orchestrator should ask "what tools are available?" not "is this action in my hardcoded map?"**
|
||||
Reference in New Issue
Block a user