Files
zopu-code/docs/TECH.md
-Puter d8383a788e feat: Slice 1 polish - MiMo V2.5 model, visible reasoning, image attachments, mobile keyboard fix
- Switch conversation agent to xiaomi/mimo-v2.5 (multimodal: text + image)
- Render native reasoning parts as live 'Thinking trace' (streaming open,
  collapsed after completion); inline <think> extraction for streaming models
- Image attachments: picker (up to 4, 10MB each), base64 to Flue
  AgentPromptImage, authenticated blob-URL replay for historical images
- Mobile keyboard viewport fix: visual-viewport hook, fixed shell,
  interactive-widget=resizes-content, header pinned, composer follows keyboard
- Conversation to Signal to proposed Work: Convex persistence, Effect
  validation in @code/work-os, Work cards with exact source provenance
- Streamdown markdown + Mermaid chart rendering in chat messages
- Flue tool turns hidden, reasoning-containing turns remain visible
- Frontend regression tests: keyboard viewport, responsive shell,
  attachment overflow, authenticated images, reasoning traces, transforms
- .env.example updated to xiaomi/mimo-v2.5 config
2026-07-27 12:22:54 +05:30

664 lines
15 KiB
Markdown

# Zopu Work OS — Technical Architecture
> **Related:** `agent-context.md` defines agent rules; `dev-loop.md` defines orchestration; `glossary.md` defines terms.
> **Status:** Working technical source of truth
> **Audience:** engineers and implementation agents
> **Read after:** `product.md`
> **Principle:** durable intent/state is separate from disposable execution.
## 1. System topology
```text
Web / Buzz / integrations
Application API + Convex durable data
Rivet actor system
├── ProjectActor
├── WorkActor
├── AttemptActor
├── VerificationActor (later split)
├── IntegrationActor (later)
└── ResultActor (later)
FLUE orchestration/application agents
├── Signal/Definition
├── Design/Slice planning
├── Resolver decisions
├── Verification design/evaluation
└── Learning synthesis
Ports
├── HarnessRuntime
├── SandboxRuntime
├── SourceControl
├── ArtifactStore
├── PreviewRuntime
├── SecretBroker
├── EventJournal
└── RuntimePolicy
Adapters
├── OMP/OpenCode/Codex/Pi/custom FLUE
├── CubeSandbox/AgentOS/persistent machine/Docker
├── Git forge
├── object storage
└── deployment/preview providers
```
### Ownership
| Layer | Owns |
|---|---|
| Work OS | durable intent, lifecycle, evidence, policy |
| Rivet actors | serialized ownership, recovery, timers, leases |
| FLUE | programmable orchestration and domain-specific agents |
| Harness | bounded coding/tool loop |
| Sandbox/runtime | filesystem, processes, network, isolation |
| Git | source revision history |
| Artifact store | durable outputs/evidence |
| External systems | collaboration, delivery, monitoring |
No harness, sandbox, or chat transcript owns Work state.
## 2. Domain boundaries
Recommended packages:
```text
packages/
├── signals/
├── work/
├── design/
├── planning/
├── kits/
├── resolver/
├── verification/
├── artifacts/
├── results/
├── knowledge/
├── runtimes/
├── harnesses/
└── integrations/
```
Each domain SHOULD separate:
```text
domain/ pure state, schemas, invariants
application/ use cases and orchestration
ports/ Effect services
adapters/ infrastructure implementations
```
Avoid package proliferation in v0; logical boundaries may begin as folders.
## 3. Core records
Minimal durable model:
```ts
type Id = string
interface Message {
id: Id
projectId: Id
content: string
createdAt: number
}
interface Signal {
id: Id
projectId: Id
sourceType: string
sourceId: string
sourcePayloadRef?: string
summary: string
fingerprint: string
status: "candidate" | "accepted" | "dismissed"
}
interface Work {
id: Id
projectId: Id
title: string
objective: string
risk: "low" | "medium" | "high"
status: WorkStatus
definitionVersion?: number
designVersion?: number
createdAt: number
updatedAt: number
}
interface Step {
id: Id
workId: Id
sliceId?: Id
kind: "design" | "implement" | "verify" | "integrate" | "publish" | "observe"
objective: string
dependsOn: readonly Id[]
status: StepStatus
}
interface Run {
id: Id
workId: Id
stepId: Id
kitVersion: string
status: RunStatus
}
interface Attempt {
id: Id
runId: Id
number: number
harness: string
runtime: string
sourceRevision: string
status: AttemptStatus
startedAt?: number
endedAt?: number
}
interface Artifact {
id: Id
workId: Id
stepId?: Id
runId?: Id
attemptId?: Id
type: string
uri?: string
contentHash?: string
sourceRevision?: string
environmentId?: string
metadata: unknown
}
interface Question {
id: Id
workId: Id
stepId?: Id
attemptId?: Id
prompt: string
recommendation?: string
alternatives: readonly string[]
status: "open" | "answered" | "withdrawn"
answer?: string
}
```
All external/model payloads MUST be decoded with schemas at boundaries.
## 4. Work state machine
```text
Proposed
→ Defining
→ AwaitingDefinitionApproval
→ Designing
→ AwaitingDesignApproval
→ Ready
→ ExecutingSlice
→ VerifyingSlice
→ AwaitingSliceReview
→ IntegratedVerification
→ Review
→ Releasing
→ Observing
→ Completed
```
Side/terminal states:
```text
NeedsInput | Blocked | Replanning | Failed | Cancelled
```
Transitions require explicit commands plus evidence. State MUST NOT be inferred from the latest text message.
## 5. Actor model
Start small.
### ProjectActor
Owns:
- project configuration;
- repository/runtime policies;
- active Work index;
- integration endpoints;
- project-level Signal routing.
### WorkActor
Initial central aggregate:
- Work Definition and versions;
- Design Packet and versions;
- slice/step graph;
- Resolver state;
- approvals/questions;
- artifact references;
- result.
It serializes lifecycle transitions and commands.
### AttemptActor
Owns one execution attempt:
- runtime lease/sandbox ID;
- harness session;
- scoped credentials;
- event stream/checkpoints;
- cancellation;
- attempt timeout;
- normalized outcome.
### VerificationActor
Split from WorkActor after v0 verification is stable. Owns plan, checks, evidence, verdict.
### IntegrationActor
Added when multiple slices/branches exist. Owns branch composition, conflicts, integrated revision, integrated checks.
### ResultActor
Added after delivery observation exists. Owns rollout observation, original success signals, final outcome, learning proposals.
Do not create an actor per document or tool call. Create actors for durable identity, serialized ownership, independent failure/recovery, or timers.
## 6. Commands, events, and idempotency
Use command/event vocabulary rather than mutable agent prose.
Example commands:
```text
ProcessMessage
AcceptSignal
CreateWork
ApproveDefinition
ApproveDesign
StartNextSlice
RecordHarnessEvent
CompleteAttempt
StartVerification
RecordVerificationCheck
CreateRepairAttempt
PublishPullRequest
RecordHumanDecision
CompleteWork
```
Example events:
```text
SignalCreated
SignalAttached
WorkProposed
DefinitionGenerated
DefinitionApproved
DesignGenerated
DesignApproved
SliceStarted
AttemptStarted
AttemptCompleted
VerificationPassed
VerificationFailed
QuestionOpened
QuestionAnswered
PullRequestCreated
WorkCompleted
```
Every side-effecting command MUST include an idempotency key. Suggested key:
```text
<workId>:<commandType>:<logicalVersion>
```
External artifact creation stores provider ID before transition completion to prevent duplicate PRs/deployments.
## 7. Effect service boundaries
Keep domain/application code provider-neutral.
```ts
interface HarnessRuntime {
open(input: OpenHarnessInput): Effect.Effect<HarnessSession, HarnessError>
prompt(id: string, content: string): Effect.Effect<void, HarnessError>
events(id: string): Stream.Stream<HarnessEvent, HarnessError>
approve(input: PermissionDecision): Effect.Effect<void, HarnessError>
abort(id: string): Effect.Effect<void, HarnessError>
close(id: string): Effect.Effect<void, HarnessError>
}
interface SandboxRuntime {
create(spec: SandboxSpec): Effect.Effect<SandboxLease, SandboxError>
exec(lease: SandboxLease, cmd: Command): Effect.Effect<CommandResult, SandboxError>
readFile(lease: SandboxLease, path: string): Effect.Effect<Uint8Array, SandboxError>
writeFile(lease: SandboxLease, path: string, body: Uint8Array): Effect.Effect<void, SandboxError>
pause(lease: SandboxLease): Effect.Effect<void, SandboxError>
resume(id: string): Effect.Effect<SandboxLease, SandboxError>
terminate(lease: SandboxLease): Effect.Effect<void, SandboxError>
}
interface SourceControl {
prepareWorktree(input: WorktreeInput): Effect.Effect<Worktree, GitError>
diff(worktree: Worktree): Effect.Effect<DiffArtifact, GitError>
commit(input: CommitInput): Effect.Effect<CommitArtifact, GitError>
push(input: PushInput): Effect.Effect<BranchArtifact, GitError>
createPullRequest(input: PullRequestInput): Effect.Effect<PullRequestArtifact, GitError>
}
interface VerificationRuntime {
execute(plan: VerificationPlan, env: EnvironmentRef):
Effect.Effect<VerificationResult, VerificationError>
}
```
Also define:
```text
ArtifactStore | SecretBroker | EventJournal | PreviewRuntime | RuntimePolicy
```
Use `Layer` for adapters, `Scope` for leases/processes, `Stream` for events, `Schedule` for bounded retries, and supervised fibers for long-running consumers. Pin Effect version and isolate beta API churn.
## 8. FLUE responsibilities
Use FLUE for domain-specific agents/workflows:
- conversation response and Signal proposal;
- Work Definition compiler;
- impact analysis;
- architecture/program design;
- vertical-slice planning;
- Resolver decision support;
- verification-plan generation;
- maintainability review;
- result/learning synthesis.
FLUE output MUST be typed proposals/commands, validated by application policy.
FLUE MUST NOT directly:
- mutate arbitrary database tables;
- bypass actor lifecycle;
- grant itself tools;
- merge/deploy outside policy;
- mark Work complete without evidence.
## 9. Harness strategy
Harness is replaceable:
```text
HarnessRuntime
├── OmpHarnessLive
├── OpenCodeHarnessLive
├── CodexHarnessLive
├── PiHarnessLive
└── FlueHarnessLive
```
v0 selects one harness. Prefer mature coding harnesses for repository exploration/edit/test loops; use custom FLUE agents for product-specific reasoning.
Zopu owns Work, branches, worktrees, budgets, approvals, artifacts, and completion. Harness owns one bounded implementation loop.
## 10. Runtime strategy
### CubeSandbox
Best for full Linux execution:
- native binaries;
- Bun/Node/Python;
- browsers;
- databases/services;
- project builds/tests;
- pause/resume;
- strong isolation.
Use E2B-compatible SDK behind `SandboxRuntime`. OMP runs as a process inside the Cube microVM.
### AgentOS
Best for:
- lightweight durable agent environments;
- ACP-integrated software;
- actor-adjacent orchestration;
- context/files/networking that fit runtime limits.
Use an attached full sandbox when native/heavy tooling is needed.
### Persistent project machine
Later optimization for long-lived caches, large repositories, and developer-customized environments. Use Git worktrees per active Work. Do not make it the only isolation boundary.
### Runtime selection
The Resolver requests capabilities:
```text
writable repo, Bun, browser, PostgreSQL, 8 GB RAM, network policy
```
`RuntimePolicy` chooses provider. Product logic never branches on Cube/AgentOS directly.
## 11. Repository isolation
Initial safe model:
```text
one project repository mirror
one Git worktree per active Work/slice
one mutating attempt per worktree
```
Rules:
- attempts receive scoped worktrees;
- parallel mutating attempts use separate branches;
- credentials are short-lived and repository-scoped;
- host home directories are never mounted;
- model credentials are run-scoped;
- untrusted code runs in sandbox;
- output commits record base and candidate SHA.
## 12. Planning and design artifacts
Required for standard work:
```text
WorkDefinition
ImpactMap
DesignPacket
VerticalSlicePlan
VerificationPlan
```
Program design SHOULD include:
- expected file-tree delta;
- expected call-flow delta;
- key types/signatures;
- dependency direction;
- invariants;
- security/data boundaries;
- known deviations.
The verifier compares candidate code against this design, but metrics are evidence, not absolute truth.
## 13. Verification architecture
Verification layers:
```text
Static
├── format/lint/typecheck
├── dependency policy
├── secret scan
└── static security
Behavior
├── unit
├── integration
├── contract
└── property tests where useful
Product
├── browser/user flow
├── screenshots/video
├── accessibility
└── visual checks
Operational
├── build/start/health
├── migration/rollback
├── logs
└── resource limits
Design
├── expected vs actual files/interfaces
├── dependency graph delta
├── design deviations
└── maintainability review
```
A `VerificationResult` MUST bind checks to candidate commit and environment.
Repair loop:
```text
failure evidence → bounded repair attempt → clean verification rerun
```
Tests added by the implementation SHOULD fail on the base revision and pass on the candidate when practical.
## 14. Delivery and integration
Publishing order:
```text
slice checks pass
→ integrated candidate created
→ impacted checks on integrated SHA
→ commit/push
→ PR
→ review package
```
Never create a “verified PR” from a different SHA than the verified candidate.
Review package:
- original intent;
- approved definition/design;
- slice narrative;
- meaningful diffs;
- screenshots/video;
- verification evidence;
- deviations/risks;
- exact commit/PR.
Manual merge remains policy in v0.
## 15. Preview and release
Preview is an artifact, not an open random port. `PreviewRuntime` may use:
- sandbox-exposed app;
- agentOS Apps for compatible generated HTTP apps;
- external staging/deployment.
Production release requires explicit policy, rollout plan, health signals, and rollback trigger.
## 16. Security baseline
- private control-plane endpoints;
- authenticated Cube/Rivet APIs;
- scoped runtime tokens;
- no long-lived model/Git secrets in workspace files;
- network egress policy;
- artifact access authorization;
- immutable audit trail for approvals;
- tool allowlist per Kit;
- destructive tools denied by default;
- merge/deploy human-gated initially;
- cleanup of terminated sandboxes and credentials.
## 17. Observability
Record product-level events, not only infrastructure logs.
Required dimensions:
```text
projectId workId sliceId runId attemptId actorId
kitVersion harness runtime model baseSha candidateSha
```
Track:
- state-transition latency;
- attempt duration/outcome;
- retries/replans;
- verification checks;
- human wait time;
- token/compute cost;
- duplicate side effects;
- abandoned/stale Work;
- post-release failures.
Raw harness logs are retained as artifacts; UI consumes normalized events.
## 18. Initial deployment shape
```text
Web + API + Convex
Bun/Effect daemon
Rivet cluster/actors
├── FLUE agents/workflows
└── SandboxRuntime
└── CubeSandbox on dedicated/VDS host
└── OMP + repo + tests
```
Keep Cube control APIs private; expose only authorized preview paths. Rivet and execution daemon may share the VPS initially but remain separate deployable processes.
## 19. Technical acceptance for v0
The architecture is proven when one real repository supports:
```text
message → Signal → approved Work → approved Design Packet
→ one slice → isolated harness run → independent verification
→ verified commit → real PR → human response/resume
```
With:
- durable recovery;
- no duplicate Work/PR;
- bounded retries;
- exact evidence;
- cancellation;
- explicit terminal states.