The Turn Lifecycle
One request to the full runtime becomes one GMI turn. Five stages own it, in this order; every line names the code that runs it. When the model requests a tool the host executes, processRequest() yields that chunk and returns; the host continues the same turn through resumeExternalToolRequest().
| Stage | Owner | What happens |
|---|---|---|
| 1. Facade | AgentOS.processRequest() | Fills selectedPersonaId from defaultPersonaId when the request has none, applies the self-improvement session overrides and skill prompt context, negotiates the language, and evaluates the input guardrails through evaluateInputGuardrails; a blocked input ends the request before any turn starts. It does not validate the input (phase 1 of the pipeline does), and it performs no authentication or rate limiting; the host does both before calling it. |
| 2. Orchestrator | AgentOSOrchestrator | Registers the stream and runs the preparation pipeline, then the GMI turn, and converts the result. |
| 3. Preparation | TurnExecutionPipeline.prepareTurn() | Twelve phases: input validation, GMI acquisition (GMIManager.getOrCreateGMIForSession()), stream context, GMI input construction, turn planning, adaptive execution policies, organization context and long-term memory policy, inbound message persistence, rolling summary compaction, prompt profile routing, long-term memory retrieval, and history assembly with metadata persistence. |
| 4. The GMI turn | GMI.processTurnStream() | Sentiment scoring when the persona enables it; the RAG trigger; memory context from CognitiveMemoryBridge when cognitive memory is attached; the prompt from the PromptEngine; the streaming model call; the tool loop through ToolOrchestrator, five iterations by default; history and memory updates; the metaprompts. |
| 5. Delivery | GMIChunkTransformer, wrapOutputGuardrails, StreamingManager | The orchestrator pushes the converted chunks into the streaming manager; the facade reads them back through an AsyncStreamClientBridge registered as one client, evaluates the output guardrails on that stream, and yields the result. Other clients of the stream receive the chunks unguarded. |
What a GMI emits
TEXT_DELTA for streamed text, TOOL_CALL_REQUEST when the model asks for a tool, USAGE_UPDATE with token counts, REASONING_STATE_UPDATE, SYSTEM_MESSAGE, RAG_SOURCES_AVAILABLE, LATENCY_REPORT, UI_COMMAND, a FINAL_RESPONSE_MARKER, or ERROR (IGMI.ts). The transformer maps them to AgentOSResponseChunkType values; an error anywhere in the turn ends the stream with an ERROR chunk whose message carries the original failure and whose code is GMI_PROCESSING_ERROR unless the error carries its own; an input blocked by a guardrail ends the request before the turn starts.
What persists between turns
The GMI's conversation history (a window of 20 messages unless the persona sets conversationContextConfig.maxMessages), its reasoning trace (500 entries by default, configurable per persona), its mood and user context in working memory, and, when attached, the cognitive memory traces it encoded. Across a restart, a storage adapter keeps agent-tier forged tools and the inbound messages the pipeline persists, a durable memory brain keeps the cognitive traces, and everything else is rebuilt.