Skip to main content

Self-Extension: Forging and Self-Improvement

An agent can add to itself in three ways, two on the runtime and one in agency(), each behind its own switch: it can forge a tool, it can adjust its own personality, skills, workflows and parameters, and a hierarchical agency's manager can spawn a specialist.

Forging a tool​

With emergent: true on the runtime config, the tool orchestrator registers the forge_tool meta-tool (ToolOrchestrator, ForgeToolMetaTool). A request names the tool, its input and output JSON Schemas, test cases and an implementation: compose chains existing tools through ComposableToolBuilder; sandbox runs agent-written JavaScript in a node:vm context through SandboxedToolForge, and is rejected until emergentConfig.allowSandboxTools is true. The EmergentCapabilityEngine builds the candidate, runs its tests, submits it to the EmergentJudge (one model call scoring safety, correctness, determinism and boundedness; approval needs safety and correctness to pass; with no judge model call configured every request is rejected) and registers an approved tool at the session tier of the EmergentToolRegistry.

Session-tier tools live in memory for the session (mirrored to SQLite when a storage adapter exists). After five uses (totalUses, whether or not they succeeded) with a judge confidence of at least 0.8, a two-reviewer promotion panel can move a tool to the agent tier, persisted in the agentos_emergent_tools table; the shared tier needs an explicit promote() call, which records an approver only when the caller passes approvedBy. Forge, promotion and removal decisions go to the agentos_emergent_audit_log table when a storage adapter is configured. Raw sandbox source is not stored unless persistSandboxSource is set.

The sandbox bans eval, Function, require, import, process, child_process and the file-writing calls at validation time, exposes fetch, fs.readFile and crypto only when a request's allowlist names them, enforces a wall clock of sandboxTimeoutMs (5000 ms by default; a request's own timeoutMs takes precedence), reads files only under the process working directory, and lets fetch reach any host: the runtime constructs the forge without fsReadRoots or fetchDomainAllowlist, and emergentConfig has no key for either.

Self-improvement tools​

With emergentConfig.selfImprovement.enabled, four more tools are registered: adapt_personality (HEXACO deltas clamped to the 0 to 1 range and a per-session budget, recorded in the PersonalityMutationStore when a storage adapter exists), manage_skills, create_workflow and self_evaluate (scores a response with a model call, records adjustments to runtime parameters in its own session state, and reports on them). Stored mutations are not reloaded into a later GMI; a trait change lasts for the instance that made it.

Spawning a specialist​

In an agency running the hierarchical strategy with emergent planning enabled, the manager gets a spawn_specialist tool: EmergentAgentForge turns the manager's spec into an agent config, EmergentAgentJudge reviews it, and the new agent joins the roster as a delegate_to_<role> tool for the manager's next turn (hierarchical.ts). At most five specialists per run by default; spawned specialists are reset between send() calls.

What is not there​

No weights are updated anywhere. self_evaluate keeps its scores and adjustments in its own session state; the runtime's session manager can carry a temperature or verbosity override into later requests of the same session (SelfImprovementSessionManager.setRuntimeParam()), but no shipped tool writes it, so no evaluation result changes a later turn on its own. No loop rewrites the agent's own code or prompts persistently. Emergent Capabilities is the full guide.