AgentOS documentation: getting started guides, full API reference, runtime recipes, extension catalog, and benchmark results for the open-source TypeScript AI agent runtime.
Runtime agent spawning
A manager with a researcher and a writer gets a task neither of them covers. It callsspawn_specialist, the LLM judge approves the new agent's spec, and the specialist joins the roster as a delegate_to_<role> tool for the manager's next turn.
npm install @framers/agentos
import { agent } from '@framers/agentos';
// Personality is six 0-1 trait values. Each trait above 0.65 or below 0.35
// adds one line of direction to the system prompt; a trait in between, or
// one left out (0.5), adds none.
const tutor = agent({
provider: 'openai',
model: 'gpt-4o',
instructions: 'You are a patient programming tutor.',
personality: {
honesty: 0.85, // direct, transparent, no flattery
emotionality: 0.70, // tone-aware without being clinical
extraversion: 0.50, // in between: adds no line
agreeableness: 0.75, // warm, encouraging
conscientiousness: 0.90, // structured, thorough, follow-through
openness: 0.85, // creative, exploratory framing
},
});
// Each session keeps its own conversation history in process memory
// (bounded to about 120,000 tokens), so two session ids share nothing.
const session = tutor.session('user-42');
// Every send() passes the session's earlier turns to the model.
await session.send('My exam is on distributed systems next Thursday.');
await session.send('I struggle with consensus algorithms.');
const reply = await session.send('What should I focus on this week?');
console.log(reply.text);
// The transcript the session holds, and the tokens it has used.
console.log(session.messages());
const usage = await session.usage();
console.log(`Total tokens: ${usage.totalTokens}`);
Seven cooperating layers. API surface at the top, channels and providers at the floor, cognition and memory in the middle. Click to zoom.
Text, images, video, music, SFX, embeddings, and speech from one API. Cloud and local backends share the same surface, with fallback chains and provider preferences that order, filter or weight them.
mission() compiles a goal into a linear step graph from a plan template (research, Q&A or creative), with anchor nodes for verification and human review spliced into its phases.
Agents forge new tools at runtime — compose (chain existing tools) or sandbox (generated JavaScript in node:vm or QuickJS, with allowlists; off by default). LLM-as-judge review, tiered promotion, portable YAML export.
Voice pipeline with VAD, STT, endpoint detection and TTS, and Twilio, Telnyx and Plivo call providers with a media-stream transport for phone calls.
Three authoring APIs — AgentGraph, workflow() DSL, mission() — compile to one IR. judgeNode for evaluation, checkpoints to resume from, streaming events.
Ebbinghaus decay, spreading activation, Baddeley-style working memory, GraphRAG retrieval and consolidation, plus 8 neuroscience-grounded mechanisms that run with a cognitiveMechanisms config and that HEXACO traits modulate.
Five guardrail packs: PII redaction (regex, NLP, NER and an LLM judge), ML classifiers (ONNX toxic-bert, an LLM judge or keywords), topicality (embedding similarity to allowed and blocked topics), code safety (OWASP-style rules) and grounding (NLI against the retrieved sources).
Test cases scored by built-in or custom scorers and an LLM judge with criteria presets. Two runs compared side by side; reports in JSON, Markdown or HTML.
Three tiers held to token budgets: category summaries (200 tokens) → the top 5 matches (800) → full schemas for the top 2 (2,000), plus a discover_capabilities tool for active search.
Signed event ledger (Ed25519 signatures over a SHA-256 hash chain), revision snapshots, tombstones for deletes, an autonomy guard, and Merkle roots anchored outside the database.
Telegram, Discord, Slack, WhatsApp, Twitter/X, LinkedIn, Bluesky, Mastodon, and custom adapters. Multi-channel routing, social publishing, browser automation, and adapter APIs.
Sealed storage policy with the signed ledger, revisions, tombstones and anchors, and a design guide for toolset pinning, secret rotation and forgetting. A sealed agent's changes are tamper-evident.
generateVideo(), analyzeVideo(), detectScenes(), generateMusic(), generateSFX() APIs. 3 video providers (Runway, Replicate, Fal) + 8 audio providers. Fallback chains and scene detection.
SKILL.md prompt modules for research, developer tools, communication, productivity, security, media, and creative workflows. Semantic discovery finds the right skill per turn.
Bounded self-modification: adapt_personality (HEXACO mutation with per-session budgets), manage_skills, create_workflow, self_evaluate. Recorded mutations decay when adapt_personality runs; the live trait keeps its change.