Skip to main content

Interface: FallbackProviderEntry

Defined in: packages/agentos/src/api/generateText.ts:273

A fallback provider entry specifying an alternative provider (and optionally model) to try when the primary provider fails with a retryable error.

See​

GenerateTextOptions.fallbackProviders

Properties​

cache?​

optional cache: false | { ttl?: "5m" | "1h"; }

Defined in: packages/agentos/src/api/generateText.ts:302

Per-hop prompt-cache disposition applied ONLY when THIS entry serves the call (same override shape as FallbackProviderEntry.effort). false sends zero cache_control on the hop; { ttl } re-times the hop's markers. Omitted -> the hop inherits the call-level GenerateTextOptions.cache.

The canonical chains (buildFallbackChain / buildPolicyAwareFallbackChain) pin cache: false on every leg: rescue traffic is sporadic and one-shot-shaped, so cache writes on a fallback hop rarely earn their reads back (the claude-sonnet-5 leg measured 0.45x write amortization in wilds prod, 2026-07-13..20 — the writes cost ~2x what the reads saved). A caller-supplied entry keeps full control: omit to inherit, or set a ttl to cache a hop deliberately.


effort?​

optional effort: string

Defined in: packages/agentos/src/api/generateText.ts:286

Per-hop reasoning depth applied ONLY when THIS entry serves the call, forwarded as output_config.effort (Anthropic) / reasoning_effort (OpenAI). Lets a chain run a fallback at a different depth than the primary — e.g. a gpt-5.6-sol frontier fallback at 'max' while the primary keeps its own (or no) effort, so arming the chain is dormant for the primary call. Omitted -> the hop inherits the call-level effort.


maxTokensHeadroom?​

optional maxTokensHeadroom: number

Defined in: packages/agentos/src/api/generateText.ts:317

Extra output tokens granted ONLY when THIS entry serves the call: the hop runs with maxTokens = <the original call's maxTokens> + headroom. For models whose hidden reasoning shares the output cap. Every current Gemini 3.x model thinks, and its thinking tokens count against maxOutputTokens; a rescue hop inherits a budget sized for the primary, and without room for the thinking the visible reply comes back cut short or empty (measured 2026-09-30: gemini-3.1-pro-preview at 400 tokens spent 382 thinking and returned 60 characters).

Ignored when the call set no maxTokens. Never compounds: every hop is granted its own headroom over the ORIGINAL budget, and an entry without headroom runs at the original.


model?​

optional model: string

Defined in: packages/agentos/src/api/generateText.ts:277

Model identifier override. When omitted, the provider's default text model is used.


provider​

provider: string

Defined in: packages/agentos/src/api/generateText.ts:275

Provider identifier (e.g. "openai", "anthropic", "openrouter").