Interface: FallbackProviderEntry
Defined in: packages/agentos/src/api/generateText.ts:273
A fallback provider entry specifying an alternative provider (and optionally model) to try when the primary provider fails with a retryable error.
See
GenerateTextOptions.fallbackProviders
Properties
cache?
optionalcache:false| {ttl?:"5m"|"1h"; }
Defined in: packages/agentos/src/api/generateText.ts:302
Per-hop prompt-cache disposition applied ONLY when THIS entry serves the
call (same override shape as FallbackProviderEntry.effort).
false sends zero cache_control on the hop; { ttl } re-times the
hop's markers. Omitted -> the hop inherits the call-level
GenerateTextOptions.cache.
The canonical chains (buildFallbackChain /
buildPolicyAwareFallbackChain) pin cache: false on every leg:
rescue traffic is sporadic and one-shot-shaped, so cache writes on a
fallback hop rarely earn their reads back (the claude-sonnet-5 leg
measured 0.45x write amortization in wilds prod, 2026-07-13..20 — the
writes cost ~2x what the reads saved). A caller-supplied entry keeps
full control: omit to inherit, or set a ttl to cache a hop deliberately.
effort?
optionaleffort:string
Defined in: packages/agentos/src/api/generateText.ts:286
Per-hop reasoning depth applied ONLY when THIS entry serves the call,
forwarded as output_config.effort (Anthropic) / reasoning_effort
(OpenAI). Lets a chain run a fallback at a different depth than the primary
— e.g. a gpt-5.6-sol frontier fallback at 'max' while the primary keeps its
own (or no) effort, so arming the chain is dormant for the primary call.
Omitted -> the hop inherits the call-level effort.
maxTokensHeadroom?
optionalmaxTokensHeadroom:number
Defined in: packages/agentos/src/api/generateText.ts:317
Extra output tokens granted ONLY when THIS entry serves the call: the hop
runs with maxTokens = <the original call's maxTokens> + headroom. For
models whose hidden reasoning shares the output cap. Every current Gemini
3.x model thinks, and its thinking tokens count against
maxOutputTokens; a rescue hop inherits a budget sized for the primary,
and without room for the thinking the visible reply comes back cut short
or empty (measured 2026-09-30: gemini-3.1-pro-preview at 400 tokens
spent 382 thinking and returned 60 characters).
Ignored when the call set no maxTokens. Never compounds: every hop is
granted its own headroom over the ORIGINAL budget, and an entry without
headroom runs at the original.
model?
optionalmodel:string
Defined in: packages/agentos/src/api/generateText.ts:277
Model identifier override. When omitted, the provider's default text model is used.
provider
provider:
string
Defined in: packages/agentos/src/api/generateText.ts:275
Provider identifier (e.g. "openai", "anthropic", "openrouter").