2026-09-25 · America/Los_Angeles · 社区动态 · #5
What should an LLM app do when a provider times out mid-response?
I’m trying to understand how developers design fallback behavior when an LLM provider fails after a response has started streaming. If some tokens have already reached the user, switching to another model seems harder than retrying a request that failed before any output. The second model may produce a different answer, and the user could end up with a confusing mix of both. How do you handle this in practice? Do you stop and show an error, restart the whole answer with another model, or let the user choose whether to retry? I’d especially like to hear how you handle conversation state and…
热度 46.0 / 100;排名与评分保留该期记录。