Check for existing issues
What happened?
AttributeError: LiteLLMCompletionStreamingIterator object has no attribute completed_response on mid-stream provider error (Anthropic via Responses API)
What happened
When the Anthropic provider returns a mid-stream error (e.g. HTTP 529 Overloaded / a 5xx) while streaming a Responses API completion, litellm's fallback/cleanup path raises:
AttributeError: 'LiteLLMCompletionStreamingIterator' object has no attribute 'completed_response'
This AttributeError masks the real error (the underlying MidStreamFallbackError / provider overload) and escapes the normal streaming-error contract. Because the surfaced exception is now a bare AttributeError rather than a stream/incomplete error, any downstream retry or model-fallback logic keyed on litellm's streaming exceptions never runs, and the call fails hard instead of falling back.
Root cause (as far as we can tell)
On a mid-stream provider error, litellm raises MidStreamFallbackError and then attempts to recover partial token usage by reading source_iterator.completed_response during cleanup.
For Anthropic-served-via-the-Responses-API, the streaming iterator is LiteLLMCompletionStreamingIterator. That class overrides __init__ without calling super().__init__(), so the completed_response attribute the base class would set is never initialized. When the recovery path reads source_iterator.completed_response, it raises AttributeError.
Approximate locations in our installed copy (line numbers may drift between versions — please confirm against your tree):
- Partial-usage recovery reads
source_iterator.completed_response — litellm/router.py, in _extract_partial_responses_usage
LiteLLMCompletionStreamingIterator.__init__ does not call super().__init__(), so completed_response is never set — litellm/…/streaming_iterator.py
Impact
- The real provider error (Anthropic 529 Overloaded) is hidden behind an unrelated
AttributeError, which makes triage harder.
- The masked exception is not recognized as a streaming/incomplete error, so configured retries and model fallbacks are bypassed — a transient, recoverable provider overload becomes an unrecoverable hard failure.
- We see this recur under Anthropic capacity pressure, so it is not a rare edge case for high-volume users.
Steps to reproduce
- Use litellm with an Anthropic model routed through the Responses API streaming path (in our case
anthropic/claude-sonnet-4-6), with a retry/fallback policy configured.
- Trigger (or simulate) a mid-stream provider error — e.g. an HTTP 529 Overloaded / 5xx raised part-way through the stream, after some chunks have been received.
- Observe that instead of the expected
MidStreamFallbackError (and instead of the fallback engaging), the call raises AttributeError: 'LiteLLMCompletionStreamingIterator' object has no attribute 'completed_response'.
A minimal repro can be built by mocking the underlying Anthropic stream to yield one or more chunks and then raise an overloaded/5xx error mid-iteration, with a LiteLLMCompletionStreamingIterator as the source iterator.
Expected behavior
The mid-stream provider error should surface as the intended MidStreamFallbackError (or a proper stream/incomplete error), so that:
- the original cause (provider overload) is visible, and
- retry / model-fallback logic runs as configured.
The partial-usage recovery should not throw AttributeError when completed_response was never initialized.
Suggested fix
Either of:
- Have
LiteLLMCompletionStreamingIterator.__init__ call super().__init__() (or otherwise initialize completed_response) so the attribute always exists; and/or
- Make
_extract_partial_responses_usage defensive, e.g. getattr(source_iterator, "completed_response", None), so a missing attribute degrades to "no partial usage" instead of raising and masking the original error.
Environment
- litellm version: 1.90.0 (please confirm — reporter to fill in exact installed version)
- Provider: Anthropic
- Model:
anthropic/claude-sonnet-4-6
- API surface: Responses API (streaming)
- Python:
- OS:
Observed traceback (sanitized)
error: 'LiteLLMCompletionStreamingIterator' object has no attribute 'completed_response'
error_type: AttributeError
model: anthropic/claude-sonnet-4-6
api_type: responses
# original cause: mid-stream AnthropicError: Overloaded (HTTP 529) -> MidStreamFallbackError, masked by the AttributeError above
Steps to Reproduce
- Use litellm with an Anthropic model routed through the Responses API streaming path (e.g.
anthropic/claude-sonnet-4-6), with retries / model-fallback configured.
- Trigger a mid-stream provider error — an HTTP 529 Overloaded (or 5xx) raised part-way through the stream, after one or more chunks have already been received. This can be simulated by mocking the Anthropic stream to yield a chunk and then raise an overloaded/5xx error mid-iteration, with a
LiteLLMCompletionStreamingIterator as the source iterator.
- Observe that the call raises
AttributeError: 'LiteLLMCompletionStreamingIterator' object has no attribute 'completed_response' instead of the expected MidStreamFallbackError, and the configured retries / fallback never run.
Relevant log output
error: 'LiteLLMCompletionStreamingIterator' object has no attribute 'completed_response'
error_type: AttributeError
model: anthropic/claude-sonnet-4-6
api_type: responses
original cause: mid-stream AnthropicError: Overloaded (HTTP 529) -> MidStreamFallbackError, masked by the AttributeError above
What part of LiteLLM is this about?
SDK (litellm Python package)
What LiteLLM version are you on ?
v1.90.0
Twitter / LinkedIn details
@otto_the_agent
Check for existing issues
What happened?
AttributeError:
LiteLLMCompletionStreamingIteratorobject has no attributecompleted_responseon mid-stream provider error (Anthropic via Responses API)What happened
When the Anthropic provider returns a mid-stream error (e.g. HTTP 529 Overloaded / a 5xx) while streaming a Responses API completion, litellm's fallback/cleanup path raises:
This
AttributeErrormasks the real error (the underlyingMidStreamFallbackError/ provider overload) and escapes the normal streaming-error contract. Because the surfaced exception is now a bareAttributeErrorrather than a stream/incomplete error, any downstream retry or model-fallback logic keyed on litellm's streaming exceptions never runs, and the call fails hard instead of falling back.Root cause (as far as we can tell)
On a mid-stream provider error, litellm raises
MidStreamFallbackErrorand then attempts to recover partial token usage by readingsource_iterator.completed_responseduring cleanup.For Anthropic-served-via-the-Responses-API, the streaming iterator is
LiteLLMCompletionStreamingIterator. That class overrides__init__without callingsuper().__init__(), so thecompleted_responseattribute the base class would set is never initialized. When the recovery path readssource_iterator.completed_response, it raisesAttributeError.Approximate locations in our installed copy (line numbers may drift between versions — please confirm against your tree):
source_iterator.completed_response—litellm/router.py, in_extract_partial_responses_usageLiteLLMCompletionStreamingIterator.__init__does not callsuper().__init__(), socompleted_responseis never set —litellm/…/streaming_iterator.pyImpact
AttributeError, which makes triage harder.Steps to reproduce
anthropic/claude-sonnet-4-6), with a retry/fallback policy configured.MidStreamFallbackError(and instead of the fallback engaging), the call raisesAttributeError: 'LiteLLMCompletionStreamingIterator' object has no attribute 'completed_response'.A minimal repro can be built by mocking the underlying Anthropic stream to yield one or more chunks and then raise an overloaded/5xx error mid-iteration, with a
LiteLLMCompletionStreamingIteratoras the source iterator.Expected behavior
The mid-stream provider error should surface as the intended
MidStreamFallbackError(or a proper stream/incomplete error), so that:The partial-usage recovery should not throw
AttributeErrorwhencompleted_responsewas never initialized.Suggested fix
Either of:
LiteLLMCompletionStreamingIterator.__init__callsuper().__init__()(or otherwise initializecompleted_response) so the attribute always exists; and/or_extract_partial_responses_usagedefensive, e.g.getattr(source_iterator, "completed_response", None), so a missing attribute degrades to "no partial usage" instead of raising and masking the original error.Environment
anthropic/claude-sonnet-4-6Observed traceback (sanitized)
Steps to Reproduce
anthropic/claude-sonnet-4-6), with retries / model-fallback configured.LiteLLMCompletionStreamingIteratoras the source iterator.AttributeError: 'LiteLLMCompletionStreamingIterator' object has no attribute 'completed_response'instead of the expectedMidStreamFallbackError, and the configured retries / fallback never run.Relevant log output
What part of LiteLLM is this about?
SDK (litellm Python package)
What LiteLLM version are you on ?
v1.90.0
Twitter / LinkedIn details
@otto_the_agent