Skip to content

fix(agent): propagate length stops and guard incomplete tool calls - #3195

Open
wuyak wants to merge 1 commit into
agentscope-ai:mainfrom
wuyak:codex/s18-output-truncation-v2
Open

wuyak wants to merge 1 commit into
agentscope-ai:mainfrom
wuyak:codex/s18-output-truncation-v2

Conversation

@wuyak

@wuyak wuyak commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

AgentScope Version

2.0.10dev

Description

Closes #3194.

Model providers report why generation stopped, but AgentScope currently drops that information and leaves ChatResponse.finished_reason as COMPLETED. If generation reaches an output-length limit while producing tool arguments, the existing JSON repair can make the partial arguments syntactically valid and allow the tool to run.

For example, assume the agent exposes this tool:

def write_file(path: str, content: str):
    with open(path, "w") as file:
        file.write(content)

The model may intend to generate:

{
  "path": "/tmp/result.txt",
  "content": "unfinished task, please continue..."
}

If the provider stops generation at its output limit, AgentScope may receive:

{"path": "/tmp/result.txt", "content": "unfinished

The existing repair path can turn it into:

{"path": "/tmp/result.txt", "content": "unfinished"}

and execute the equivalent of:

write_file(
    path="/tmp/result.txt",
    content="unfinished",
)

JSON repair can restore syntax, but it cannot recover content the model never generated.

Changes

  • Added FinishedReason.LENGTH.

  • Preserved the provider's original value in ChatResponse.metadata["raw_finish_reason"].

  • Preserved explicit finish reasons and metadata through streaming accumulation.

  • Read finish reasons inside the existing parsing paths of the supported adapters:

    • OpenAI Chat, DashScope, DeepSeek, Moonshot, and Volcengine: length
    • Anthropic: max_tokens and model_context_window_exceeded
    • Gemini: MAX_TOKENS
    • xAI: REASON_MAX_LEN and REASON_MAX_CONTEXT
    • Ollama: length
  • Added an opt-in ReAct guard:

    ReActConfig(reject_truncated_tool_call=True)

    The option defaults to False. When enabled, AgentScope rejects only the final tool call when the response is explicitly marked FinishedReason.LENGTH and strict json.loads cannot parse its arguments. The rejected call is closed with the existing ERROR ToolResult; the existing ReAct loop decides what happens next.

Behavior boundaries

  • Earlier complete tool calls in the same response still execute.
  • A final tool call with valid JSON still enters the normal tool path, even when the provider reports a length limit.
  • Unknown raw finish reasons are retained for diagnostics but are not classified as length limits.
  • The guard does not infer truncation from token usage or missing business fields.
  • This PR does not add global stopping, retries, continuation, argument completion, warnings, hints, or context-compression behavior.
  • OpenAI Responses is intentionally unchanged because its existing incomplete-response path has different error semantics.
  • The guard depends on an explicit provider signal. In a separate qwen3.7-flash probe, the provider reported normal completion/tool use for a token-limited malformed tool call, so the guard correctly did not activate.

Added tests

This PR adds 14 test methods to AgentScope's existing Agent and model-adapter test files, covering:

  • extraction, retention, and normalization of representative length-limit values for each changed adapter;
  • finish-reason propagation through OpenAI Chat, xAI, and the common stream accumulator;
  • retention of unknown raw reasons without classifying them as length limits;
  • preservation of the existing JSON repair and execution behavior while the option is disabled;
  • rejection of only the final incomplete tool call while the option is enabled;
  • execution of earlier complete tool calls in the same response;
  • execution of valid JSON arguments despite a length-limit reason;
  • continuation of the existing ReAct state machine after an error ToolResult is produced.
PYTHONPATH=src python -m pytest -q \
  tests/agent_basic_test.py \
  tests/model_base_test.py \
  tests/model_anthropic_test.py \
  tests/model_dashscope_test.py \
  tests/model_deepseek_test.py \
  tests/model_gemini_test.py \
  tests/model_moonshot_test.py \
  tests/model_ollama_test.py \
  tests/model_openai_chat_test.py \
  tests/model_volcengine_test.py \
  tests/model_xai_test.py

Result:

228 passed, 32 subtests passed

git diff --check also passes.

Real-model validation

I also ran a temporary, side-effect-free probe against deepseek-v4.1-flash through Alibaba Cloud Model Studio's OpenAI-compatible Chat API (DeepSeek API, model documentation). The probe used this branch, max_tokens=64, the guard enabled, a two-iteration ReAct limit, and an in-memory tool.

Observed behavior:

  1. The first model call returned length with incomplete JSON tool arguments.
  2. AgentScope produced an ERROR ToolResult and did not execute the tool.
  3. The existing ReAct loop made a second real model call.
  4. Under a controlled prompt, the second call returned TRUNCATION_ERROR_OBSERVED.

The tool execution count remained zero. This confirms the provider signal reaches the guard and that the existing loop can continue after the error result; the controlled prompt does not claim a general model recovery strategy.

Checklist

  • An issue has been created for this PR
  • I have read the CONTRIBUTING.md
  • Docstrings are in Google style
  • Related documentation has been updated (e.g. links, examples, etc.) in documentation repository
  • Code is ready for review

@github-actions github-actions Bot added the pr/awaiting-review Waiting for maintainers to review this PR label Oct 9, 2026
@wuyak wuyak closed this Oct 9, 2026
@wuyak
wuyak deleted the codex/s18-output-truncation-v2 branch October 9, 2026 06:49
@wuyak
wuyak restored the codex/s18-output-truncation-v2 branch October 9, 2026 06:50
@wuyak wuyak reopened this Oct 9, 2026
@wuyak
wuyak force-pushed the codex/s18-output-truncation-v2 branch from ea5bc8a to 4939d50 Compare October 9, 2026 06:50

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

pr/awaiting-review Waiting for maintainers to review this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Provider length-limit stop reasons are lost, so incomplete tool calls may run

1 participant