AgentScope is an open-source project. To involve a broader community, we recommend asking your questions in English.
Prerequisites
Background / Description
Agent._compress_context_tool decides what to tell the model by comparing the
number of context messages before and after the call:
n_msgs = len(self.state.context)
try:
await self.compress_context(context_config=context_config)
except Exception as e:
return ToolChunk(..., state=ToolResultState.ERROR)
if len(self.state.context) == n_msgs:
text = ("The context is not long enough to compress, so it remains "
"unchanged.")
else:
text = "Context compressed successfully."
The message count is not a reliable proxy for "did we compress". In
_split_context_for_compression, when the very first message is itself over
the reserve budget it is chosen as the boundary message and then split by
content block:
if len(boundary_msg_to_compress.content) > 0:
msgs_to_compress += [boundary_msg_to_compress]
if len(boundary_msg_to_reserve.content) > 0:
msgs_to_reserve = [boundary_msg_to_reserve] + msgs_to_reserve
and _apply_change then sets:
self.state.summary = new_summary
self.state.context = msgs_to_reserve
With msg_index == 0 and a boundary message that splits into two non-empty
halves, msgs_to_reserve becomes [boundary_msg_to_reserve] + context[1:],
which is the same length as the original context. So a compression that
really happened - a summary generated, tokens spent, the context rewritten -
is reported to the model as "remains unchanged".
The knock-on effect is the worse part: the tool's own text tells the model the
context is unchanged, so a model that calls the tool again gets told the same
thing again, having paid for a summary generation each time. It is a loop
bait.
Error Messages
No exception is raised; the report is wrong. Observed with a context whose
first message alone exceeds the reserve budget and carries several content
blocks:
before_summary = agent.state.summary
chunk = await agent._compress_context_tool()
print(chunk.content[0].text)
# The context is not long enough to compress, so it remains unchanged.
print(agent.state.summary != before_summary)
# True <- a summary *was* produced
Steps to Reproduce
- Code:
model = RecordingStructuredMockModel(context_size=<small>)
model.set_structured_response(StructuredResponse(content={...five keys...}))
agent = Agent(
name="Friday",
system_prompt="You are helpful.",
model=model,
context_config=ContextConfig(
trigger_ratio=0.9,
reserve_ratio=0.3,
context_buffer_ratio=0.4,
compression_tool_enabled=True,
),
state=AgentState(
session_id="123",
context=[
# first message alone is over the reserve budget and has
# several content blocks, so it is chosen and then split
UserMsg("User", [TextBlock(text="a" * N),
TextBlock(text="b" * N),
TextBlock(text="c" * N)]),
UserMsg("User", "d" * N),
UserMsg("User", "e" * N),
],
),
toolkit=Toolkit(),
)
before = agent.state.summary
chunk = await agent._compress_context_tool()
assert agent.state.summary != before, "no summary was produced"
assert "remains unchanged" in chunk.content[0].text, "but it was reported as such"
-
Run: python repro.py (the exact context size that lands on the split
boundary depends on the tokenizer, which is why the accompanying test sweeps
sizes rather than hard-coding one).
-
See: the second assertion fires - a summary was produced and the tool still
reported "remains unchanged".
Expected: the tool reports "Context compressed successfully."
Environment
- AgentScope Version: 2.0.10dev (main at
6109dd0)
- Python Version: 3.11
- OS: Linux
Reproduced with the repo's own RecordingStructuredMockModel, so no
credentials and no network are needed.
AgentScope is an open-source project. To involve a broader community, we recommend asking your questions in English.
Prerequisites
Background / Description
Agent._compress_context_tooldecides what to tell the model by comparing thenumber of context messages before and after the call:
The message count is not a reliable proxy for "did we compress". In
_split_context_for_compression, when the very first message is itself overthe reserve budget it is chosen as the boundary message and then split by
content block:
and
_apply_changethen sets:With
msg_index == 0and a boundary message that splits into two non-emptyhalves,
msgs_to_reservebecomes[boundary_msg_to_reserve] + context[1:],which is the same length as the original context. So a compression that
really happened - a summary generated, tokens spent, the context rewritten -
is reported to the model as "remains unchanged".
The knock-on effect is the worse part: the tool's own text tells the model the
context is unchanged, so a model that calls the tool again gets told the same
thing again, having paid for a summary generation each time. It is a loop
bait.
Error Messages
No exception is raised; the report is wrong. Observed with a context whose
first message alone exceeds the reserve budget and carries several content
blocks:
Steps to Reproduce
Run:
python repro.py(the exact context size that lands on the splitboundary depends on the tokenizer, which is why the accompanying test sweeps
sizes rather than hard-coding one).
See: the second assertion fires - a summary was produced and the tool still
reported "remains unchanged".
Expected: the tool reports "Context compressed successfully."
Environment
6109dd0)Reproduced with the repo's own
RecordingStructuredMockModel, so nocredentials and no network are needed.