Repository navigation
Replies: 3 comments
|
@qbc2016 |
0 replies
|
@iluv7 The overall direction looks good to me. |
0 replies
|
Thanks for taking a look! I’ll move forward with the implementation based on this direction and submit a PR with the corresponding tests. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
This discussion proposes an implementation direction for #2566: context compression for
RealtimeAgent, including whether a compressed checkpoint should trigger a proactive session reconnect.Proposed approach
Realtime context differs from the standard
Agentcontext because the active conversation is maintained by the remote session. Compressing onlystate.contextdoes not shorten the context already held by the provider; the summary takes effect only after a reconnect.The proposed flow is:
Providers for which proactive reconnect is not yet reliable can fall back to preparing the checkpoint and using it only after a natural disconnect.
Should we proactively reconnect?
Without a proactive reconnect, local compression does not change the active remote context. The provider may continue growing or automatically truncate the earliest history before the checkpoint is ever used. A proactive reconnect makes the compressed context effective immediately, at the cost of connection latency and additional lifecycle complexity.
To check whether the reconnect overhead is acceptable, we ran 20 sequential trials against DashScope
qwen-audio-3.0-realtime-flash. Each trial contained a baseline turn and a reconnect turn. During reconnect, a fixed 2.35-second PCM input was buffered locally and sent after the new session was established.Results:
connect()p50connect()p95Although establishing a session took about 1.1 seconds, most of that time overlapped with the user audio, so the observed user-visible overhead remained below 500 ms. No audio loss was observed in these trials.
This is only a smoke test. Twenty successful trials do not establish a 99% success rate. Roughly 300 zero-failure trials would be needed to support a sub-1% failure rate at about 95% confidence. We also need short-input tests, because 300 ms or 500 ms of speech may not hide the reconnect as effectively as the 2.35-second input used here.
Based on the current result, our preference is to support proactive reconnect at a safe boundary, with a per-provider fallback to natural reconnect.
Open question 1: when should compression start?
Realtime providers expose different types of limits, so the trigger should not depend on a single universal token counter.
input_tokens + output_tokens. Input tokens should not be summed across turns because the latest input count normally already includes conversation history.Available dimensions can be normalized into a pressure value:
A possible default is to start background compression around 70% and schedule rollover no later than around 85%. The thresholds should remain configurable, while provider adapters normalize usage and expose their applicable limits.
Open question 2: what happens to messages created during summarization?
The summary must not replace the live context directly. Otherwise, messages created while the summarizer is running may be lost.
Instead, compression should use an immutable snapshot and a revision/watermark:
For example:
If the snapshotted prefix changes due to interruption handling or message aggregation, the candidate summary should be discarded and retried.
Reconnect boundary and audio handling
The reconnect should happen only when the current user turn and assistant response are complete, no tool call or confirmation is pending, and no assistant audio is still playing.
During reconnect:
Each connection should also have a session epoch so late terminal or error events from the old session cannot affect the new one. In addition, adapter
connect()methods need a clear ready guarantee; returning immediately after sendingsession.updateis not a sufficiently strong contract for reliable backlog replay.Validation plan
Unit and adapter tests should cover:
Feedback is especially welcome on the provider capability boundary and whether proactive reconnect should initially be opt-in or enabled by default for validated providers.
All reactions