Skip to content

feat(realtime): support context management for the realtime agent #2566

Description

@DavdGao

Background
Context management for RealtimeAgent differs fundamentally from the standard Agent class:

  • Runtime context is maintained on the API side, not locally
  • On session recovery, context must be re-injected — either merged into the system prompt/instructions, or replayed as text messages to the model. This implies a summarization step: a compressed summary of prior context should be prepended to the system prompt when resuming a session
  • Compressed context segments should be offloaded to a local workspace, consistent with how Agent handles this
  • Truncation strategy for tool results should align with the existing Agent behavior

Two open questions to resolve before implementation:

  • Token counting: Realtime APIs may not expose token usage in the same way as standard chat APIs — the counting mechanism needs to be clarified first, as it gates the truncation/compression trigger
  • Compression timing: Unlike Agent, compression in RealtimeAgent can potentially run concurrently with the ongoing conversation — the concurrency model needs to be decided

Changes

  • Add a ContextConfig class for RealtimeAgent, mirroring Agent's context_config, with the following fields:
    • context_length: maximum allowed context size
    • tool_result_limit: cap on individual tool result size
    • compression_prompt: a ChatModelBase instance used for compression
    • compression_prompt: prompt template for compression
    • compression_schema: structured output BaseModel for compression result
    • compression_summary: the summary template
  • Add a workspace parameter to RealtimeAgent for offloading truncated and compressed tool results

Activity

  1. iluv7 commented on Sep 9, 2026

    @iluv7
    Contributor

    Hi, I’d like to work on this issue. I’ll first investigate the Realtime API’s token-usage capabilities and the concurrency model for context compression. I’ll discuss the possible technical approaches to the two open questions here before proposing an implementation aligned with the existing Agent context-management behavior.
    Please assign this issue to me if the direction sounds good. Thanks!

  2. radianceded commented on Sep 9, 2026

    @radianceded

    Hi @maintainers — I'd like to take this one. I've read through RealtimeAgent (agent/_realtime/_agent.py), Agent's context pipeline (agent/_agent.py, agent/_config.py) and the realtime event/metrics code, and I can propose concrete answers to both open questions before starting implementation.

    Q1 — Token counting. We already have a per-turn signal: ResponseDoneEvent.input_tokens/output_tokens (realtime/_events.py), surfaced into TurnMetrics in the downlink pump. I propose maintaining a running estimate as:

    • base = provider-reported usage accumulated per completed turn (cumulative when a provider reports session-cumulative counts, summed deltas when it reports per-turn);
    • fallback = character-based estimation over the transcript delta in state.context for any turn where the provider omitted usage (fields default to 0), so the trigger still works on providers without usage reporting.

    The estimate would live alongside TurnMetrics and gate both truncation and compression against context_length.

    Q2 — Compression timing / concurrency. A "rolling checkpoint" model, so compression never blocks the audio loop:

    1. At trigger_ratio, spawn one background compression task (single-flight, cancelled on close()) over an immutable snapshot of state.context up to a monotonic sequence number.
    2. The summary commits only at a turn boundary (after ResponseDoneEvent, before the next user speech start): the summarized prefix is atomically replaced by the summary message, and summarized_upto_seq is recorded. Turns newer than the checkpoint stay verbatim.
    3. On reconnect, connect() (which currently unconditionally appends the raw transcript as "## Conversation so far", connect() L263-272) prepends the latest committed checkpoint summary to the instructions instead, followed by only the post-checkpoint verbatim turns. Short sessions with no checkpoint keep today's behavior as the fallback.

    This matches the issue's "merged into the system prompt on session recovery" option, and the compressed segment is offloaded via the existing Offloader protocol Agent already uses — so the new workspace parameter accepts the same objects.

    Planned changes

    • New RealtimeContextConfig mirroring Agent's ContextConfig: context_length (explicit count rather than a ratio of an opaque server-side session), tool_result_limit (applied where tool results enter state.context, same truncation semantics as Agent), compression model/prompt/schema/summary-template fields, and trigger_ratio/reserve_ratio rebased onto context_length.
    • workspace: Offloader | None parameter on RealtimeAgent; truncated and compressed segments offloaded after commit.
    • Tests in tests/realtime_agent_test.py following the existing fake-model style: threshold triggering, reconnect injection with/without checkpoint, tool-result truncation, offload calls.

    Two small clarifications I'll need: (1) in the issue's field list compression_prompt appears twice — I read it as compression_model: ChatModelBase + compression_prompt: str template; and (2) whether field naming should mirror Agent (summary_template/summary_schema) or follow the issue's compression_schema/compression_summary literally. I'll default to mirroring Agent unless told otherwise.

    For context, I build semantix, an open-source agent kernel that does exactly this class of work — semantic slicing, context compression and cross-session reuse — so the summarization/offload design here is familiar ground. Happy to post the design doc as a follow-up comment before the PR if that's useful. Could you assign this to me?

  3. wuwenbo0626 commented on Sep 24, 2026

    @wuwenbo0626
    Contributor

    I see that this issue is currently unassigned, but two contributors previously expressed interest. Are either of those claims still active, or is the issue available for a new contributor to take? If it is available, I would be interested in working on it and will align on the two open design questions before starting the implementation.

  4. radianceded commented on Sep 24, 2026

    @radianceded

    My claim is still active — the design for both open questions is in my comment above (token counting via the existing ResponseDoneEvent usage signals with a character-based fallback; a rolling-checkpoint compression model that only commits at turn boundaries).

    To move this forward rather than keep the issue parked, I'll start implementing on a fork and share a draft PR here — the two open questions will be answered concretely in code, so maintainers can judge on a working artifact instead of prose.

    @wuwenbo0626 glad to see the interest — if you'd like to contribute, the token-accounting groundwork (per-provider usage normalization + estimation fallback) is a well-scoped slice that could land first, independent of the compression core. I'll be posting the draft PR shortly.

  5. github-actions commented on Sep 24, 2026

    @github-actions
    Contributor

    It's yours, @radianceded. If there is no pull request and no word from you by 2026-10-08 the claim is released so nobody is blocked — a comment here renews it.

  6. radianceded commented on Sep 25, 2026

    @radianceded

    Draft PR is up: #2838 — the observation layer (usage estimate with provenance/staleness, presence-preserving events across the four adapters, tracker wired into the agent lifecycle) plus RealtimeContextConfig and the workspace offloader parameter. Marked as draft; the compression checkpoint slice (stable-prefix verification, turn-boundary commit, reconnect injection) lands in this same PR next.

    This secures the claim past 2026-10-08. @wuwenbo0626 the token-accounting slice offer from my earlier comment still stands if you want in.

  7. radianceded commented on Sep 25, 2026

    @radianceded

    Compression slice is up on #2838 (9fdd437), implementing the rolling-checkpoint design from my claim comment:

    • Trigger: at a completed-turn boundary, when the usage estimate from the observation layer crosses trigger_ratio of context_length, one background compression task snapshots the settled transcript and summarizes it — single-flight, cancelled on close(), never blocking the audio loop.
    • Commit protocol: the summary lands only when the snapshot prefix is still byte-identical (id list + SHA-256 content digest) and no reply/tool work is in flight. Any drift — barge-in truncation, a late tool result — discards the attempt silently and the next completed turn retries. Chained compressions fold the previous summary into the prompt, so repeated passes stay cumulative rather than myopic.
    • Offload: committed prefixes go through the workspace Offloader (offload_context), and the summary cites the returned path in a system-reminder, matching the text-mode agent's convention.
    • tool_result_limit: enforced in _report_tool before a result enters the session context (chars ≈ tokens/4, the same conservative basis as the usage fallback), with the full result offloaded via offload_tool_result and the truncation reminder citing the path.

    Also exports RealtimeContextConfig from agentscope.agent since it is a public constructor argument. Ten new tests cover trigger gating (threshold and missing estimate), quiet-boundary commit, rejection of modified/shrunk prefixes and in-flight tools, and truncation with and without a workspace.

    Natural next step is moving #2838 out of draft once the reconnect-side summary injection has been exercised a bit more — feedback welcome in the meantime.

  8. radianceded commented on Sep 25, 2026

    @radianceded

    Follow-up: while exercising the reconnect side before moving out of draft, found and fixed a real bug in the fallback path (59e3d2b) — for providers without history replay, the 32k newest-first transcript budget dropped the compression summary first, since it led the rendering. The summary is now pinned ahead of the budget; only verbatim turns are dropped, oldest-first. Pinned by a unit test plus an end-to-end test on the reconnect instructions.

    #2838 is now out of draft — both slices (observation + compression) are complete. Feedback welcome.

  9. wuwenbo0626 commented on Sep 25, 2026

    @wuwenbo0626
    Contributor

    @radianceded I took the token-estimation slice you offered and opened a PR directly against your branch: radianceded#1

    It adds the missing fallback for responses without input_tokens: estimate from the local provider-facing transcript, mark the observation as ESTIMATED, and keep any earlier provider-reported count as a conservative lower bound. Provider usage still replaces the estimate whenever it is available.

    I kept the change out of the compression, checkpoint, reconnect, and adapter-normalization paths. The targeted realtime tests pass (58 passed), and pre-commit passes on the touched files. Please let me know if you would prefer a different estimation boundary or want me to rebase after further changes to #2838.

  10. radianceded commented on Sep 25, 2026

    @radianceded

    Small update: @wuwenbo0626 implemented the transcript-estimation fallback I'd sketched in the claim (for providers that never report usage) and I've reviewed and merged it into #2838's branch — the estimate is a monotone lower bound flagged ESTIMATED, so compression now works on providers without usage reporting without weakening the reported-numbers semantics. Details in his PR's review thread.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Help WantedFurther help required.RoadmapThe development plan

Type

No type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions