Repository navigation
feat(realtime): support context management for the realtime agent #2566
Description
Activity
Hi, I’d like to work on this issue. I’ll first investigate the Realtime API’s token-usage capabilities and the concurrency model for context compression. I’ll discuss the possible technical approaches to the two open questions here before proposing an implementation aligned with the existing Agent context-management behavior.
Please assign this issue to me if the direction sounds good. Thanks!Hi @maintainers — I'd like to take this one. I've read through
RealtimeAgent(agent/_realtime/_agent.py),Agent's context pipeline (agent/_agent.py,agent/_config.py) and the realtime event/metrics code, and I can propose concrete answers to both open questions before starting implementation.Q1 — Token counting. We already have a per-turn signal:
ResponseDoneEvent.input_tokens/output_tokens(realtime/_events.py), surfaced intoTurnMetricsin the downlink pump. I propose maintaining a running estimate as:- base = provider-reported usage accumulated per completed turn (cumulative when a provider reports session-cumulative counts, summed deltas when it reports per-turn);
- fallback = character-based estimation over the transcript delta in
state.contextfor any turn where the provider omitted usage (fields default to 0), so the trigger still works on providers without usage reporting.
The estimate would live alongside
TurnMetricsand gate both truncation and compression againstcontext_length.Q2 — Compression timing / concurrency. A "rolling checkpoint" model, so compression never blocks the audio loop:
- At
trigger_ratio, spawn one background compression task (single-flight, cancelled onclose()) over an immutable snapshot ofstate.contextup to a monotonic sequence number. - The summary commits only at a turn boundary (after
ResponseDoneEvent, before the next user speech start): the summarized prefix is atomically replaced by the summary message, andsummarized_upto_seqis recorded. Turns newer than the checkpoint stay verbatim. - On reconnect,
connect()(which currently unconditionally appends the raw transcript as "## Conversation so far",connect()L263-272) prepends the latest committed checkpoint summary to the instructions instead, followed by only the post-checkpoint verbatim turns. Short sessions with no checkpoint keep today's behavior as the fallback.
This matches the issue's "merged into the system prompt on session recovery" option, and the compressed segment is offloaded via the existing
OffloaderprotocolAgentalready uses — so the newworkspaceparameter accepts the same objects.Planned changes
- New
RealtimeContextConfigmirroringAgent'sContextConfig:context_length(explicit count rather than a ratio of an opaque server-side session),tool_result_limit(applied where tool results enterstate.context, same truncation semantics asAgent), compression model/prompt/schema/summary-template fields, andtrigger_ratio/reserve_ratiorebased ontocontext_length. workspace: Offloader | Noneparameter onRealtimeAgent; truncated and compressed segments offloaded after commit.- Tests in
tests/realtime_agent_test.pyfollowing the existing fake-model style: threshold triggering, reconnect injection with/without checkpoint, tool-result truncation, offload calls.
Two small clarifications I'll need: (1) in the issue's field list
compression_promptappears twice — I read it ascompression_model: ChatModelBase+compression_prompt: strtemplate; and (2) whether field naming should mirrorAgent(summary_template/summary_schema) or follow the issue'scompression_schema/compression_summaryliterally. I'll default to mirroringAgentunless told otherwise.For context, I build semantix, an open-source agent kernel that does exactly this class of work — semantic slicing, context compression and cross-session reuse — so the summarization/offload design here is familiar ground. Happy to post the design doc as a follow-up comment before the PR if that's useful. Could you assign this to me?
I see that this issue is currently unassigned, but two contributors previously expressed interest. Are either of those claims still active, or is the issue available for a new contributor to take? If it is available, I would be interested in working on it and will align on the two open design questions before starting the implementation.
My claim is still active — the design for both open questions is in my comment above (token counting via the existing
ResponseDoneEventusage signals with a character-based fallback; a rolling-checkpoint compression model that only commits at turn boundaries).To move this forward rather than keep the issue parked, I'll start implementing on a fork and share a draft PR here — the two open questions will be answered concretely in code, so maintainers can judge on a working artifact instead of prose.
@wuwenbo0626 glad to see the interest — if you'd like to contribute, the token-accounting groundwork (per-provider usage normalization + estimation fallback) is a well-scoped slice that could land first, independent of the compression core. I'll be posting the draft PR shortly.
Reacted by WuWenbogithub-actions commented
on Sep 24, 2026 on Sep 24, 2026 – with GitHub ActionsContributorMore actionsIt's yours, @radianceded. If there is no pull request and no word from you by 2026-10-08 the claim is released so nobody is blocked — a comment here renews it.
Draft PR is up: #2838 — the observation layer (usage estimate with provenance/staleness, presence-preserving events across the four adapters, tracker wired into the agent lifecycle) plus
RealtimeContextConfigand theworkspaceoffloader parameter. Marked as draft; the compression checkpoint slice (stable-prefix verification, turn-boundary commit, reconnect injection) lands in this same PR next.This secures the claim past 2026-10-08. @wuwenbo0626 the token-accounting slice offer from my earlier comment still stands if you want in.
Reacted by WuWenbo- added a commit that references this issue
on Sep 25, 2026 Compression slice is up on #2838 (9fdd437), implementing the rolling-checkpoint design from my claim comment:
- Trigger: at a completed-turn boundary, when the usage estimate from the observation layer crosses
trigger_ratioofcontext_length, one background compression task snapshots the settled transcript and summarizes it — single-flight, cancelled onclose(), never blocking the audio loop. - Commit protocol: the summary lands only when the snapshot prefix is still byte-identical (id list + SHA-256 content digest) and no reply/tool work is in flight. Any drift — barge-in truncation, a late tool result — discards the attempt silently and the next completed turn retries. Chained compressions fold the previous summary into the prompt, so repeated passes stay cumulative rather than myopic.
- Offload: committed prefixes go through the workspace
Offloader(offload_context), and the summary cites the returned path in asystem-reminder, matching the text-mode agent's convention. tool_result_limit: enforced in_report_toolbefore a result enters the session context (chars ≈ tokens/4, the same conservative basis as the usage fallback), with the full result offloaded viaoffload_tool_resultand the truncation reminder citing the path.
Also exports
RealtimeContextConfigfromagentscope.agentsince it is a public constructor argument. Ten new tests cover trigger gating (threshold and missing estimate), quiet-boundary commit, rejection of modified/shrunk prefixes and in-flight tools, and truncation with and without a workspace.Natural next step is moving #2838 out of draft once the reconnect-side summary injection has been exercised a bit more — feedback welcome in the meantime.
- Trigger: at a completed-turn boundary, when the usage estimate from the observation layer crosses
Follow-up: while exercising the reconnect side before moving out of draft, found and fixed a real bug in the fallback path (59e3d2b) — for providers without history replay, the 32k newest-first transcript budget dropped the compression summary first, since it led the rendering. The summary is now pinned ahead of the budget; only verbatim turns are dropped, oldest-first. Pinned by a unit test plus an end-to-end test on the reconnect instructions.
#2838 is now out of draft — both slices (observation + compression) are complete. Feedback welcome.
@radianceded I took the token-estimation slice you offered and opened a PR directly against your branch: radianceded#1
It adds the missing fallback for responses without
input_tokens: estimate from the local provider-facing transcript, mark the observation asESTIMATED, and keep any earlier provider-reported count as a conservative lower bound. Provider usage still replaces the estimate whenever it is available.I kept the change out of the compression, checkpoint, reconnect, and adapter-normalization paths. The targeted realtime tests pass (58 passed), and pre-commit passes on the touched files. Please let me know if you would prefer a different estimation boundary or want me to rebase after further changes to #2838.
Small update: @wuwenbo0626 implemented the transcript-estimation fallback I'd sketched in the claim (for providers that never report usage) and I've reviewed and merged it into #2838's branch — the estimate is a monotone lower bound flagged
ESTIMATED, so compression now works on providers without usage reporting without weakening the reported-numbers semantics. Details in his PR's review thread.Reacted by WuWenbo
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsTodo
Background
Context management for
RealtimeAgentdiffers fundamentally from the standardAgentclass:Agenthandles thisAgentbehaviorChanges
ContextConfigclass forRealtimeAgent, mirroringAgent'scontext_config, with the following fields:context_length: maximum allowed context sizetool_result_limit: cap on individual tool result sizecompression_prompt: aChatModelBaseinstance used for compressioncompression_prompt: prompt template for compressioncompression_schema: structured outputBaseModelfor compression resultcompression_summary: the summary templateworkspaceparameter toRealtimeAgentfor offloading truncated and compressed tool results