Support tool callbacks in MCP sampling by EronWright · Pull Request #2998 · docker/docker-agent

EronWright · 2026-06-04T02:11:21Z

Summary

Closes the tool callbacks functional gap in MCP sampling support — a follow-up to #2815, addressing one of the remaining items from #2809.

When an MCP server includes a tools array in a sampling/createMessage request, the host now drives its model with those tools and returns any tool_use blocks back to the server as ToolUseContent. The server remains responsible for executing the tool and continuing the loop in a follow-up sampling request.

sequenceDiagram
     participant H as cagent
      participant S as MCP Server
      participant L as LLM

      activate H
      H->>+S: tools/call {name, arguments}

      note over S: needs LLM inference

      S->>+H: sampling/createMessage<br/>{messages, tools: [...]}
      H->>+L: chat completion
      L-->>-H: ToolUseContent<br/>stopReason: "toolUse"
      H-->>-S: CreateMessageResult<br/>{tool_use, stopReason: "toolUse"}

      note over S: executes tool locally

      S->>+H: sampling/createMessage<br/>{messages + tool_use + tool_result, tools: [...]}
      H->>+L: chat completion
      L-->>-H: TextContent<br/>stopReason: "endTurn"
      H-->>-S: CreateMessageResult<br/>{text, stopReason: "endTurn"}

      S-->>-H: tool result
      deactivate H

What's new

New SamplingWithToolsHandler type and SampleableWithTools interface — additive, parallel to the existing SamplingHandler / Sampleable. No breaking changes to the basic sampling path merged in feat(mcp): add sampling/createMessage support #2815.
MCP toolset wires both handler types. At Initialize, exactly one of the SDK's mutually exclusive ClientOptions.CreateMessage* fields is populated — prefer with-tools when registered, fall back to basic.
Capability handshake advertises sampling.tools so servers know the host can receive tool-enabled requests.
Runtime handler (pkg/runtime/sampling.go):
- Converts V2 multi-block messages: text, image/audio, tool_use → assistant ToolCalls, tool_result → MessageRoleTool rows (parallel tool_results expand to multiple chat.Message rows).
- Converts []*mcp.Tool → []tools.Tool with a no-op handler (the server, not the host, executes).
- Drives model.CreateChatCompletionStream, aggregates streamed tool calls.
- Builds result Content with TextContent + ToolUseContent blocks; stopReason: "toolUse" when tool calls are present.
New limits: maxSamplingTools=64, maxSamplingToolCalls=32.
End-to-end test (e2e/sampling_test.go): mounts an in-process gomcp.NewServer on an httptest server via StreamableHTTPHandler. The server exposes one tool (ask_with_calculator) whose handler drives a real sampling-with-tools loop against the connecting cagent. The Gemini side is recorded once and replayed on subsequent runs, so the test runs offline in CI.

Out of scope (separate gaps from #2809)

Human-in-the-loop approval UI
Model-preference hints

Test plan

Adds a parallel SamplingWithToolsHandler alongside the existing SamplingHandler so MCP servers can include a tools array in sampling/createMessage requests. The host drives its model with those tools and returns any tool_use blocks as ToolUseContent; the server remains responsible for executing the tool and continuing the loop in a follow-up sampling request. The initialize handshake now advertises sampling.tools capability, and the MCP toolset selects the appropriate go-sdk handler (basic vs. with-tools) based on which handler is registered.

Mounts an in-process gomcp.NewServer on an httptest server via StreamableHTTPHandler. Its one tool, ask_with_calculator, runs a sampling loop: sends sampling/createMessage with a calculator tool, gets a tool_use back from the host LLM, "executes" the calculator, sends a follow-up sampling request carrying the tool_result, and returns the final text. The Gemini side is recorded once and replayed on subsequent runs, so the test runs offline in CI.

aheritier · 2026-06-08T06:57:30Z

/review

docker-agent

Assessment: 🟡 NEEDS ATTENTION

Two medium-severity findings in the newly-added sampling-with-tools code. The core stream-aggregation logic, capability handshake, content building, and limits enforcement are all well-structured and correctly tested.

- Reject tool_use blocks under a non-assistant role explicitly rather than silently dropping the message. The MCP spec places tool_use on assistant turns, but a malformed server would previously have its message disappear from the converted chat history with no error. - Document the ordering dependency between SetSampling*Handler and Initialize in both stdio and remote MCP client setups, so future maintainers don't reorder them and silently lose the sampling handler on reconnect.

Use errors.New for the static loop-exhausted message and strconv.FormatInt for the calculator result instead of fmt.Errorf / fmt.Sprintf %d.

dgageot · 2026-06-10T15:19:19Z

Sorry @EronWright, our project requires all the commit to be signed and verified. Could you sign them? Thanks!

EronWright marked this pull request as ready for review June 7, 2026 00:56

EronWright requested a review from a team as a code owner June 7, 2026 00:56

aheritier added the area/providers For features/issues/fixes related to LLM providers (Bedrock, LiteLLM, Qwen, custom, etc.) label Jun 7, 2026

docker-agent reviewed Jun 8, 2026

View reviewed changes

Comment thread pkg/runtime/sampling.go Outdated

Comment thread pkg/tools/mcp/remote.go

aheritier removed the area/providers For features/issues/fixes related to LLM providers (Bedrock, LiteLLM, Qwen, custom, etc.) label Jun 8, 2026

EronWright marked this pull request as draft June 8, 2026 15:03

EronWright requested a review from docker-agent June 8, 2026 20:35

EronWright marked this pull request as ready for review June 8, 2026 20:38

aheritier added area/testing Test infrastructure, CI/CD, test runners, evaluation area/providers/gemini Google Gemini provider support labels Jun 8, 2026

Fix perfsprint lint in sampling e2e test

b1fad0b

Use errors.New for the static loop-exhausted message and strconv.FormatInt for the calculator result instead of fmt.Errorf / fmt.Sprintf %d.

dgageot approved these changes Jun 10, 2026

View reviewed changes

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Support tool callbacks in MCP sampling#2998

Support tool callbacks in MCP sampling#2998
EronWright wants to merge 4 commits into
docker:mainfrom
EronWright:sampling-tools

EronWright commented Jun 4, 2026 •

edited

Loading

Uh oh!

aheritier commented Jun 8, 2026

Uh oh!

docker-agent left a comment

Uh oh!

Uh oh!

Uh oh!

dgageot commented Jun 10, 2026

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

4 participants

Conversation

EronWright commented Jun 4, 2026 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Summary

What's new

Out of scope (separate gaps from #2809)

Test plan

Uh oh!

aheritier commented Jun 8, 2026

Uh oh!

docker-agent left a comment

Choose a reason for hiding this comment

Assessment: 🟡 NEEDS ATTENTION

Uh oh!

Uh oh!

Uh oh!

dgageot commented Jun 10, 2026

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

4 participants

EronWright commented Jun 4, 2026 •

edited

Loading