Why
- An MCP tool call may return a large amount of data, which can quickly fill the available context window.
- The agent may not be able to control how much data a server returns by adjusting tool arguments alone. Fetching a single pull request can still return a very large
descriptionfield.
How it works
Two complementary thresholds govern when tool responses are written to the sandbox instead of staying in context: A single tool response that’s too large. If an individual tool response exceeds the per-call threshold (default 6,000 tokens), the full result is written to a file in the sandbox and replaced in context with a short preview (default the first and last 100 characters) plus the file path. Parallel tool calls returning together. When the agent fires several tool calls in parallel, no single response may be over the per-call threshold but their combined size can still flood the context window. If the combined tool-call content exceeds the total threshold (default 10,000 tokens), the harness offloads responses one at a time — starting with the largest — until the total drops below the limit. The smallest responses stay inline and the largest end up on disk. In every case the offloaded content remains accessible — the agent can read,grep, or parse the saved file from the sandbox whenever it needs the original data, typically using Code Mode to extract just the fields it needs.
Example
User
Can you please list the mcp servers I have access to
Assistant
Let me get more information about the user.
get_me
get_me
Tool CallTool Response
Assistant
Let me list the MCP servers you have access to.
list_mcp_servers
list_mcp_servers
Tool CallTool Response (blocked) — Full result saved to
/tmp/tool-id-xyz-outputAssistant
Let me get a preview of the data and print the schema.
sandbox
sandbox
Tool CallTool Response
Assistant
Now I understand the schema of the MCP servers output file. Let me write code to parse the MCP server names.
sandbox
sandbox
Tool CallTool Response
Assistant
Hey John, here are the MCP servers you have access to:
- test-realtime-global
- slack-mcp
- google-calendar-mcp
- …
Configuration
Turn offloading on or off per agent:config.context_management in the agent spec reference.