Context Management — ML Junction docs
Context fitting reserves output space and trims input only when necessary. It is atomic around semantic units: system instructions, the current user turn, and assistant/tool chains can be preserved together.
Field · Default · Meaning mode · auto_fit · auto_fit trims safely; strict rejects an oversized input input_budget_ratio · 0.70 · Fraction of available context targeted for input min_output_reserve_tokens · 4096 · Minimum context reserved for generation preserve_system · true · Retain system and developer instructions preserve_current_user · true · Retain the latest user request preserve_tool_chains · true · Keep tool calls with their corresponding results
{
"model": "gemini-3.5-flash",
"messages": [
{ "role": "user", "content": "Analyze the supplied long document." }
],
"requirements": {
"min_context_tokens": 500000
},
"routing": {
"preset": "long_context"
},
"context": {
"mode": "strict",
"min_output_reserve_tokens": 12000
}
}
Canonical URL: https://mljunction.com/docs/context-management