Context Management — ML Junction docs

Context fitting reserves output space and trims input only when necessary. It is atomic around semantic units: system instructions, the current user turn, and assistant/tool chains can be preserved together.

Field · Default · Meaning
mode · auto_fit · auto_fit trims safely; strict rejects an oversized input
input_budget_ratio · 0.70 · Fraction of available context targeted for input
min_output_reserve_tokens · 4096 · Minimum context reserved for generation
preserve_system · true · Retain system and developer instructions
preserve_current_user · true · Retain the latest user request
preserve_tool_chains · true · Keep tool calls with their corresponding results
{
  "model": "gemini-3.5-flash",
  "messages": [
    { "role": "user", "content": "Analyze the supplied long document." }
  ],
  "requirements": {
    "min_context_tokens": 500000
  },
  "routing": {
    "preset": "long_context"
  },
  "context": {
    "mode": "strict",
    "min_output_reserve_tokens": 12000
  }
}

Canonical URL: https://mljunction.com/docs/context-management