A long agent session accumulates facts that stopped being true, and tool output nobody will read again. The model then has to reconcile two values for the same key — and you pay for every token of it. contextgc compiles the transcript down to what is currently true, deterministically, before it reaches the model.
pip install contextgc
0 dependencies
no model call
MIT
system: You are an airline support agent.
user: I need to change the cabin on reservation
0U4NPP. I am in economy class.
assistant: Confirming reservation 0U4NPP.
<contextgc-state>{"assert":{"active_reservation":
"0U4NPP","cabin_class":"economy"}}</contextgc-state>
user: Upgrade me to business class.
assistant: Upgrading you now.
<contextgc-state>{"assert":{"cabin_class":
"business"}}</contextgc-state>
user: Actually, keep it in economy, I changed
my mind.
assistant: Reverting to economy.
user: Thanks.
[ACTIVE_AGENT_STATE]
- active_reservation = "0U4NPP" (turn 2 [declared])
- cabin_class = "economy" (turn 6 [declared])
system: You are an airline support agent.
assistant: Confirming reservation 0U4NPP.
user: Upgrade me to business class.
user: Actually, keep it in economy, I changed
my mind.
assistant: Reverting to economy.
user: Thanks.
The business upgrade is gone because the user reversed it, and the
register states the value that survived. Nothing was inferred and nothing was
invented — the agent said which key it had changed, as a side-effect of a turn
it was already making.
Not a benchmark of the compiler grading itself. Three public corpora, real model
output, and a hand-labelled sample for the precision figure. Every number below
ships with the n it came from.
| Corpus | Sample | Token reduction | Turns retired | Payloads compacted | Precision |
|---|---|---|---|---|---|
| Coding SWE-agent bug-fixing, 20 repos | 40 transcripts | 66.5% | 161 | 668 | 100% n=36, CI 90–100% |
| Airline support APIGen-MT-5k, reservation tool use | 60 conversations | 59.0% | 9 | 138 | 100% n=72, CI 95–100% |
| Retail support APIGen-MT-5k, order and exchange tool use | 80 conversations | 62.5% | 2 | 275 | 100% n=47, CI 92–100% |
Reduction transfers across all three; supersession does not. Coding re-asserts a key 77 times across 40 transcripts, retail twice across 80 — so the two support corpora gain most of their reduction from compacting tool output, not from retiring turns. Reported separately because averaging them would hide it.
Each one is a total function over the transcript. There is no model in the loop, so the same input always compiles to the same output.
Regexes cannot resolve “send it to the new place instead”. So the agent reports what it concluded, as a structured side-effect of the turn it was already making — no extra model call, no extra latency.
<contextgc-state>{
"assert": {"destination_address": "Gate 2"},
"pin": {"allergy": "peanut"},
"revoke": ["order_id"],
"unsure": {"rider": "maybe west"}
}</contextgc-state>
Entity state is last-write-wins. When a key is re-asserted, the earlier turn is retired from the context instead of competing with the newer one.
destination_address: "Gate 2" gate_code: "4921"
If a turn is superseded for one key but is still the only support for another, it stays. Retiring it would hand the model a value with nothing behind it.
retired: [1, 5] still required: turn 2 supports refund_claim orphaned facts: none
Bulk rows are sampled, but any row carrying an allergen, severity, PII or expiry signal is always retained — and the omission is reported rather than performed quietly.
[inventory: 15 rows. Showing 4.
SAFETY-FLAGGED ROWS RETAINED: [{"name":
"Peanut Butter 500g", "warning":
"PEANUT_ALLERGEN"}] 11 non-safety
rows omitted (truncated=true).]
The failure modes are documented because you will hit them, and a tool that only lists its strengths is not one you can rely on.
StateDAG.register_entity_schema().Nones, and error messages where a test identifier
belongs. Checking the key is not checking the value — so a schema may now
state a value contract per slot, which drops 24 of the 55 declarations in that
capture. It removes an invisible failure; it does not prove the rest.chars/4, not a BPE tokenizer. Read them
as a ratio, not a bill. Exact figures need tiktoken against your
model.No key, no account, nothing stored. The console compiles server-side and shows the exact messages the model would receive, with the reasoning behind every number.