Stop your agent contradicting itself.

A long agent session accumulates facts that stopped being true, and tool output nobody will read again. The model then has to reconcile two values for the same key — and you pay for every token of it. contextgc compiles the transcript down to what is currently true, deterministically, before it reaches the model.

pip install contextgc 0 dependencies no model call MIT
What the agent was sent
system: You are an airline support agent.

user: I need to change the cabin on reservation
      0U4NPP. I am in economy class.
assistant: Confirming reservation 0U4NPP.
  <contextgc-state>{"assert":{"active_reservation":
  "0U4NPP","cabin_class":"economy"}}</contextgc-state>
user: Upgrade me to business class.
assistant: Upgrading you now.
  <contextgc-state>{"assert":{"cabin_class":
  "business"}}</contextgc-state>
user: Actually, keep it in economy, I changed
      my mind.
assistant: Reverting to economy.

user: Thanks.
What the model receives — 39.6% smaller
[ACTIVE_AGENT_STATE]
  - active_reservation = "0U4NPP" (turn 2 [declared])
  - cabin_class = "economy" (turn 6 [declared])

system: You are an airline support agent.

assistant: Confirming reservation 0U4NPP.

user: Upgrade me to business class.

user: Actually, keep it in economy, I changed
      my mind.

assistant: Reverting to economy.

user: Thanks.

The business upgrade is gone because the user reversed it, and the register states the value that survived. Nothing was inferred and nothing was invented — the agent said which key it had changed, as a side-effect of a turn it was already making.

Measured on 180 real agent trajectories

Not a benchmark of the compiler grading itself. Three public corpora, real model output, and a hand-labelled sample for the precision figure. Every number below ships with the n it came from.

Corpus Sample Token reduction Turns retired Payloads compacted Precision
Coding
SWE-agent bug-fixing, 20 repos
40 transcripts66.5%161668100% n=36, CI 90–100%
Airline support
APIGen-MT-5k, reservation tool use
60 conversations59.0%9138100% n=72, CI 95–100%
Retail support
APIGen-MT-5k, order and exchange tool use
80 conversations62.5%2275100% n=47, CI 92–100%

Reduction transfers across all three; supersession does not. Coding re-asserts a key 77 times across 40 transcripts, retail twice across 80 — so the two support corpora gain most of their reduction from compacting tool output, not from retiring turns. Reported separately because averaging them would hide it.

Four things it does

Each one is a total function over the transcript. There is no model in the loop, so the same input always compiles to the same output.

It lets the agent declare its own state

Regexes cannot resolve “send it to the new place instead”. So the agent reports what it concluded, as a structured side-effect of the turn it was already making — no extra model call, no extra latency.

<contextgc-state>{
  "assert": {"destination_address": "Gate 2"},
  "pin":    {"allergy": "peanut"},
  "revoke": ["order_id"],
  "unsure": {"rider": "maybe west"}
}</contextgc-state>

It tracks what each turn asserted

Entity state is last-write-wins. When a key is re-asserted, the earlier turn is retired from the context instead of competing with the newer one.

destination_address: "Gate 2"
gate_code:          "4921"

It never retires the evidence

If a turn is superseded for one key but is still the only support for another, it stays. Retiring it would hand the model a value with nothing behind it.

retired:        [1, 5]
still required: turn 2 supports refund_claim
orphaned facts: none

It compacts tool output, keeping the safety bits

Bulk rows are sampled, but any row carrying an allergen, severity, PII or expiry signal is always retained — and the omission is reported rather than performed quietly.

[inventory: 15 rows. Showing 4.
SAFETY-FLAGGED ROWS RETAINED: [{"name":
 "Peanut Butter 500g", "warning":
 "PEANUT_ALLERGEN"}]  11 non-safety
 rows omitted (truncated=true).]

What it does not do

The failure modes are documented because you will hit them, and a tool that only lists its strengths is not one you can rely on.

  • It does not make models reliable. It removes contradictions from the input. It cannot add judgement, and it will not stop a model ignoring a rule you pinned.
  • Read-path state tracking is regular expressions. There is no coreference resolution, so a fact expressed only through pronouns will not be tracked. Missed extraction is the main failure mode: the state is silently incomplete rather than loudly wrong. Extend it with StateDAG.register_entity_schema().
  • The write path is unproven, not working. On 6 real trajectories a 14B model names a schema slot 96% of the time and emits well-formed JSON every turn — and every resulting addition is still a value of the wrong shape: three literal Nones, and error messages where a test identifier belongs. Checking the key is not checking the value — so a schema may now state a value contract per slot, which drops 24 of the 55 declarations in that capture. It removes an invisible failure; it does not prove the rest.
  • Token counts are chars/4, not a BPE tokenizer. Read them as a ratio, not a bill. Exact figures need tiktoken against your model.
  • Retired-turn recall is lexical — token, subword and phrase overlap, not semantics. “What did we decide about billing?” will not find a turn that only says “invoice”.

Paste a transcript and see what changes

No key, no account, nothing stored. The console compiles server-side and shows the exact messages the model would receive, with the reasoning behind every number.

Open the console