Reduce repeated context and token use in legal AI workflows
How Relex uses summary-first MCP reads, ontology digests and saved conclusions to reduce unnecessary legal context. Includes a fair measurement method.
Why do legal AI conversations keep using so much context?
Repeatedly uploading a case bundle, replaying long transcripts and rebuilding instructions can consume input context without adding new facts. A persistent workspace lets the next session start from saved legal work and fetch the source detail it actually needs. Persistence alone does not make model processing free or remove context limits.
How does Relex reduce unnecessary context?
The hosted MCP exposes search and execute. Search returns relevant operations and compact parameters instead of requiring the whole API catalog upfront. A single-matter read defaults to a summary rather than the full aggregated timeline; the case ontology defaults to a bounded digest. Agents can request more detail. Concluding a steering session saves its result onto the continuing matter record so the next assistant can consult it.
Does this guarantee a lower bill?
No. It creates a practical route to fewer unnecessary input tokens, not a published benchmark or guaranteed saving. Tool calls, output and reasoning usage, extra source reads, provider caching and Relex usage all affect total cost. A fixed-price assistant subscription may not become cheaper. De-identification protects identities; it is not itself a token-compression method.
How can a firm measure the benefit?
Use the same authorized or synthetic matter, task, model and output requirements. Compare repeated full-context loading with summary-first retrieval and saved conclusions across several sessions. Record input, cached input, output and any reported reasoning tokens, tool calls, latency, total charges and missing or unsupported findings. Count the initial setup and every retrieval, not just the final prompt. Retain source access and human review; do not trade away material facts to improve a token score.