We measure context and orchestration overhead in bounded workflows. These results show what happened in specific first-party tests, not a promise about every model, task or invoice.
COSMOS keeps useful state between sessions, retrieves a focused working set and returns compact evidence. Different layers save different resources: input tokens, transported bytes and tool calls must be measured separately.
Measured first-party result
Input context
Before
69,393
After
4,547
Reduction
93.45%
input tokens
Scope & method
One exact-output synthetic canary compared inherited and isolated minimal provider profiles. Both returned the required exact answer; cache states differed. Input tokens only, not billed savings or whole-job quality.
What this does not prove
Cache states differed. This does not isolate retrieval quality or establish equivalent provider billing.
Evidence reviewed: · Trials: 1
Measured first-party result
Compact result payload
Before
4,290
After
712
Reduction
83.4%
result bytes
Scope & method
Ten local development trials of five file comparisons compared already-compact results with a batched procedure. The serialized successful result bytes were 4,290 and 712 per trial. Same-snapshot equality outcomes matched; setup, failures and procedure construction were excluded.
What this does not prove
Result bytes are not full wire traffic or model tokens. October results are a local development candidate, not a release claim.
Evidence reviewed: · Trials: 10
Measured first-party result
Signed diagnostic traffic
Before
121,896
After
82,455
Reduction
32.36%
wire bytes
Scope & method
Ten neutral local diagnostic scenarios measured 121,896 and 82,455 actual response wire bytes in aggregate. Four scenarios needed more bytes and one extra call. Shared setup was excluded; final-model-answer quality was not compared.
What this does not prove
Wire bytes do not establish model-token usage or actual costs. October results are local development candidates, not a release claim.
Evidence reviewed: · Trials: 10
Measured first-party result
Diagnostic calls
Before
59
After
49
Reduction
16.95%
tool calls
Scope & method
The same ten local diagnostic scenarios measured 59 and 49 tool calls in aggregate. The total fell, but individual scenarios did not all improve.
What this does not prove
Fewer calls do not by themselves prove lower latency, lower cost or equal performance on other tasks.
Evidence reviewed: · Trials: 10
Our measurement commitment
For a pilot we agree on a baseline, representative tasks, quality checks and success criteria before comparing results. We report cache assumptions, output, retries and tool usage; we do not turn a single canary into a universal claim.
Real workflows, clearly bounded
These are first-party development and operational examples. They are not independent customer endorsements, and an operational example is not automatically a measured token-saving case.
Operational pattern
Marketplace evidence → accounting verification
COSMOS keeps source records in the trusted boundary, prepares a candidate, runs parity checks and verifies approved accounting actions by reading back the external state. The lesson is controlled execution and less repeated investigation; no token percentage is claimed for this workflow.
Development pattern
Focused development handoffs
A specialist receives the relevant files, constraints and checks rather than the entire conversation history. Structured results return to one reviewable task. Context reduction is evaluated per task, with tests protecting the required outcome.
Illustrative workflow
Research → translation → review
Research, drafting, translation and editorial review use focused roles and source references. Publication remains a separate approved action. Savings depend on the source volume, handoffs and selected providers.
MCP / API / IDE
Your environment. A shared control plane.
Connect the tools you already use through MCP and configured provider adapters. Protocol compatibility is not the same as an end-to-end verified native integration.
Local adoption or configured adapter; acceptance required
Local adoption and configured hosts
VS Code
Native Codex
Cursor
Claude / Claude Code
VS Code and native Codex have recorded local adoption paths. Cursor and native Claude / Claude Code have harness and MCP configuration adapters; a configured adapter is not host acceptance. A fresh host chat, version, permissions and workload are verified separately.
Protocol-compatible or candidate; pilot validation required
MCP-compatible hosts and candidates
MCP-compatible hosts
OpenClaw
Other MCP-compatible environments can use the governed MCP surface subject to transport, permissions and tool admission. OpenClaw is a candidate for a scoped external integration pilot, not a verified bundled integration; its gateway and runtime are not embedded.
Synthetic evidence, adapter contracts and planned routing
Providers and custom pipelines
Z.AI / GLM
OpenRouter
Python / CLI / MCP pipelines
Z.AI / GLM passed an exact-output synthetic canary; this does not authorize private-data routing or establish general model quality. OpenRouter routing and evaluation are planned, not currently admitted by the closed provider profiles. Custom pipelines can use Python, CLI and MCP contracts around individually validated domain adapters.
Deploy around your boundary
We discuss local, private-server and managed deployment around your data and operational constraints. A scoped demo establishes the required IDE, model, tools, permissions and evidence before a rollout.
Public reproducibility is next
A public GitHub repository with a sanitized benchmark corpus and reproduction instructions is in the backlog. It is not published yet. Today we can discuss the measurement method and agree on a repeatable pilot; no broken repository link is presented.
A SMALLER FIRST STEP
Start with one workflow. Define what “done” means.
A pilot discussion starts with the work, not a platform migration. Agree on the scope, access boundaries and a useful test before deciding what to build.
01
Map the handoffs
Name the sources, agent roles and decisions that need a person.
02
Define the proof
Choose a baseline, acceptance checks and the limits of the result.
03
Review the next step
Use the evidence to decide whether to refine, expand or stop.
A useful first candidate
A recurring task with accessible evidence and a result you can inspect. Keep the first scope small enough to review.
Keep outside the first scope
Unrestricted access, unsupervised irreversible actions or a promise that every model answer will be correct.
Prepare a short enquiry
Choose a starting point and describe one outcome. Copy the brief when you are ready; you decide what to share.
This builder works only on this page. Nothing is stored, analysed or automatically sent. Do not enter passwords or sensitive information.
Will COSMOS guarantee the same savings for my team?
No. We measure the actual workload and check the required outcome. Context, cache, model pricing, tools and retries can change the result.
Can I use my current IDE and models?
COSMOS supports native provider paths, MCP-compatible environments and configured adapters. The exact host, model, permissions and deployment are validated for your workflow.
Where is the public benchmark repository?
Public GitHub publication is in the backlog. No public repository or independent reproduction is claimed on this page.
Measure your workflow with us.
Bring one task, a baseline and the outcome that must remain correct. We can show the workflow, discuss your stack and define a bounded demo or pilot.