Savings calculator

See what you'd stop paying for.

Enter what you spend on model calls each month. We apply the reduction we actually measured on production — not a projection.

$

Calculated at a 90% reduction against carrying your full project context on every call. We measured 89–93% once a project holds around a hundred facts, and 83% averaged across every run including nearly empty stores. We apply the round, conservative number.

You'd pay instead
$10,000

$90,000 saved every month

Over a year$1,080,000
Reduction applied90%
Answer qualityunchanged
Why the bill looks like this

A model has no memory, so you buy it one every time.

Every call to every model starts from nothing. The model does not remember your codebase, your decisions, your customers, or the conversation you had with it an hour ago. The only way it can know any of that is if you send it again — and you are billed for every word of it, every single time. That resend is not a small line item. On a mature project it is most of the bill.

The resend

You pay for the same paragraphs on every request.

The context that makes your agent useful — conventions, architecture, prior decisions — is not stored anywhere the model can reach. It rides along on each call, and it is charged at full price each time, even though nothing about it changed.

The wrong curve

Cost grows with what the project knows, not with what you asked.

A one-line question against a mature project costs many times what the same question costs on day one. The work did not get harder. The payload did. Left alone, the price of every interaction rises for as long as the project succeeds.

The waste

Almost none of what you send is relevant to the question asked.

When everything must travel on every call, the overwhelming majority of each payload has nothing to do with the request. You are paying to transmit context the model will not use — and burying the part it needs inside it.

What changes

Your project keeps what it knows. Each request carries only its share.

Iconia gives the work a memory that outlives the conversation. What your project knows is held once, and each request is accompanied by the part of it that request actually needs — not the whole archive. The model still receives everything necessary to answer well. It simply stops being handed everything else.

Smaller calls

The payload stops tracking the archive.

What you send reflects the question, not the accumulated history of the project. This is where the money goes back in your pocket, and it is the reason the saving widens rather than narrows as you keep working.

Better answers

Knowing beats guessing.

On questions answerable only from project knowledge, models with no memory got 13% right. With Iconia, the same models on the same questions got 100%. Cheaper is only worth having if the answers hold — these got better, not worse.

Any model

What you know stops being locked to who you rent it from.

The same memory serves Claude, GPT, Gemini, Llama and models you run yourself. Switch providers for price or capability without rebuilding the context that makes your agent worth using.

The shape of the saving

It compounds, because the alternative does too.

The reduction is not a fixed discount. It is the gap between a payload that grows and one that does not — so the longer a project runs, the wider it gets. Measured against carrying full context, at three project sizes:

What the project knowsCarry everythingWith Iconia ReductionAnswers correct
A new project · 8 facts$0.0079$0.0058 19–45%9 / 9
A month in · 40 facts$0.0310$0.0071 75–86%9 / 9
A real codebase · 100 facts$0.0718$0.0072 89–93%9 / 9

Read the second column: carrying everything rose ninefold as the project matured. The Iconia column barely moved. That flatness is the whole product — and note the last column, which never moved either.

Measured on production

Every number, and what it was measured against.

A number without its baseline is not evidence. Each of these carries the comparison it came from, and every one is reproducible.

Answers correct
100%
Up from 13% with no memory. Five models, thirty questions, on facts knowable only from the project.
Cost reduction
90%+
Against carrying full context on every call, once a project holds around a hundred facts.
Recall at 1,000 memories
100%
The right memory ranked first — not merely present — at every store size tested.
Added latency
<10ms
Flat as the store grows. Below the threshold where a person notices anything at all.
Models verified
Anthropic · OpenAI
Google · Meta
Plus models you host yourself. Same memory, same result, whichever you run.
Setup
2 lines
Point at Iconia, add your key. No SDK, no rewrite, no migration, no change to how you work.
Start free — keep your own model bill

1,000 memories free. Your provider keys stay yours.

Read this before you quote us

What these numbers do not say.

The figures above are real and reproducible. They are also bounded, and we would rather you learn the bounds here than discover them later.

A new project saves less.

The saving comes from what you would otherwise resend. If your project knows almost nothing yet, there is little to avoid sending — 19–45% at eight facts. The calculator assumes a project with real accumulated knowledge.

Small local models are not covered.

Memory reaches a 7B model correctly and it answers well — until a large agent tool payload is also present, at which point the model cannot hold everything at once. That is a limit of the model, not of the memory, but it is a real limit.

The samples are modest.

Six to nine questions per condition, consistent and repeatable across models and runs. Directionally strong; not a published benchmark suite, and we will not present it as one.

Fewer round trips is unproven.

We have measured that each call gets smaller. Whether an agent also finishes a task in fewer steps is plausible, untested, and therefore not claimed.

Questions

The things people ask before they try it.

Am I paying you instead of paying the model provider?

You choose. Send your own provider key alongside your Iconia key and your provider bills you directly, at the lower amount — your key stays yours and we never substitute ours for it. Send only your Iconia key and our provider account covers the call instead. Either way Iconia is priced as a flat subscription, not a cut of your inference, so we have no incentive to send more tokens than your question needs.

Does a smaller payload mean worse answers?

It did not. Correctness rose from 13% to 100% on project-specific questions, and held at 9 out of 9 at every project size we tested. Removing what a request does not need makes the part it does need easier to use, not harder.

What happens as I store more? Does it get slower or less accurate?

Neither, in testing from 8 to 1,000 memories. The right memory ranked first at every size, and the time added stayed under ten milliseconds and flat. Growing stores are where most memory products degrade, so we measured that one hardest.

Do I have to change how I work?

You point your existing setup at Iconia and keep your tools, your models and your code. There is no library to adopt and nothing to migrate. If you turn Iconia off, you are exactly where you started.

Is my project's knowledge used to train anything?

No. What you store is yours, readable by your account, and exportable whenever you want it. Nothing is harvested from your traffic — what gets remembered is what you choose to store.

Stop paying twice for what you already know.

Get your key