← back to blog
CostsMeasured·September 28, 2026·5 min read

LLM cost per request: what 204 production calls actually cost

We logged every model call our product made for 90 days. The useful finding was not the total. It was that two operations in the same product have opposite cost shapes, so the advice that saves money on one wastes it on the other.

Maxfounder of Sndwich · @SattyEth

Pricing pages are honest and almost useless. Anthropic will tell you that Sonnet is $3 per million input tokens and $15 per million output. True, and it tells you nothing about what a feature costs, because it does not know how many tokens your feature spends or which side of the ledger they land on.

So we read our own log instead. Every model call our product makes is written to a table at the moment it happens, with the task name, the model the router picked, the input and output token counts, and the cost computed at log time. This is what 204 real calls over 90 days look like.

LLM cost per request is not one number

The total is small and boring: $1.21 across 204 calls between June 25 and September 23, 2026. Divide it out and you get about half a cent per call, which is the sort of average that hides everything worth knowing.

Here is the same data split by what the call was actually doing. The bar is the share of the cost that went to output, the text the model wrote, as opposed to input, the text we sent it.

Write a post in a voice432 in / 581 out
87%
Clone a writing style691 in / 758 out
85%
Rewrite for another network303 in / 310 out
84%
Suggest a topic280 in / 117 out
68%
Reply to someone's post963 in / 27 out
12%
Share of the cost that is output, by operation. Same product, same 90 days. The reply is the odd one out and it is the one that runs most often.
Writing a post costs $0.0100 and is 87% output. Answering someone else's post costs $0.0011 and is 88% input. The money is on opposite sides of the same product.

Why the shape flips

It is obvious once you see the token counts, and invisible until you do. Writing a post means sending a short brief and getting back a long piece of text: 432 tokens in, 581 out. The model is doing the expensive thing, and output is priced five times higher than input.

Answering someone means the opposite. To write one good sentence under a stranger's post you have to send the post, the thread, the voice you are answering in, and the rules about what not to say. That came to 963 tokens in for 27 tokens out, a ratio of about 35 to 1. You are paying to read, not to write.

Output-heavy workPosts, rewrites, style clones. Cost scales with how much text you ask for, so the lever is the number of variants you generate.
Input-heavy workReplies, classification, moderation. Cost scales with the context you attach, so the lever is what you stop sending.
The cheap model is not always cheaperReplies run on Haiku at $1/$5. That is why they cost a tenth of a post despite sending twice the tokens.

What this changes about cutting LLM cost per request

The standard advice is to batch: put several jobs in one call and save on the repeated prompt. Our earlier measurement on the post path says that batching six style variants into one call saves only the input, which is 13% of that operation. The six posts get written either way. Batching made the call count smaller and the bill nearly the same.

On the reply path the same advice is right, for the opposite reason. There, the prompt is the bill. Trimming the thread context, dropping the parts of the voice description the model does not need for a one-line answer, and reusing a cached prefix all hit the 88% directly.

Measure by operation, not by model.A per-token price cannot tell you where your money is. Log task, model, input and output on every call, and the answer falls out in an afternoon.
On output-heavy work, cut variants.Generate the one thing the user asked for. Generating six and showing one is paying five times for nothing.
On input-heavy work, cut context.Every token of instruction you send with a one-line reply is charged on every reply, forever.

The honest limits of this number

This is one product's log, not a benchmark. 204 calls is enough to see the shape and not enough to argue about the third decimal. The reply path has 58 calls behind it and the style-clone path has nine. The router picks the model per task, so a task's cost is entangled with that choice rather than being a property of the work.

One more thing we found while pulling the data, offered as a warning: the same table held 99,999 rows from a quota load test, a single afternoon in June. They were 96% of the recorded spend. Any dashboard reading that table without filtering them would have reported a cost forty times the real one. If you log every call, label your test rows.

The method costs nothing and does not need a vendor. A table with five columns, written at the moment of the call, answers questions that no pricing page can. For how the writing side of this works in practice, see the piece on replies as a growth channel.

This is the meter we run on

Sndwich logs every model call it makes for you, with the task and both token counts. The numbers above came out of that table, not an estimate.

Cost per call, recorded. Task, model, input, output and price, written when the call happens.
One variant by default. A post is written in the style you picked, not in six and thrown away.
Your own key, optionally. Bring an Anthropic key and the calls stop counting against a plan.
Start free →
ShareTelegramX

Read next