Pricing pages are honest and almost useless. Anthropic will tell you that Sonnet is $3 per million input tokens and $15 per million output. True, and it tells you nothing about what a feature costs, because it does not know how many tokens your feature spends or which side of the ledger they land on.
So we read our own log instead. Every model call our product makes is written to a table at the moment it happens, with the task name, the model the router picked, the input and output token counts, and the cost computed at log time. This is what 204 real calls over 90 days look like.
LLM cost per request is not one number
The total is small and boring: $1.21 across 204 calls between June 25 and September 23, 2026. Divide it out and you get about half a cent per call, which is the sort of average that hides everything worth knowing.
Here is the same data split by what the call was actually doing. The bar is the share of the cost that went to output, the text the model wrote, as opposed to input, the text we sent it.
Why the shape flips
It is obvious once you see the token counts, and invisible until you do. Writing a post means sending a short brief and getting back a long piece of text: 432 tokens in, 581 out. The model is doing the expensive thing, and output is priced five times higher than input.
Answering someone means the opposite. To write one good sentence under a stranger's post you have to send the post, the thread, the voice you are answering in, and the rules about what not to say. That came to 963 tokens in for 27 tokens out, a ratio of about 35 to 1. You are paying to read, not to write.
What this changes about cutting LLM cost per request
The standard advice is to batch: put several jobs in one call and save on the repeated prompt. Our earlier measurement on the post path says that batching six style variants into one call saves only the input, which is 13% of that operation. The six posts get written either way. Batching made the call count smaller and the bill nearly the same.
On the reply path the same advice is right, for the opposite reason. There, the prompt is the bill. Trimming the thread context, dropping the parts of the voice description the model does not need for a one-line answer, and reusing a cached prefix all hit the 88% directly.
The honest limits of this number
This is one product's log, not a benchmark. 204 calls is enough to see the shape and not enough to argue about the third decimal. The reply path has 58 calls behind it and the style-clone path has nine. The router picks the model per task, so a task's cost is entangled with that choice rather than being a property of the work.
One more thing we found while pulling the data, offered as a warning: the same table held 99,999 rows from a quota load test, a single afternoon in June. They were 96% of the recorded spend. Any dashboard reading that table without filtering them would have reported a cost forty times the real one. If you log every call, label your test rows.
The method costs nothing and does not need a vendor. A table with five columns, written at the moment of the call, answers questions that no pricing page can. For how the writing side of this works in practice, see the piece on replies as a growth channel.
This is the meter we run on
Sndwich logs every model call it makes for you, with the task and both token counts. The numbers above came out of that table, not an estimate.
11 minHow to create an AI model for Instagram and TikTok: face swap video, voice clone, first clip step by step
9 minInstagram action blocks: why they land and what to do
9 minHow to grow on X in 2026: replies get reach, posts do not