LLM bills tend to grow in a predictable way. Small, repetitive decisions, such as which team a ticket belongs to, whether a request is in scope, or whether an answer passed a check, get sent to the most capable model available, and each one pays for generated tokens to return a single value. At low volume nobody notices. At high volume those decisions can account for a meaningful share of the bill.
Jev, TypeSafe AI's decision model, takes that work off the LLM. It answers closed questions about text with a choice, a score, or a yes/no probability, and it never writes prose, so there are no generated tokens to pay for. The saving is real but narrower than the headlines suggest, because it applies only to the decisions in a pipeline and the routing around them.
This guide covers where most LLM budgets leak and how Jev reduces LLM costs.
Key takeaways
- Only decisions get cheaper: Jev affects the share of spend that is closed decisions and the routing around it, not calls that write.
- Four ways: Replace closed-decision calls, route requests by difficulty, trim what executors read, and replace LLM judges.
- One step is not a bill: The 96% figure covers a classifier, and LiteLLM's own end-to-end routing results run from 27% to 51%.
- Against a cheap model the gap is small: The savings are large against flagship models and modest in dollars against cheap tiers.
- Test before moving traffic: Run candidates on a labeled sample and compare total spend afterward, since cheaper decisions tend to get made more often.
Where most LLM budgets bleed money
LLM bills rarely grow because of one expensive feature. They grow because small, repetitive work gets sent to the most capable model available. Four leaks show up most often:

- Closed decisions sent to a frontier model: Classifying a ticket, tagging a message, or answering a yes/no check needs one value back. A general model produces it by generating tokens, and often reasoning tokens first, and output tokens are typically priced at several times the rate of input tokens. The expensive part of the call goes on a one-word answer.
- Every request sent to the same model: When one model handles everything, a greeting and a hard multi-step analysis are billed at the same rate per token. Much of the traffic in many products is simple, so much of the capability being paid for goes unused.
- The schema tax: An agent with several tools sends the definition of every tool with every turn, so input tokens are spent reading tools the request will never use. Turns that need no tool at all pay the same toll, and the more tools an agent registers, the larger that fixed cost per turn.
- Reviewing by sample: When every quality check is an LLM call, checking every output adds a second paid call to each response, so teams review a small sample instead.
How much each leak costs depends on a team's traffic.
What Jev is and why its classification-only design matters
Jev is TypeSafe AI's System One model. It answers closed questions about text with a choice from a list, a score on a scale, or a yes/no probability, and it never writes prose. It is hosted, closed-weight, and text only.
A general model acts like an overqualified general contractor who reads the whole blueprint just to say which subcontractor to call, and Jev is built to be the one-line answer.
That design is what changes the cost:
- No generated tokens: Pricing reflects it, at an early-access $0.042 per million input tokens with output free.
- A single pass: Jev scores every option in one pass, and TypeSafe reports response times of 70 to 500 milliseconds.
- A closed list: It can't return a label outside the list or malformed output, so there are no retries for broken JSON.
- Probabilities on every answer: Only the uncertain cases need to reach a pricier model or a person.
How Jev reduces LLM costs
Jev reduces LLM costs in four ways, each tied to one of the leaks above. How large each saving is depends on how much of a pipeline's spend sits in that leak.

1. It takes closed decisions out of the LLM bill
The most direct saving is replacing LLM calls whose only job is to return a label, a score, or a yes/no.
Those calls pay for generated tokens, and often reasoning tokens, to produce one value. A decision model returns the value and charges for input only.
The per-model math for a million decisions is in Jev vs ChatGPT, and a different hypothetical shows how much the result depends on what is being replaced.
Take 2 million requests a month, with 40% of them closed decisions, so 800,000 calls of about 500 input tokens and 20 output tokens each, at the GPT-6 list prices used in that article:
- On a flagship-tier model (GPT-6 Sol): About $960 a month, against about $17 on Jev, a saving of roughly $943.
- On a cheap-tier model (GPT-6 Luna): About $48 a month, against about $17 on Jev, a saving of about $31.
These figures exclude reasoning tokens, review costs, and mistakes. The saving is only as large as the share of spend that is decisions. Calls that write stay exactly where they are.
2. It routes each request to the cheapest model that can handle it
Many requests don't need the flagship model. Model routing uses a small classifier to judge how hard each request is, then sends it to a matching tier, and Jev can be that classifier, returning easy, medium, or hard in one fast call. LiteLLM's Auto Router lists Jev as a classifier option. In LiteLLM's benchmark, Jev was 5.43 times faster than Haiku at median and 96.12% lower in registry-priced classifier cost, though LiteLLM makes the router and notes that cost savings stop at the classifier boundary.
The bigger saving comes from the routing itself. LiteLLM reports 51% saved across 272,876 requests in one production deployment against an all-flagship baseline, and those results aren't specific to Jev. Jev's part is making the routing decision cheap and fast. Requests it's unsure about can fall back to a stronger tier, as covered in running Jev and an LLM in one workflow.
3. It stops the executor from reading everything
A tool-calling agent sends the definition of every registered tool with every turn, so the model pays input tokens to read tools it won't use, and does the same on turns that need no tool at all. A decision model in front changes that: Jev picks the tool, or none, and the executor sees only that tool's definition.
4. It replaces LLM judges on evals and review
Checking every output is expensive when each check is an LLM call, so teams end up reviewing only a small sample. A decision model can run the same checks on every output, since each check is a label, a score, or a yes/no returned in one fast call, and several checks can share a single request for only marginally more.
How much this saves depends on which judge it replaces. In LangChain's judging test, the gap was large against flagship LLM judges and small against the cheapest one, where the case rests more on speed and consistency than on dollars. The test is LangChain's own, and LangChain is an integration partner.
Auditing your LLM spend in one place
To use the four ways above, you first need to know which of your LLM calls they apply to and what those calls cost. Four steps find that out:
- Group spend by endpoint or feature: Most gateways and dashboards can break cost down by tag, which often points to the right calls without reading a prompt.
- Mark the calls that fit: Flag the ones that match any of the four kinds above.
- Size the saving on a sample: Multiply those calls by the price difference between what they cost now and the cheaper option.
- Test before moving traffic: Run them on a few hundred labeled cases, send uncertain ones to the current model or a person, and compare total spend afterward, since cheaper decisions tend to get made more often.
Grouping, marking, and sizing are analysis over data a team already has, and that is where PromptQL can help. It connects to warehouses, databases, and SaaS apps, so if call logs and billing exports sit in a connected source, you can ask in plain English which endpoints drive most of the spend.
The numbers are computed by code in a secure sandbox against the real data, and in workspaces where Jev has been provisioned, PromptQL can run the sorting of logged prompts, as laid out in using Jev without writing code. Permissions are enforced at the data layer using the access of the person asking.
Conclusion
Cutting an LLM bill starts with knowing which calls are decisions, because only those, plus the routing around them, can move to something cheaper. The headline figures measure single steps, so the real saving is smaller than they suggest, and it depends on how much of a bill is decisions and which model runs them today. Teams that size it first, test it on their own data, and keep a stronger path for the uncertain cases are the ones that keep the savings they measure.
Frequently Asked Questions
How much can Jev actually save on LLM costs?
It depends on how much of a bill is closed decisions and which model handles them today. Against a flagship model the decision slice can cost a small fraction of what it does now, about $960 against about $17 in this article's hypothetical, while against a cheap tier the dollar gap is small. Figures like 96% describe a single step, not a whole bill.
Does Jev replace an LLM?
No. Jev returns choices, scores, and probabilities, not text, so writing, summarizing, and coding stay with an LLM. It handles the decisions around those calls.
Is Jev's pricing stable?
Not yet guaranteed. The $0.042 per million input tokens price is early-access pricing, and rate limits can change without notice, according to a third-party summary of TypeSafe's documentation, so build a fallback to an LLM and re-check pricing before moving a large share of traffic.
Does cheaper mean less accurate?
Not necessarily, but it can be on harder tasks. An arXiv audit of the first nine days after release found accuracy gaps on harder tasks, so test on labeled cases and weigh any accuracy loss against the saving, not unit price alone.
Can Jev plug into an existing LLM gateway?
In some setups, yes. LiteLLM's Auto Router lists Jev as a classifier option, and one author's test reached Jev through OpenRouter's Decisions API. Check the gateway's current documentation for setup details.