Pricing

Simple, per-user pricing

Prices are shown per month. Start free with a few coworkers, add a usage budget when you need more, and move to dedicated infrastructure when your organization does.

Playground

Always free

$0/ user / month

 

A free workspace to build multiplayer bots

Team

Starts at

$40/ paid user / month

Every user starts free. Pay only for users who need more in $40 increments, billed weekly.

Shared context: Knowledge, Skills, Semantic layer

Enterprise

Starts at

$1,000/ month

Consumption based with a monthly minimum commit

No limit on users

Shared context: Knowledge, Skills, Semantic layer

Single-tenant deployment

Custom

Priced to fit

Talk to us

Contract pricing

BYOC or self-hosted deployment

Bring your own models

Dedicated forward-deployed engineers

Your model choice changes your mileage.

OLUs are the normalized unit we use to measure usage across models. Frontier models like Claude Fable 5 and GPT-6 Astra consume the most OLUs; open-weight models like Kimi K3, DeepSeek V4 and GLM 5.2 do the same work for a fraction of the OLUs. Switch to different models for different bots — or even mid-thread.

Model

Type

OLU multiplier (vs Opus 4.6)

Input tokens / OLU

Output tokens / OLU

Cached tokens / OLU

Cache creation tokens / OLU

Notes

Open-weight models · served via Fireworks

DeepSeek V4 Flash 0731
Fireworks
Open-weight

0.025x (40x cheaper)

952,400

595,200

6.0M

1.2M

Served via Fireworks. 07/31 snapshot variant of DeepSeek V4 Flash.
Llama 3.3 70B
Fireworks
Open-weight

0.036x (28x cheaper)

493,800

617,300

6.2M

617,300

Meta; served via Fireworks.
DeepSeek V4.1 Flash
Fireworks
Open-weight

0.04x (25x cheaper)

606,100

252,500

23.8M

757,600

Multimodal DeepSeek model with native vision; served via Fireworks. $0.22/$0.007/$0.66 per 1M in/cached/out.
Qwen 3.7 Plus
Fireworks
Open-weight

0.071x (14x cheaper)

333,300

104,200

4.2M

416,700

Alibaba, vision-capable; served via Fireworks.
Mistral Large 3
Fireworks
Open-weight

0.077x (13x cheaper)

333,300

83,300

4.2M

416,700

Mistral AI; served via Fireworks.
Kimi K2.7
Fireworks
Open-weight

0.17x (5.8x cheaper)

140,400

41,700

1.8M

175,400

Moonshot AI 1T-param MoE; served via Fireworks.
GLM-5.2
Fireworks
Open-weight

0.23x (4.3x cheaper)

95,200

37,900

1.2M

119,000

Z.ai (Zhipu), MIT-licensed; served via Fireworks.
DeepSeek V4 Pro 0813
Fireworks
Open-weight

0.24x (4.2x cheaper)

101,000

42,100

3.6M

126,300

Official GA snapshot (0813) of DeepSeek V4 Pro; served via Fireworks. Supersedes the preview. $1.32/$0.044/$3.96 per 1M in/cached/out.
DeepSeek V4 Pro (Fireworks)
Fireworks
Open-weight

0.25x (4x cheaper)

76,600

47,900

1.1M

95,800

Preview (deepseek-v4-pro), still billed at 0.25x / $1.74/$0.145/$3.48 per 1M in/cached/out; superseded by 0813. Served via Fireworks (what PromptQL uses for open-weight).
Kimi K3
Fireworks
Open-weight

0.6x (1.7x cheaper)

44,400

11,100

555,600

44,400

Moonshot AI flagship; served via Fireworks.

Proprietary models

GPT-5.6 Luna
OpenAI
Proprietary

0.04x (25x cheaper)

133,300

33,300

1.7M

166,700

Cost-efficient GPT-5.6 tier.
Gemini 3.7 Flash
Google
Proprietary

0.15x (6.7x cheaper)

177,800

44,400

2.2M

222,200

Google Gemini Flash tier; ~2x cheaper than Gemini 3.6 Flash.
Gemini 3.8 Flash
Google
Proprietary

0.15x (6.7x cheaper)

177,800

44,400

2.2M

222,200

Latest Google Gemini Flash tier; same introductory pricing as Gemini 3.7 Flash through Dec 31, 2026.
Claude Haiku 4.5
Anthropic
Proprietary

0.2x (5x cheaper)

133,300

33,300

1.7M

133,300

Fast tier.
GPT-5
OpenAI
Proprietary

0.3x (3.4x cheaper)

106,700

16,700

1.3M

133,300

GPT-5 base.
GPT-5.1
OpenAI
Proprietary

0.3x (3.4x cheaper)

106,700

16,700

1.3M

133,300

Same rate as GPT-5.
Grok 4.5
SpaceXAI
Proprietary

0.37x (2.7x cheaper)

66,700

22,200

666,700

SpaceXAI (xAI) frontier coding/agentic model. No cache-creation pricing.
Grok 4.6
SpaceXAI
Proprietary

0.37x (2.7x cheaper)

66,700

22,200

666,700

SpaceXAI (xAI) frontier coding/agentic model. Same pricing as Grok 4.5. No cache-creation pricing.
Gemini 3.1 Pro
Google
Proprietary

0.42x (2.4x cheaper)

66,700

13,900

833,300

83,300

Google flagship Pro.
GPT-5.2
OpenAI
Proprietary

0.42x (2.4x cheaper)

76,200

11,900

952,400

95,200

Flagship step-up.
GPT-5.6 Terra
OpenAI
Proprietary

0.42x (2.4x cheaper)

53,300

11,100

666,700

66,700

Balanced GPT-5.6 tier.
GPT-5.4
OpenAI
Proprietary

0.52x (1.9x cheaper)

53,300

11,100

666,700

66,700

Reasoning/coding/agentic gains.
Claude Sonnet 4.5
Anthropic
Proprietary

0.6x (1.7x cheaper)

44,400

11,100

555,600

44,400

Balanced tier. (Sonnet 4.6 is blocklisted in PromptQL.)
Claude Opus 4.6
Anthropic
Proprietary

1.0x (baseline)

26,700

6,700

333,300

26,700

Anchor (1.0x).
Claude Opus 4.8
Anthropic
Proprietary

1.0x (baseline)

26,700

6,700

333,300

26,700

Current GA Opus; PromptQL Powerful tier. Same rate as 4.6/4.7.
Claude Opus 5
Anthropic
Proprietary

1.0x (baseline)

26,700

6,700

333,300

26,700

Latest GA Opus. Same rate as 4.8/4.6.
GPT-5.5
OpenAI
Proprietary

1.0x (baseline)

26,700

5,600

333,300

33,300

Latest flagship. >272K-context billed 2x in / 1.5x out.
GPT-5.6 Sol
OpenAI
Proprietary

1.0x (baseline)

26,700

5,600

333,300

33,300

GPT-5.6 flagship. GA July 9, 2026.
Claude Fable 5
Anthropic
Proprietary

2x (2x cost)

13,300

3,300

166,700

13,300

New frontier flagship (Jun 2026); most capable, highest cost.
GPT-6 Astra
OpenAI
Proprietary

2x (2x cost)

13,300

3,300

166,700

13,300

OpenAI GPT-6; $10/$50 per MTok in/out. Formula (10 + 0.05*50)/6.25 = 2.00.
GPT-5 Pro
OpenAI
Proprietary

7.22x (7.2x cost)

8,900

1,400

11,100

11,100

Pro reasoning. No prompt-cache discount.
GPT-5.2 Pro
OpenAI
Proprietary

10.11x (10.1x cost)

6,300

992

7,900

7,900

Pro reasoning. No prompt-cache discount.
GPT-5.4 Pro
OpenAI
Proprietary

13.53x (13.5x cost)

4,400

926

5,600

5,600

Pro reasoning. No prompt-cache discount.
GPT-5.5 Pro
OpenAI
Proprietary

13.53x (13.5x cost)

4,400

926

5,600

5,600

Top GPT pro tier. No prompt-cache discount.

Anchor: Claude Opus 4.6 = 1.0×. Lower multiplier = cheaper, i.e. more tokens of work per OLU (per $0.20). PromptQL serves open-weight models via Fireworks (DeepSeek is shown for both its first-party Official API and the Fireworks-hosted rate PromptQL bills on). PromptQL gives you access to every model available, and the lineup expands over time — pick your model per thread. The GPT *-pro tiers run higher mainly because they have no prompt-cache discount. Token figures are representative; your exact mileage depends on the task.

Frequently asked questions

Playground is a free workspace where you build bots and share them, within basic usage limits. On Team, every user starts free with a basic usage limit, and paid usage plans start at $40 per user per month, billed weekly as $10 per user per week. Enterprise is a $1,000 per month minimum commit with no limit on users and a single-tenant deployment. Custom covers BYOC, self-hosted and bring-your-own-model deployments.
A usage budget. Each Team user on a paid plan gets $40 of usage a month, billed as $10 each week, and you can assign a higher plan to anyone who needs it in $40 per month increments. The budget is spent as that user's bots consume OLUs. Invoices are weekly, and unused credits roll over to the next week.
An Operational Language Unit (OLU) is PromptQL's normalized unit for AI usage. It combines input, output and cached tokens, then applies the selected model's multiplier, so usage across different token types and models can be measured in one unit. More complex or longer work generally consumes more OLUs.
On Team, usage is billed at cost with no margin during the introductory period: $10 buys $10 worth of tokens, VM runtime and any other billable usage. Enterprise starts at $1,000 per month, fully allocated to usage, with OLU rates set by contract.
Yes. Your price per OLU is fixed by your plan or contract, while the model multiplier changes how many OLUs the model's token usage consumes. A 0.20× model uses one-fifth the OLUs of a 1.0× model for the same token mix. The model → OLU table above shows the multiplier and representative tokens per OLU for every model.
How much context a bot reads, how much it writes, the mix of input, output and cached tokens, and the model you choose. You control those dials by picking the right model and keeping each bot focused on the context it needs. A running VM also consumes OLUs, as do built-in integrations like Exa, Gemini or XAI.
All model, infrastructure and sandbox-hosting charges. That covers the LLM tokens plus the orchestration and the isolated computer each bot runs on: VM, files, artifacts and schedules. There is no separate LLM pass-through or infrastructure bill.
Yes. Connect your own OpenAI Codex subscription under bring-your-own-model and PromptQL charges $0 for usage on that connected model. Usage on every other model is deducted from your OLU budget as usual.
Yes. On Team, every user starts on the free tier with a basic usage limit, and you can move anyone to a paid plan from $40 a month ($10 a week), in $40 per month increments. When a user reaches their limit, their bots pause until the budget resets or you raise the plan. Enterprise sets per-user weekly limits with no cap on the number of users.
Start on Free: create a project, invite coworkers, connect your data and start a bot. Upgrade to Team when you need a bigger usage budget, or talk to us about Enterprise when you need dedicated infrastructure and organization-wide controls.