You pay for the run,
not the cluster.
Compute is billed per GPU-hour at the rate the scheduler actually won it for. No reserved capacity you are not using, no minimum, no seat licence.
| Free | Flow | Reserved | Enterprise | |
|---|---|---|---|---|
| Platform fee | $0 | $0 / mo | $400 / mo | Talk to us |
| Compute | $40 credit | $1.90 / GPU-hr | $1.35 / GPU-hr | Committed rate |
| Model size | Up to 8B | Every base | Every base | Every base + private |
| Concurrent runs | 1 | 2 | 12 | Unlimited |
| Spot recovery & checkpointing | Yes | Yes | Yes | Yes |
| Weight export | Yes | Yes | Yes | Yes |
| Guaranteed capacity window | No | No | 4 h | Contracted |
| Private VPC deployment | No | No | No | Yes |
| Support | Community | Email, 1 working day | Shared Slack | Named engineer |
How the billing works
Why the rate moves
The scheduler bids for interruptible capacity across regions. When the market is quiet you pay less, and we pass that through rather than pocketing the spread. The prices above are ceilings, not averages.
What a real invoice looks like
An 8B LoRA tune on 240M tokens finishes in about four hours on 24 GPUs, so roughly 95 GPU-hours: about $180 on Flow, $130 on Reserved. Ingestion, tokenisation and every evaluation you run afterwards are free.
Leaving
Export the weights and go. There is no egress fee, no licence that lapses, and no request
form: penstock export works on the last day of your account exactly as it
did on the first.