Pricing
GPU pricing is dynamic per-hour, shown live in the catalog, metered hourly and billed from your credit balance.
GPU pricing is dynamic. There are no fixed GPU price tables — the catalog shows the current per-hour price for each machine, based on real-time hardware availability.
How you are charged
You pay the per-hour rate shown on the offer at the moment you launch it. Usage is metered hourly and deducted from your credit balance.
| Detail | Value |
|---|---|
| Rate | Per-hour, shown on each offer in the catalog |
| Metering | Hourly, while the instance is running |
| Billed from | credit balance |
Reading live prices
To see what a machine costs right now, open the catalog screen — every offer shows its current per-hour price.
GPU-hours vs. hosted inference
A GPU instance or model server bills per hour it is running, whether or not it is actively handling work — so it pays off for steady, predictable load, and for anything that isn't a simple chat request. Inference instead bills per token, with no instance to leave running, which is usually cheaper for spiky or low-volume use against a model NevTan already hosts. Stop or destroy instances and model servers you're done with — see Launch an instance and Model servers — since the hourly rate keeps accruing until you do.