GPU Instances · Pricing
Docs / GPU Instances

Pricing

GPU pricing is dynamic per-hour, shown live in the catalog, metered hourly and billed from your credit balance.

GPU pricing is dynamic. There are no fixed GPU price tables — the catalog shows the current per-hour price for each machine, based on real-time hardware availability.

How you are charged

You pay the per-hour rate shown on the offer at the moment you launch it. Usage is metered hourly and deducted from your credit balance.

DetailValue
RatePer-hour, shown on each offer in the catalog
MeteringHourly, while the instance is running
Billed fromcredit balance

Reading live prices

To see what a machine costs right now, open the catalog screen — every offer shows its current per-hour price.

Important
Because prices are live, the cost of an equivalent GPU can change between visits to the catalog. Confirm the per-hour price on the offer before you launch.

GPU-hours vs. hosted inference

A GPU instance or model server bills per hour it is running, whether or not it is actively handling work — so it pays off for steady, predictable load, and for anything that isn't a simple chat request. Inference instead bills per token, with no instance to leave running, which is usually cheaper for spiky or low-volume use against a model NevTan already hosts. Stop or destroy instances and model servers you're done with — see Launch an instance and Model servers — since the hourly rate keeps accruing until you do.

Note
See credits for how balances, top-ups, and deductions work.