GPU cloud pricing
Hourly rates for renting GPUs across serverless and dedicated providers. Every number is taken from the vendor's public pricing page — no averages, no estimates. Per-second and per-minute rates are converted to hourly equivalents.
| GPU ↕ | VRAM | RunPod Pods ↕ | RunPod Serverless ↕ | Modal ↕ | Fal.ai ↕ | Baseten ↕ | Replicate ↕ | Lambda Cloud ↕ | CoreWeave Inference ↕ | Nebius AI Cloud ↕ | Hyperstack ↕ | Novita AI ↕ | Verda (formerly DataCrunch) ↕ | Salad Container Engine ↕ |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| T4 | 16 GB | — | — | $0.59/hr* | — | $0.63/hr* | $0.81/hr* | — | — | — | — | — | — | — |
| L4 | 24 GB | $0.39/hr | $0.69/hr* | $0.80/hr* | — | $0.85/hr* | — | — | — | — | — | — | — | — |
| A10 / A10G | 24 GB | — | — | $1.10/hr* | — | $1.21/hr* | — | $1.29/hr* | — | — | — | — | — | — |
| RTX 4090 | 24 GB | $0.69/hr | $1.10/hr | — | — | — | — | — | — | — | — | $0.33/hr* | — | $0.20/hr* |
| RTX 5090 | 32 GB | $0.99/hr | $1.58/hr | — | — | — | — | — | — | — | — | $0.72/hr* | — | $0.29/hr* |
| RTX A6000 / A40 | 48 GB | $0.44/hr* | $1.22/hr* | — | — | — | — | $1.09/hr* | — | — | $0.50/hr* | — | $0.61/hr* | — |
| L40S | 48 GB | $0.99/hr | $1.75/hr* | $1.95/hr* | — | — | $3.51/hr* | — | $2.25/hr* | $1.55/hr* | — | $0.55/hr* | $1.37/hr* | — |
| RTX PRO 6000 | 96 GB | $1.99/hr | $3.49/hr | $3.03/hr* | $2.99/hr* | — | — | — | $2.50/hr* | $1.80/hr* | $1.85/hr* | — | $1.89/hr* | — |
| A100 40 GB | 40 GB | — | — | $2.10/hr* | — | — | — | $1.99/hr* | — | — | — | — | $1.29/hr* | — |
| A100 80 GB | 80 GB | $1.39/hr* | $2.72/hr | $2.50/hr* | — | $4.00/hr* | $5.04/hr* | — | $2.70/hr* | — | $1.35/hr* | $1.60/hr* | $1.79/hr* | — |
| H100 80 GB | 80 GB | $2.89/hr* | $4.55/hr | $3.95/hr* | $4.50/hr* | $6.50/hr* | $5.49/hr* | $3.29/hr* | $6.16/hr* | $3.85/hr* | $2.50/hr* | $3.39/hr* | $3.25/hr* | — |
| H200 | 141 GB | $4.39/hr | $5.93/hr | $4.54/hr* | $4.50/hr* | — | — | — | $6.31/hr* | $4.50/hr* | $3.99/hr* | — | $4.00/hr* | — |
| B200 | 180 GB | $5.89/hr | $8.64/hr | $6.25/hr* | $6.25/hr* | $9.98/hr* | — | $6.99/hr* | $8.60/hr* | $7.15/hr* | $6.00/hr* | — | $6.11/hr* | — |
| B300 | 288 GB | $7.39/hr | $9.98/hr | $7.10/hr* | $8.50/hr* | — | — | — | — | $7.85/hr* | — | — | $7.50/hr* | — |
Green = cheapest listed rate for that GPU. * = hover for the vendor's raw per-second/per-minute rate or a caveat. — = the vendor does not list that GPU publicly.
Sign up (affiliate, same price — some include credit): RunPod +$5 credit · Vultr +$300 trial · DigitalOcean GPU — header links above stay direct, always.
// billing models differ — read this before comparing
- RunPod Pods — per-second, always-on container (Secure Cloud)
- RunPod Serverless — per-second, scale-to-zero workers
- Modal — per-second, scale-to-zero ($30/mo free credit)
- Fal.ai — per-hour list price for custom deployments; committed-use discounts advertised
- Baseten — per-minute, dedicated deployments, scale-to-zero
- Replicate — per-second, private model deployments
- Lambda Cloud — per-hour on-demand GPU instances; prices vary by instance GPU count
- CoreWeave Inference — per-hour single-GPU inference rate; inference platform customers only (contact account executive)
- Nebius AI Cloud — per-second VM compute, shown as on-demand GPU-hour; H100/H200/B200/B300/RTX PRO rates include prescribed vCPU and RAM, while L40S is the minimum all-in 1-GPU preset
- Hyperstack — per-minute on-demand GPU VM, shown per GPU-hour; fixed CPU, RAM, root disk, and ephemeral disk are included, while public IPs and shared storage are separate
- Novita AI — per-second on-demand GPU instance, settled hourly; listed one-GPU compute prices include the product's fixed vCPU and RAM, with storage above the free container-disk quota billed separately
- Verda (formerly DataCrunch) — prepaid pay-as-you-go GPU instances in 10-minute increments, with unused terminated time refunded in the next billing period; listed one-GPU price includes fixed CPU and RAM, while storage is separate
- Salad Container Engine — per-second managed container instances at Lowest (formerly Batch) priority; GPU rates include selected vCPU and RAM, allocation and image-download time are unbilled, and distributed nodes are interruptible
// methodology
Prices are read from each vendor's public pricing page and converted to hourly equivalents (per-second × 3600, per-minute × 60), rounded to cents. "As low as" and committed-use rates are noted but never used in the comparison cells. We verify the whole table twice a week and stamp the date above; the raw JSON with source URLs is downloadable. If a number is wrong, the fix is one verification away — tell us.
Want cost per workload instead of cost per hour? Use the LLM hosting cost calculator.
// signing up? these links support HostFleet
The source links in the table always go to official pricing pages, never through us. The signup links below are affiliate links — clearly labeled, same price for you, and some include free credit: