GPU rental · on request
Systems with four NVIDIA RTX 5090 cards — 32 GB GDDR7 each — joining the AxForge fleet as dedicated EU-hosted machines. On the way now; reserve a system and an engineer confirms timing with you.
Specifications
| System | 4× NVIDIA RTX 5090 in one dedicated machine |
|---|---|
| Memory | 32 GB GDDR7 per card |
| Availability | On request — reservable now, not rentable today |
| Tenancy | Dedicated machine, your traffic only |
| Pricing | €0.54 / GPU-hr · billed monthly (launch pricing) |
Four cards in one box: run one model per card, or spread a workload across the system. An engineer maps your models onto the machine with you at reservation.
Community data
On the same-workload llama.cpp CUDA scoreboard the RTX 5090 decodes at 290.0 t/s with 14,073 t/s prompt processing — the fastest single-stream result on the board, edging even the 96 GB RTX 6000 Pro (274.2 t/s) when the model fits in 32 GB. Roughly 3.8× an RTX 3060 — not 10×: the expensive cards earn their keep on capacity and batch throughput, not just one user's tokens.
Same-workload numbers from the llama.cpp CUDA scoreboard (Llama 2 7B Q4_0, tg128, full GPU offload) — community measurements, not ours. We have not measured this card on our own nodes yet; when it joins the fleet, we publish our own numbers. Full ladder on the GPU catalogue.
Community data
Measured by Hardware Corner with llama.cpp at 4k context, one RTX 5090:
| Model | Quant | Decode @ 4k |
|---|---|---|
| Qwen3 30B-A3B MoE | Q4 | 226.1 t/s |
| Qwen3 8B | Q4 | 200.4 t/s |
| Gemma 4 26B | Q4 | 180.3 t/s |
| Qwen3.5 35B MoE | MXFP4 | 165.2 t/s |
| Qwen3 14B | Q4 | 123.8 t/s |
| Qwen3 32B dense | Q4 | 61.4 t/s |
| Qwen3.5 27B dense | Q4 | 58.8 t/s |
MoE vs dense — again
226 t/s for a 30B MoE vs 61 t/s for a 32B dense: same lesson as GB10 — active parameters and memory traffic beat headline parameter count.
Software moves the numbers
A newer llama.cpp run of the same 30B MoE reaches 365.7 t/s on a short tg128 profile — a different benchmark shape, so we don't mix it into the table, but it shows how far the serving stack moves the result.
Numbers above are single-card, single-stream — rent one card or a multi-GPU system, as your workload needs. Community measurements, not ours — when the systems join the fleet, we publish our own numbers.
Data & privacy
Like every AxForge machine: your model, your traffic, our hardware — prompts never persisted. Only request metadata (token counts, timestamps, status) is kept for billing and operations. Full policy at axforge.ai/privacy.
FAQ
Not yet — the RTX 5090 systems are on the way to the fleet. You can reserve one now; an engineer confirms timing with you before you commit.
Four NVIDIA RTX 5090 cards with 32 GB GDDR7 each, rented as one dedicated machine serving only your traffic.
€0.54 per GPU-hour, billed monthly for the dedicated 4-GPU system (launch pricing). We watch the market and price under it: that is 90% of the lowest comparable on-demand listing we found on 2026-08-26.
In an AxForge EU region. Stockholm and Málaga are live today, with more EU regions in deployment; placement is confirmed with your reservation.
No. Your model, your traffic, our hardware — prompts never persisted. Only request metadata is kept for billing and operations — see the privacy policy.