GPU rental · on request

RTX 5090 rental in Europe — 4-GPU dedicated systems

Systems with four NVIDIA RTX 5090 cards — 32 GB GDDR7 each — joining the AxForge fleet as dedicated EU-hosted machines. On the way now; reserve a system and an engineer confirms timing with you.

On request 4× 32 GB GDDR7 €0.54 / GPU-hr · billed monthly

Specifications

What is coming

System4× NVIDIA RTX 5090 in one dedicated machine
Memory32 GB GDDR7 per card
AvailabilityOn request — reservable now, not rentable today
TenancyDedicated machine, your traffic only
Pricing€0.54 / GPU-hr · billed monthly (launch pricing)

Four cards in one box: run one model per card, or spread a workload across the system. An engineer maps your models onto the machine with you at reservation.

Community data

The fastest single-stream card on the board

On the same-workload llama.cpp CUDA scoreboard the RTX 5090 decodes at 290.0 t/s with 14,073 t/s prompt processing — the fastest single-stream result on the board, edging even the 96 GB RTX 6000 Pro (274.2 t/s) when the model fits in 32 GB. Roughly 3.8× an RTX 3060 — not 10×: the expensive cards earn their keep on capacity and batch throughput, not just one user's tokens.

Same-workload numbers from the llama.cpp CUDA scoreboard (Llama 2 7B Q4_0, tg128, full GPU offload) — community measurements, not ours. We have not measured this card on our own nodes yet; when it joins the fleet, we publish our own numbers. Full ladder on the GPU catalogue.

Community data

Modern models at 4k context

Measured by Hardware Corner with llama.cpp at 4k context, one RTX 5090:

ModelQuantDecode @ 4k
Qwen3 30B-A3B MoEQ4226.1 t/s
Qwen3 8BQ4200.4 t/s
Gemma 4 26BQ4180.3 t/s
Qwen3.5 35B MoEMXFP4165.2 t/s
Qwen3 14BQ4123.8 t/s
Qwen3 32B denseQ461.4 t/s
Qwen3.5 27B denseQ458.8 t/s

MoE vs dense — again

226 t/s for a 30B MoE vs 61 t/s for a 32B dense: same lesson as GB10 — active parameters and memory traffic beat headline parameter count.

Software moves the numbers

A newer llama.cpp run of the same 30B MoE reaches 365.7 t/s on a short tg128 profile — a different benchmark shape, so we don't mix it into the table, but it shows how far the serving stack moves the result.

Numbers above are single-card, single-stream — rent one card or a multi-GPU system, as your workload needs. Community measurements, not ours — when the systems join the fleet, we publish our own numbers.

Data & privacy

Dedicated, EU-hosted

Like every AxForge machine: your model, your traffic, our hardware — prompts never persisted. Only request metadata (token counts, timestamps, status) is kept for billing and operations. Full policy at axforge.ai/privacy.

FAQ

RTX 5090 rental — common questions

Can I rent an RTX 5090 server from AxForge today?

Not yet — the RTX 5090 systems are on the way to the fleet. You can reserve one now; an engineer confirms timing with you before you commit.

What is in an AxForge RTX 5090 system?

Four NVIDIA RTX 5090 cards with 32 GB GDDR7 each, rented as one dedicated machine serving only your traffic.

What does RTX 5090 rental cost?

€0.54 per GPU-hour, billed monthly for the dedicated 4-GPU system (launch pricing). We watch the market and price under it: that is 90% of the lowest comparable on-demand listing we found on 2026-08-26.

Where will the RTX 5090 systems be hosted?

In an AxForge EU region. Stockholm and Málaga are live today, with more EU regions in deployment; placement is confirmed with your reservation.

Do you keep prompts on a rented RTX 5090 system?

No. Your model, your traffic, our hardware — prompts never persisted. Only request metadata is kept for billing and operations — see the privacy policy.

Need a machine scoped to your workload?

Talk to an engineer Get an API key

Explore

All GPUs

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms