Enterprise GPU Cloud.
Built in the Netherlands.
Dedicated NVIDIA GPU capacity for AI training and inference, delivered with EU data residency and direct hardware ownership.
The NovaServe Token Factory
The industry is shifting how it measures AI infrastructure. Instead of GPU-hours, the unit that matters is the token: every word generated, every inference step, every agent action. NVIDIA has termed this shift the "AI factory" — compute built to produce tokens at scale, the way a power plant produces electricity.
The NovaServe Token Factory puts a metered, pay-per-token API layer on top of our own GPU fleet. You get inference priced by output, not by idle capacity — running on hardware we own, hosted in the Netherlands.
Metered by the token
Pay per million tokens processed, not per GPU-hour reserved. Costs track usage directly.
Hosted in the Netherlands
Inference runs exclusively on our own Dutch infrastructure, not shared with third-party tenants.
Open-weight models
Serve leading open models out of the box, or bring your own fine-tuned weights.
Dedicated throughput
Reserved-capacity tiers remove multi-tenant queueing for latency-sensitive workloads.
Indicative pricing. Final rates depend on model, context length and committed volume.
Pricing by GPU tier
Hourly on-demand pricing across our current fleet. All GPUs are dedicated, not shared or virtualised.
| Specification | RTX 4090 | RTX 5090 | RTX Pro 6000 | H200 Flagship |
|---|---|---|---|---|
| VRAM | 24 GB GDDR6X | 32 GB GDDR7 | 96 GB GDDR7 ECC | 141 GB HBM3e |
| Memory bandwidth | 1.0 TB/s | 1.79 TB/s | 1.6 TB/s | 4.8 TB/s |
| FP16 performance | ~330 TFLOPS | ~419 TFLOPS | ~503 TFLOPS | ~989 TFLOPS |
| FP8 performance | ~660 TFLOPS | ~838 TFLOPS | ~1,006 TFLOPS | ~1,979 TFLOPS |
| Best for | Rendering, dev & test | Generative media, fine-tuning | Inference at scale | LLM training & large-context inference |
| Indicative price | from €0.65/hron-demand | from €1.05/hron-demand | from €1.85/hron-demand | from €3.40/hron-demand |
Prices shown are indicative on-demand rates per GPU. Reserved and enterprise pricing, with volume discounts, is available on request.
Matched to the workload
We size the hardware to the job, not the other way round.
LLM training & fine-tuning
Large-context model training and fine-tuning workloads that need maximum memory bandwidth.
Inference at scale
High-throughput serving for production inference endpoints with predictable latency.
Rendering & generative media
Image, video and 3D rendering pipelines that benefit from strong single-GPU throughput.
Research & experimentation
Short-lived, flexible capacity for prototyping models and testing new architectures.
Infrastructure you can audit
Built for organisations that need to know exactly where their data and compute sit.
Data stays in the EU
All compute and storage remain within the Netherlands and the EU, with no dependency on non-EU cloud infrastructure.
Hardware we own and run
We operate our own GPU fleet. No reseller layer, no capacity shared with third-party tenants.
White-glove support
A dedicated technical team for onboarding, scaling and incident response, from first call to production.
Talk to our team
Tell us about your workload and we'll come back with capacity and pricing options within one business day.