GPU Capacity.

Bare-metal NVIDIA GPU capacity billed hourly, for training, fine-tuning, inference, model serving, and BYOM workloads.

GPU Capacity console: choosing a GPU compute plan, billed hourly

Capacity

Dedicated servers, not spot.

Dedicated GPU servers allocated to the workload. Not spot capacity. Run your own models, containers, serving stack, or training jobs on Origon infrastructure.

NVIDIA HGX H100

GPU Count
8
GPU Ram
640 GB
CPU (x2)
Intel® Xeon® Platinum 8468
System RAM
2,048 GiB
NVMe Storage
61.44 TiB

NVIDIA HGX H200

GPU Count
8
GPU Ram
1,128 GB
CPU (x2)
Intel® Xeon® Platinum 8592+
System RAM
2,048 GiB
NVMe Storage
61.44 TiB

AMD MI300X

GPU Count
8
GPU Ram
1,536 GB
CPU (x2)
Intel® Xeon® Platinum 8568Y
System RAM
2,048 GiB
NVMe Storage
61.44 TiB

Workloads

The full training-to-serving path.

Training

Full training runs on dedicated GPUs.

Fine-tuning

Adapt base models to your own data.

Inference

Serve model predictions at scale.

Model serving

Host your own model endpoints.

Customer models

Run the models you bring.

BYOM workflows

Your stack on Origon capacity.

Service Pairings

Compose with the rest of AI Cloud.

Use GPU Capacity with AI Datastore, Speech, and Voice Network when the workload needs datastore, real-time speech, or telephony.

GPU Cluster composing with AI Datastore, Speech Engine, and Voice Network

Reserve GPU capacity.

Tell us your workload and timeline — we provision dedicated GPU servers matched to your training and inference needs.

Request Access

© 2026 Origon Inc.