Kinesis
Pricing

Simple pricing for every way you use the Grid

Pay for the compute you actually use when you build on the Grid. Pay one flat annual fee when you bring your own. No idle capacity on your invoice either way.

Traditional cloud

$0.00

True-Util™

$0.00

Idle Capacity

$0.00

Available to monetize on the Kinesis grid

00:0008:0016:0024:00
Reserved rate (traditional)True-Util™ (actual usage)Idle capacity · monetizable
The True-Util™ Model

One pricing model. Two ways to buy.

True-Util™ works the same way regardless of how much computing power you need. Pick the buying mode that matches your workload. The metering, caps, and telemetry are identical.

TRUE-UTIL™ SERVERLESS

Metered usage, capped at the Dedicated rate

Serverless compute on the Kinesis grid. You pay for the compute, memory, storage, and bandwidth your workloads actually use, and never more than you would pay for the equivalent Dedicated machine. Spiky, variable, or hard-to-forecast workloads save the most.

  • Best for inference, dev/staging, agencies, MVPs
  • Runs across a range of CPU and GPU classes
  • No upfront commitments
TRUE-UTIL™ DEDICATED

The whole machine, billed by the hour

Single-tenant compute that is yours for as long as you run it. Full control over the box, predictable billing, and the same Kinesis orchestration and telemetry as Serverless. For workloads where steady utilization is a given.

  • Best for steady training, production HPC, regulated workloads
  • Choice across providers, which reduces lock-in
  • Full control over configuration, performance, privacy
Pricing

Rates

Rates effective July 2026. GPU rates are per GPU per hour.

A100$1.35

Per GPU / Per Hour

1x GPU, 28 CPUs, 120GB RAM, 750GB Storage

Available in 1x, 2x, 4x GPU configurations.

H100$2.50

Per GPU / Per Hour

1x GPU, 28 CPUs, 180GB RAM, 750GB Storage

Available in 1x, 2x, 4x GPU configurations.

A100 NVLink$1.50

Per GPU / Per Hour

8x GPU node: 252 CPUs, 1920GB RAM, 6500GB Storage

Priced per GPU, sold as 8-GPU NVLink nodes.

H100 NVLink$2.75

Per GPU / Per Hour

8x GPU node: 252 CPUs, 1440GB RAM, 6500GB Storage

Priced per GPU, sold as 8-GPU NVLink nodes.

H200 SXM$4.25

Per GPU / Per Hour

8x GPU node: 176 CPUs, 1800GB RAM, 48000GB Storage

Priced per GPU, sold as 8-GPU NVLink nodes.

B200 SXM$6.50

Per GPU / Per Hour

8x GPU node: 252 CPUs, 2048GB RAM, 40000GB Storage

Priced per GPU, sold as 8-GPU NVLink nodes.

B300 SXMfrom $8.13

Per GPU / Per Hour

8x GPU node: 252 CPUs, 2048GB+ RAM, 40000GB+ Storage

Spot capacity. Price varies with market conditions.

Check availability
Compute Optimized CPU$0.035

Per vCPU / Per Hour

1 vCPU, 2GB RAM, 50GB NVMe

Available as serverless True-Util™. Available as 2x, 4x, 8x, 16x, 32x and 64x configurations.

General Purpose CPU$0.045

Per vCPU / Per Hour

1 vCPU, 4GB RAM, 50GB NVMe

Available as serverless True-Util™. Available as 2x, 4x, 8x, 16x, 32x and 64x configurations.

Memory Optimized CPU$0.055

Per vCPU / Per Hour

1 vCPU, 8GB RAM, 50GB NVMe

Available as serverless True-Util™. Available as 2x, 4x, 8x, 16x, 32x and 64x configurations.

Prices shown are for representative configurations. Actual specifications may vary by provider. GPU rates are per GPU per hour. NVLink and SXM systems are sold as 8-GPU nodes. Rates marked “from” are spot rates and vary with market conditions.

Grid Your Compute

Pricing that follows your hardware, not our meter

You already own the compute. Grid Your Compute connects your fleet to the Kinesis control plane, and the Grid places paying workloads and your own workloads across it, recovering the idle hours you are currently writing off.

Because the value we deliver scales with the hardware you bring, so does the price. The fee is a simple percentage of your hardware's market system cost per year. No per-hour meter on your own machines, and no charge that grows with how well the Grid does its job.

3% per year

One flat annual fee per GPU, set as a percentage of the current market cost of the system class you bring. A fleet of eight H100 systems pays the same whether the Grid places workloads on it around the clock or holds capacity for your own bursts.

Longer terms available

Multi-year commitments carry meaningful discounts, including partial-upfront structures that reduce the annual fee. Talk to us about term pricing for your fleet.

What's included

Placement, health monitoring, recovery, and the same telemetry and observability the public Grid runs on. Fleets not directly connected to the internet route through the Kinesis proxy layer. A generous data allowance is included, with metered transfer beyond it.

Volume and support

Larger fleets qualify for volume pricing, and dedicated support arrangements are available. Contact us and we will scope it with you.

Learn about Grid Your Compute

Where True-Util™ saves the most

The same workloads that cost the most on traditional clouds save the most on True-Util™.

AI startups & LLM inference
The pain

H100s sit idle between prompts. The bill is the same whether you served 100 requests or 100,000.

The Kinesis win

True-Util™ Shared meters inference time only. No queries, no cost. Bursty traffic caps at the Reserved rate.

Early-stage SaaS & MVPs
The pain

Overprovisioning for traffic that hasn’t shown up. Or worse — under-provisioning and falling over the first time it does.

The Kinesis win

Pay pennies at low traffic. Costs cap at the Reserved rate during spikes. Headroom without prepayment.

Dev, Staging & CI/CD
The pain

Staging servers run 24/7 to be ready, but burn nights and weekends.

The Kinesis win

True-Util™ drops the bill as activity drops. Same reservation, lower cost when the team’s asleep.

Enterprises with idle capacity
The pain

Reserved AWS instances, on-prem servers, donated lab GPUs — capacity already paid for, sitting underused.

The Kinesis win

Run the Kinesis grid on your hardware at 20% of Shared. Same orchestration, FinOps visibility, 80% less spend on what you already own.

Try it on a real app

$100 in free credit. No credit card required. Deploy your first container in under five minutes. Bring a GitHub repo, a Dockerfile, or just describe what you want.