Pay for the compute you actually use when you build on the Grid. Pay one flat annual fee when you bring your own. No idle capacity on your invoice either way.
$0.00
$0.00
$0.00
Available to monetize on the Kinesis grid
True-Util™ works the same way regardless of how much computing power you need. Pick the buying mode that matches your workload. The metering, caps, and telemetry are identical.
Serverless compute on the Kinesis grid. You pay for the compute, memory, storage, and bandwidth your workloads actually use, and never more than you would pay for the equivalent Dedicated machine. Spiky, variable, or hard-to-forecast workloads save the most.
Single-tenant compute that is yours for as long as you run it. Full control over the box, predictable billing, and the same Kinesis orchestration and telemetry as Serverless. For workloads where steady utilization is a given.
Rates effective July 2026. GPU rates are per GPU per hour.
Per GPU / Per Hour
1x GPU, 28 CPUs, 120GB RAM, 750GB Storage
Available in 1x, 2x, 4x GPU configurations.
Per GPU / Per Hour
1x GPU, 28 CPUs, 180GB RAM, 750GB Storage
Available in 1x, 2x, 4x GPU configurations.
Per GPU / Per Hour
8x GPU node: 252 CPUs, 1920GB RAM, 6500GB Storage
Priced per GPU, sold as 8-GPU NVLink nodes.
Per GPU / Per Hour
8x GPU node: 252 CPUs, 1440GB RAM, 6500GB Storage
Priced per GPU, sold as 8-GPU NVLink nodes.
Per GPU / Per Hour
8x GPU node: 176 CPUs, 1800GB RAM, 48000GB Storage
Priced per GPU, sold as 8-GPU NVLink nodes.
Per GPU / Per Hour
8x GPU node: 252 CPUs, 2048GB RAM, 40000GB Storage
Priced per GPU, sold as 8-GPU NVLink nodes.
Per GPU / Per Hour
8x GPU node: 252 CPUs, 2048GB+ RAM, 40000GB+ Storage
Spot capacity. Price varies with market conditions.
Check availabilityPer vCPU / Per Hour
1 vCPU, 2GB RAM, 50GB NVMe
Available as serverless True-Util™. Available as 2x, 4x, 8x, 16x, 32x and 64x configurations.
Per vCPU / Per Hour
1 vCPU, 4GB RAM, 50GB NVMe
Available as serverless True-Util™. Available as 2x, 4x, 8x, 16x, 32x and 64x configurations.
Per vCPU / Per Hour
1 vCPU, 8GB RAM, 50GB NVMe
Available as serverless True-Util™. Available as 2x, 4x, 8x, 16x, 32x and 64x configurations.
Prices shown are for representative configurations. Actual specifications may vary by provider. GPU rates are per GPU per hour. NVLink and SXM systems are sold as 8-GPU nodes. Rates marked “from” are spot rates and vary with market conditions.
You already own the compute. Grid Your Compute connects your fleet to the Kinesis control plane, and the Grid places paying workloads and your own workloads across it, recovering the idle hours you are currently writing off.
Because the value we deliver scales with the hardware you bring, so does the price. The fee is a simple percentage of your hardware's market system cost per year. No per-hour meter on your own machines, and no charge that grows with how well the Grid does its job.
One flat annual fee per GPU, set as a percentage of the current market cost of the system class you bring. A fleet of eight H100 systems pays the same whether the Grid places workloads on it around the clock or holds capacity for your own bursts.
Multi-year commitments carry meaningful discounts, including partial-upfront structures that reduce the annual fee. Talk to us about term pricing for your fleet.
Placement, health monitoring, recovery, and the same telemetry and observability the public Grid runs on. Fleets not directly connected to the internet route through the Kinesis proxy layer. A generous data allowance is included, with metered transfer beyond it.
Larger fleets qualify for volume pricing, and dedicated support arrangements are available. Contact us and we will scope it with you.
The same workloads that cost the most on traditional clouds save the most on True-Util™.
H100s sit idle between prompts. The bill is the same whether you served 100 requests or 100,000.
True-Util™ Shared meters inference time only. No queries, no cost. Bursty traffic caps at the Reserved rate.
H100s sit idle between prompts. The bill is the same whether you served 100 requests or 100,000.
The Kinesis winTrue-Util™ Shared meters inference time only. No queries, no cost. Bursty traffic caps at the Reserved rate.
Overprovisioning for traffic that hasn’t shown up. Or worse — under-provisioning and falling over the first time it does.
Pay pennies at low traffic. Costs cap at the Reserved rate during spikes. Headroom without prepayment.
Overprovisioning for traffic that hasn’t shown up. Or worse — under-provisioning and falling over the first time it does.
The Kinesis winPay pennies at low traffic. Costs cap at the Reserved rate during spikes. Headroom without prepayment.
Staging servers run 24/7 to be ready, but burn nights and weekends.
True-Util™ drops the bill as activity drops. Same reservation, lower cost when the team’s asleep.
Staging servers run 24/7 to be ready, but burn nights and weekends.
The Kinesis winTrue-Util™ drops the bill as activity drops. Same reservation, lower cost when the team’s asleep.
Reserved AWS instances, on-prem servers, donated lab GPUs — capacity already paid for, sitting underused.
Run the Kinesis grid on your hardware at 20% of Shared. Same orchestration, FinOps visibility, 80% less spend on what you already own.
Reserved AWS instances, on-prem servers, donated lab GPUs — capacity already paid for, sitting underused.
The Kinesis winRun the Kinesis grid on your hardware at 20% of Shared. Same orchestration, FinOps visibility, 80% less spend on what you already own.
$100 in free credit. No credit card required. Deploy your first container in under five minutes. Bring a GitHub repo, a Dockerfile, or just describe what you want.