Kinesis
Build & Run

Your priorities. Our orchestration.

Choose what matters: cost, reliability, latency, or provider diversity. Kinesis places and operates your workload to match.

Cost

Match your workload to the lowest-cost compute that meets its requirements.

Reliability

Use proven supply, with automatic recovery when a node fails.

Latency

Run close to your users and adapt placement as traffic shifts.

Multi-cloud

Distribute workloads across providers from one control plane.

Connect a repo or bring a container. Get a running app.

The right hardware for your workload

Access CPUs and GPUs across vetted providers. Kinesis handles placement and operations, with one deployment workflow wherever your app runs.

Write once, run anywhere

Run the same container across clouds, partner datacenters, and your own hardware.

Operations come standard

Monitoring, scaling, recovery, TLS, and secrets are built into each deployment.

Pay only for real usage

Serverless billing follows resource usage, capped at the equivalent Dedicated rate.

The Kinesis Difference

Built into every deployment

LESS OPS SURFACE

Networking, scaling, health checks, and rollbacks are handled by the platform.

~80%

Fewer actions than a hyperscaler

TRUE-UTIL™ PRICING

Usage-based billing, with a cap

Pay for actual resource usage, up to the equivalent Dedicated rate. Bursty workloads save the most.

usage → billed

HARDWARE THAT FITS

CPU, GPU, big-iron, on-demand

A100, H100, H200, B200. Multi-card for training, single card for inference.

PORTABLE BY CONSTRUCTION

Standard containers

Standard Dockerfiles and images keep your app portable.

OBSERVABILITY BUILT IN

Logs, metrics, and cost together

Inspect live logs, resource usage, and spend for each app in one console.

FOR THE ARCHITECTS

How placement works

See how Kinesis matches workloads to hardware and adapts as demand and availability change.

Under the hood →
BUILD ANY WAY, RUN ON KINESIS

Your code. Your intent.
A running URL.

Keep your tools and frameworks. Kinesis builds and runs your app, then handles placement, scaling, and recovery. Every deployment stays portable.

1 Bring your code

Connect your repo or push a standard container from your existing workflow.

2 Set your intent

Set your priorities for cost, reliability, latency, and provider diversity.

3 Kinesis runs it

Placement, scaling, recovery, and monitoring are handled. You get a live URL.

KINESIS DEPLOY
  • Import a repo, container, or generated app.
  • Set the priorities that guide placement.
  • Keep a standard container you can run elsewhere.
  • Get networking, autoscaling, recovery, monitoring, and rollbacks built in.
  • Meter actual resource usage with True-Util™, capped at the equivalent Dedicated rate.
Workload Deployment

Start wherever you are

Six ways to deploy. One runtime, one control plane.

GITHUB · START HERE

Connect a repo. Push to ship.

Connect your GitHub project. Each push triggers a build and deployment.

Best for: teams · production workflows · continuous delivery

REGISTRY

Bring an image from your registry.

Use a public or private registry. Choose your compute and deploy.

Best for: existing apps · proprietary environments · full control

DOCKERFILE

Upload a Dockerfile or ZIP. We build.

Upload your source and Dockerfile. Kinesis builds and runs the image.

Best for: reproducible builds · portability · open-source projects

IMAGE UPLOAD

Push an image file directly.

Already built locally? Upload the image and go live without setting up a registry.

Best for: quick proofs · offline builds · restricted networks

APP GALLERY

Start from a template.

Choose an LLM, vector database, web framework, or batch runner. Configure and deploy.

Best for: standard apps · ready-to-run stacks

PROMPT

Or just describe it.

Describe your app and generate a container with your model or ours. Edit it, then deploy.

Best for: prototypes · MVPs · exploring ideas

The True-Util™ Model

Serverless or Dedicated

Use Serverless for variable demand or Dedicated for sustained usage. Both run on the same grid with the same operational tools.

TRUE-UTIL™ SERVERLESS

Metered usage, capped at the Dedicated rate

Pay for compute, memory, storage, and bandwidth used, capped at the equivalent Dedicated rate. Best suited to variable demand.

  • Best for inference, dev/staging, agencies, MVPs
  • Runs across a range of CPU and GPU classes
  • No upfront commitments
TRUE-UTIL™ DEDICATED

The whole machine, billed by the hour

Reserve a machine for your workload and pay by the hour. Get full control with Kinesis orchestration and telemetry.

  • Best for steady training, production HPC, regulated workloads
  • Choice across providers, which reduces lock-in
  • Full control over configuration, performance, privacy
Pricing

Rates

Rates effective July 2026. GPU rates are per GPU per hour.

Full pricing details
A100$1.35

Per GPU / Per Hour

1x GPU, 28 CPUs, 120GB RAM, 750GB Storage

Available in 1x, 2x, 4x GPU configurations.

H100$2.50

Per GPU / Per Hour

1x GPU, 28 CPUs, 180GB RAM, 750GB Storage

Available in 1x, 2x, 4x GPU configurations.

A100 NVLink$1.50

Per GPU / Per Hour

8x GPU node: 252 CPUs, 1920GB RAM, 6500GB Storage

Priced per GPU, sold as 8-GPU NVLink nodes.

H100 NVLink$2.75

Per GPU / Per Hour

8x GPU node: 252 CPUs, 1440GB RAM, 6500GB Storage

Priced per GPU, sold as 8-GPU NVLink nodes.

H200 SXM$4.25

Per GPU / Per Hour

8x GPU node: 176 CPUs, 1800GB RAM, 48000GB Storage

Priced per GPU, sold as 8-GPU NVLink nodes.

B200 SXM$6.50

Per GPU / Per Hour

8x GPU node: 252 CPUs, 2048GB RAM, 40000GB Storage

Priced per GPU, sold as 8-GPU NVLink nodes.

B300 SXMfrom $8.13

Per GPU / Per Hour

8x GPU node: 252 CPUs, 2048GB+ RAM, 40000GB+ Storage

Spot capacity. Price varies with market conditions.

Check availability
Compute Optimized CPU$0.035

Per vCPU / Per Hour

1 vCPU, 2GB RAM, 50GB NVMe

Available as serverless True-Util™. Available as 2x, 4x, 8x, 16x, 32x and 64x configurations.

General Purpose CPU$0.045

Per vCPU / Per Hour

1 vCPU, 4GB RAM, 50GB NVMe

Available as serverless True-Util™. Available as 2x, 4x, 8x, 16x, 32x and 64x configurations.

Memory Optimized CPU$0.055

Per vCPU / Per Hour

1 vCPU, 8GB RAM, 50GB NVMe

Available as serverless True-Util™. Available as 2x, 4x, 8x, 16x, 32x and 64x configurations.

Prices shown are for representative configurations. Actual specifications may vary by provider. GPU rates are per GPU per hour. NVLink and SXM systems are sold as 8-GPU nodes. Rates marked “from” are spot rates and vary with market conditions.

Try it on a real app

Bring a repo or container and start with $100 in free credit. No credit card required.