Your priorities. Our orchestration.
Choose what matters: cost, reliability, latency, or provider diversity. Kinesis places and operates your workload to match.
Cost
Match your workload to the lowest-cost compute that meets its requirements.
Reliability
Use proven supply, with automatic recovery when a node fails.
Latency
Run close to your users and adapt placement as traffic shifts.
Multi-cloud
Distribute workloads across providers from one control plane.
Connect a repo or bring a container. Get a running app.
The right hardware for your workload
Access CPUs and GPUs across vetted providers. Kinesis handles placement and operations, with one deployment workflow wherever your app runs.
Write once, run anywhere
Run the same container across clouds, partner datacenters, and your own hardware.
Operations come standard
Monitoring, scaling, recovery, TLS, and secrets are built into each deployment.
Pay only for real usage
Serverless billing follows resource usage, capped at the equivalent Dedicated rate.
Built into every deployment
LESS OPS SURFACE
Networking, scaling, health checks, and rollbacks are handled by the platform.
Fewer actions than a hyperscaler
TRUE-UTIL™ PRICING
Usage-based billing, with a cap
Pay for actual resource usage, up to the equivalent Dedicated rate. Bursty workloads save the most.
usage → billed
HARDWARE THAT FITS
CPU, GPU, big-iron, on-demand
A100, H100, H200, B200. Multi-card for training, single card for inference.
PORTABLE BY CONSTRUCTION
Standard containers
Standard Dockerfiles and images keep your app portable.
OBSERVABILITY BUILT IN
Logs, metrics, and cost together
Inspect live logs, resource usage, and spend for each app in one console.
FOR THE ARCHITECTS
How placement works
See how Kinesis matches workloads to hardware and adapts as demand and availability change.
Under the hood →Your code. Your intent.
A running URL.
Keep your tools and frameworks. Kinesis builds and runs your app, then handles placement, scaling, and recovery. Every deployment stays portable.
1 Bring your code
Connect your repo or push a standard container from your existing workflow.
2 Set your intent
Set your priorities for cost, reliability, latency, and provider diversity.
3 Kinesis runs it
Placement, scaling, recovery, and monitoring are handled. You get a live URL.
- •Import a repo, container, or generated app.
- •Set the priorities that guide placement.
- •Keep a standard container you can run elsewhere.
- •Get networking, autoscaling, recovery, monitoring, and rollbacks built in.
- •Meter actual resource usage with True-Util™, capped at the equivalent Dedicated rate.
Start wherever you are
Six ways to deploy. One runtime, one control plane.
Connect a repo. Push to ship.
Connect your GitHub project. Each push triggers a build and deployment.
Best for: teams · production workflows · continuous delivery
Bring an image from your registry.
Use a public or private registry. Choose your compute and deploy.
Best for: existing apps · proprietary environments · full control
Upload a Dockerfile or ZIP. We build.
Upload your source and Dockerfile. Kinesis builds and runs the image.
Best for: reproducible builds · portability · open-source projects
Push an image file directly.
Already built locally? Upload the image and go live without setting up a registry.
Best for: quick proofs · offline builds · restricted networks
Start from a template.
Choose an LLM, vector database, web framework, or batch runner. Configure and deploy.
Best for: standard apps · ready-to-run stacks
Or just describe it.
Describe your app and generate a container with your model or ours. Edit it, then deploy.
Best for: prototypes · MVPs · exploring ideas
Serverless or Dedicated
Use Serverless for variable demand or Dedicated for sustained usage. Both run on the same grid with the same operational tools.
Metered usage, capped at the Dedicated rate
Pay for compute, memory, storage, and bandwidth used, capped at the equivalent Dedicated rate. Best suited to variable demand.
- Best for inference, dev/staging, agencies, MVPs
- Runs across a range of CPU and GPU classes
- No upfront commitments
The whole machine, billed by the hour
Reserve a machine for your workload and pay by the hour. Get full control with Kinesis orchestration and telemetry.
- Best for steady training, production HPC, regulated workloads
- Choice across providers, which reduces lock-in
- Full control over configuration, performance, privacy
Per GPU / Per Hour
1x GPU, 28 CPUs, 120GB RAM, 750GB Storage
Available in 1x, 2x, 4x GPU configurations.
Per GPU / Per Hour
1x GPU, 28 CPUs, 180GB RAM, 750GB Storage
Available in 1x, 2x, 4x GPU configurations.
Per GPU / Per Hour
8x GPU node: 252 CPUs, 1920GB RAM, 6500GB Storage
Priced per GPU, sold as 8-GPU NVLink nodes.
Per GPU / Per Hour
8x GPU node: 252 CPUs, 1440GB RAM, 6500GB Storage
Priced per GPU, sold as 8-GPU NVLink nodes.
Per GPU / Per Hour
8x GPU node: 176 CPUs, 1800GB RAM, 48000GB Storage
Priced per GPU, sold as 8-GPU NVLink nodes.
Per GPU / Per Hour
8x GPU node: 252 CPUs, 2048GB RAM, 40000GB Storage
Priced per GPU, sold as 8-GPU NVLink nodes.
Per GPU / Per Hour
8x GPU node: 252 CPUs, 2048GB+ RAM, 40000GB+ Storage
Spot capacity. Price varies with market conditions.
Check availabilityPer vCPU / Per Hour
1 vCPU, 2GB RAM, 50GB NVMe
Available as serverless True-Util™. Available as 2x, 4x, 8x, 16x, 32x and 64x configurations.
Per vCPU / Per Hour
1 vCPU, 4GB RAM, 50GB NVMe
Available as serverless True-Util™. Available as 2x, 4x, 8x, 16x, 32x and 64x configurations.
Per vCPU / Per Hour
1 vCPU, 8GB RAM, 50GB NVMe
Available as serverless True-Util™. Available as 2x, 4x, 8x, 16x, 32x and 64x configurations.
Prices shown are for representative configurations. Actual specifications may vary by provider. GPU rates are per GPU per hour. NVLink and SXM systems are sold as 8-GPU nodes. Rates marked “from” are spot rates and vary with market conditions.
Try it on a real app
Bring a repo or container and start with $100 in free credit. No credit card required.