CloudNova — The AI Cloud Platform
NEW CloudNova 2.0 — Serverless GPUs are here

The AI cloud built for
velocity, not overhead.

Elastic GPU clusters, sub-50ms global inference, and autoscaling that drops to zero — so your team ships models, not infrastructure.

No credit card required  ·  $200 free compute  ·  First deploy in 90 seconds

console.cloudnova.ai / inference ⌘K

Inference Overview

LIVE
1H24H7D30D
REQUESTS · 24H
8.4B
▲ 12.3%
P99 LATENCY
42ms
▼ 8ms
GPU UTILIZATION
87%
▲ 4.1%
us-east-112ms
eu-west-231ms
ap-south-144ms

POWERING PRODUCTION AI AT

HELIOSquanticaNEURALFORGELuminaOCTANEvertex_labsKite.aiDATAWISE HELIOSquanticaNEURALFORGELuminaOCTANEvertex_labsKite.aiDATAWISE
Platform

Everything between your model
and planet-scale production.

One console for the full model lifecycle — training runs, inference endpoints, spend, and security. No YAML archaeology required.

Elastic GPU clusters

Reserve H100 and B200 capacity in seconds, not quarters. Burst across a 120K-GPU pool with workload-aware scheduling that packs inference and training side by side.

cluster · prod-east · h100×512utilization 87%

Global edge inference

34 regions, one endpoint. Requests route to the nearest warm replica automatically — p99 stays flat wherever your users are.

us-east12ms
eu-west31ms
ap-south44ms

Autoscale to zero

Scale on queue depth and concurrency, drop to zero when traffic sleeps. Billing is per-second, down to the millisecond.

$0.00 /hr while idle

Observability, built in

Token-level traces, GPU flame graphs, and cost attribution per model, team, or API key — no agents to install, no exporters to babysit.

POST /v1/chat/completions · 1,204 tokens · 38ms14:02:11
trace 8f3a…c2 · embedding → rerank → generate14:02:10
autoscale event · replicas 6 → 14 · cold start 812ms14:01:58
budget alert · team/search at 82% of monthly cap13:58:40

Model registry & versioning

Version every weight, adapter, and prompt template. One-click rollbacks and canary splits keep deploys boring.

v1.4.2 LATEST v1.4.1 v1.3.9 v1.3.8

Enterprise-grade security

Isolated VPCs, customer-managed keys, and region-pinned data residency — audited continuously, not annually.

SOC 2 TYPE II HIPAA ISO 27001 GDPR
0%
Uptime SLA
credit-backed, multi-AZ
0ms
p99 global latency
measured across 34 regions
0K+
GPUs in the pool
H100 · B200 · L40S
0B
Inferences daily
and counting, 24/7
Workflow

From notebook to planet-scale
in four moves.

A CLI-first workflow that fits the way ML teams already work. Git push, and CloudNova handles the rest.

01

Connect

Point CloudNova at your repo or model registry. We resolve images, cache weights, and pre-warm the runtime.

$ nova init
02

Deploy

One command spins up warm replicas with health checks, TLS, and a stable endpoint out of the box.

$ nova deploy --gpu h100
03

Scale

Traffic-driven autoscaling with sub-second cold starts — and a hard floor of zero when nobody's calling.

min: 0 · max: 512
04

Observe

Live traces, spend alerts, and SLO burn rates streamed straight to your dashboard and pager.

$ nova logs --follow

Average time from nova init to production traffic: 11 minutes.

Customers

Teams ship faster on CloudNova.

From two-person labs to public companies — here's what happens after the first deploy.

We cut inference spend by 62% in the first month. Scale-to-zero felt like a rounding error until we saw the invoice. Closest thing to magic I've seen in infra.

MC
Maya Chen
VP Engineering · Quill AI

We migrated 40 models over a single weekend. The console is what every cloud product wishes it looked like — my team actually enjoys doing deploys now.

DO
Derek Osei
Head of Platform · Helios

p99 stayed flat at 40ms through a 30× traffic spike on launch day. Our on-call engineer finally sleeps through the night — that alone is worth the contract.

SM
Sofia Marques
CTO · Lumen Labs
Pricing

Start free. Scale when you do.

Per-second billing on every plan. No reservations, no minimums, no surprise invoices.

Starter
For side projects and prototypes
$0/mo
+ usage at cost
  • 100 GPU-hours included monthly
  • 2 concurrent deployments
  • 8 starter regions
  • Community support
Start for free
Enterprise
For regulated & high-volume workloads
Custom
volume pricing · committed capacity
  • Dedicated GPU clusters & VPC peering
  • 99.99% SLA with credit guarantees
  • Region-pinned data residency & CMK
  • Dedicated solutions engineer
  • SSO / SCIM & audit log export
Talk to sales
FAQ

Questions, answered.

Everything teams ask before their first deploy. Still curious? Talk to an engineer →

Per second, down to the millisecond, from the moment a replica is scheduled until it's released. When your endpoint scales to zero, you pay exactly $0.00 — there are no idle fees, warm-pool surcharges, or minimum commitments on Starter and Pro.

Anything — open-weight models from Hugging Face, your own fine-tunes, ONNX exports, or a custom container. CloudNova auto-detects the runtime, quantizes optionally, and exposes an OpenAI-compatible endpoint alongside your own API contract.

Under 900ms at p95 for models up to 70B parameters. We snapshot GPU memory state and keep weight caches pinned in every region, so "scale from zero" behaves like scale-from-one — your users never see the difference.

Wherever you pin it. Choose a home region per deployment, and prompts, weights, and logs never leave it. Enterprise plans add dedicated VPCs, customer-managed encryption keys, and region-locked replication for GDPR, HIPAA, and residency mandates.

Yes — nova import reads your existing serving configs (SageMaker, Vertex, AKS) and generates CloudNova equivalents in minutes. Pro teams get a guided migration review; Enterprise gets white-glove migration with an engineer on call.

Starter includes community Discord and docs. Pro adds priority tickets with a 1-hour response SLA during business hours. Enterprise gets a dedicated solutions engineer, a shared Slack channel, and 24/7 incident response with 15-minute paging.

Launch

Your model deserves a
faster home.

Spin up your first endpoint in the next two minutes. $200 in compute credits, on us.

No credit card · Cancel anytime · SOC 2 Type II

互动区域

登录后可以点赞此内容

参与互动

登录后可以点赞和评论此内容,与作者互动交流