Elastic GPU clusters
Reserve H100 and B200 capacity in seconds, not quarters. Burst across a 120K-GPU pool with workload-aware scheduling that packs inference and training side by side.
Elastic GPU clusters, sub-50ms global inference, and autoscaling that drops to zero — so your team ships models, not infrastructure.
No credit card required · $200 free compute · First deploy in 90 seconds
POWERING PRODUCTION AI AT
One console for the full model lifecycle — training runs, inference endpoints, spend, and security. No YAML archaeology required.
Reserve H100 and B200 capacity in seconds, not quarters. Burst across a 120K-GPU pool with workload-aware scheduling that packs inference and training side by side.
34 regions, one endpoint. Requests route to the nearest warm replica automatically — p99 stays flat wherever your users are.
Scale on queue depth and concurrency, drop to zero when traffic sleeps. Billing is per-second, down to the millisecond.
Token-level traces, GPU flame graphs, and cost attribution per model, team, or API key — no agents to install, no exporters to babysit.
Version every weight, adapter, and prompt template. One-click rollbacks and canary splits keep deploys boring.
Isolated VPCs, customer-managed keys, and region-pinned data residency — audited continuously, not annually.
A CLI-first workflow that fits the way ML teams already work. Git push, and CloudNova handles the rest.
Point CloudNova at your repo or model registry. We resolve images, cache weights, and pre-warm the runtime.
$ nova init
One command spins up warm replicas with health checks, TLS, and a stable endpoint out of the box.
$ nova deploy --gpu h100
Traffic-driven autoscaling with sub-second cold starts — and a hard floor of zero when nobody's calling.
min: 0 · max: 512
Live traces, spend alerts, and SLO burn rates streamed straight to your dashboard and pager.
$ nova logs --follow
Average time from nova init to production traffic: 11 minutes.
From two-person labs to public companies — here's what happens after the first deploy.
We cut inference spend by 62% in the first month. Scale-to-zero felt like a rounding error until we saw the invoice. Closest thing to magic I've seen in infra.
We migrated 40 models over a single weekend. The console is what every cloud product wishes it looked like — my team actually enjoys doing deploys now.
p99 stayed flat at 40ms through a 30× traffic spike on launch day. Our on-call engineer finally sleeps through the night — that alone is worth the contract.
Per-second billing on every plan. No reservations, no minimums, no surprise invoices.
Everything teams ask before their first deploy. Still curious? Talk to an engineer →
Per second, down to the millisecond, from the moment a replica is scheduled until it's released. When your endpoint scales to zero, you pay exactly $0.00 — there are no idle fees, warm-pool surcharges, or minimum commitments on Starter and Pro.
Anything — open-weight models from Hugging Face, your own fine-tunes, ONNX exports, or a custom container. CloudNova auto-detects the runtime, quantizes optionally, and exposes an OpenAI-compatible endpoint alongside your own API contract.
Under 900ms at p95 for models up to 70B parameters. We snapshot GPU memory state and keep weight caches pinned in every region, so "scale from zero" behaves like scale-from-one — your users never see the difference.
Wherever you pin it. Choose a home region per deployment, and prompts, weights, and logs never leave it. Enterprise plans add dedicated VPCs, customer-managed encryption keys, and region-locked replication for GDPR, HIPAA, and residency mandates.
Yes — nova import reads your existing serving configs (SageMaker, Vertex, AKS) and generates CloudNova equivalents in minutes. Pro teams get a guided migration review; Enterprise gets white-glove migration with an engineer on call.
Starter includes community Discord and docs. Pro adds priority tickets with a 1-hour response SLA during business hours. Enterprise gets a dedicated solutions engineer, a shared Slack channel, and 24/7 incident response with 15-minute paging.
Spin up your first endpoint in the next two minutes. $200 in compute credits, on us.
No credit card · Cancel anytime · SOC 2 Type II