The stack you license to run a cloud

You license a general-purpose, multi-tenant stack, deploy it on hardware you already own, and let your tenants provision compute, network and storage themselves through your portal and your API. The IaaS layer gives them virtual machines or whole bare-metal nodes, per-tenant networking, and block, S3-compatible object and POSIX file. The platform layer above it covers managed Kubernetes, Slurm, notebooks and model serving; anything past that is a product you would build on top, and we say so rather than let it sit implied in a title. What you buy is software — not capacity, not hosting, not a managed service.

What the operator gets

Hardware in, revenue out

Arrow
Your unsold capacity becomes sellable: Around 20% of the hours you already sell are lost to execution-idle inside running jobs, and one cluster serving many tenants roughly doubles the real compute share per accelerator. Modelled against published rental prices, that is about 30% additional revenue per installed accelerator with no new hardware. These are published research figures rather than our own records; the full build-up, including where the number does not apply to you, is on the monetization page.
Arrow
You serve both markets at once: You sign the long contracts that make a cluster financeable and sell the hours they leave open into the short-term on-demand market, where the average realised price per hour is markedly higher.
Arrow
What you do have to buy: Nothing changes on the accelerators or the fabric you already bought. You add control-plane and management nodes, out-of-band reachability on every node, and — only if you intend to sell high-throughput shared file — the VAST hardware behind that service.
Arrow
You reach the first invoice quickly: From signed contract, pilot and cluster onboarding included, the first invoice follows in two weeks. Bringing a further cluster into production takes 2 days.
Engineering

The parts that are hard to copy

Orchestration cut down to VMs
We cut Kubernetes back to the scheduler and the reconciliation loop and pointed it at virtual machines. Its relation to KubeVirt is direct: we run upstream KubeVirt unmodified, with our own scheduler, network and metering operators on top — we would rather name the upstream than imply we invented one. Ours is what sits around it: placement against fabric topology, the tenancy model, quota enforcement and the accelerator lifecycle.
Arrow
Tenant separation on the fabric
You isolate tenants over InfiniBand partition keys, enforced at the adapter and the switch port, with the subnet manager as the source of truth. We contribute to SONiC on white-box Broadcom hardware — platform and SAI fixes for the Broadcom white boxes we run — and run a user-space data plane, whose honest measure is a rate against a core budget and a feature set: 100 Gbit/s per node at 1500-byte frames on four cores, with per-tenant ACLs, NAT and policing.
Arrow
Eight years of receipts
We were the operator: on the market since 2018, more than 20,000 users and 5,000+ B2B customers since 2023, the first non-US company to receive Intel Gaudi, and roughly 5 PiB of VAST integrated behind one API.
Arrow
Day two

What your team still does, week to week

Arrow
Node health and RMA: You watch accelerator health, act on Xid and ECC events, burn in replacements and carry the vendor RMA. The stack cordons a failing node and moves the tenant off it; it does not open the ticket with your supplier.
Arrow
Firmware and drivers: BMC, BIOS, HCA and accelerator driver levels are yours to hold consistent, and we publish the tested matrix per release.
Arrow
Fabric: You own the subnet manager hosts, the cabling and the link error budget; we ship the tenancy configuration that runs on top.
Arrow
Storage: Ceph day-two work — OSD replacement, rebalancing, capacity headroom — stays with your team, and VAST stays on your own support contract.
Arrow
Platform: Control-plane upgrades, backup and restore of control-plane state, and the tenant-facing incident. That list is what a smaller, less specialised team actually means: Linux, network and storage operations on a rota, not a platform engineering organisation.
Demand

You still have to find the buyers

Arrow
Sellable is not sold: You get capacity that is sellable. You still have to sell it. If you have no route to buyers, no platform will invent one.
Arrow
You lose the deal-size floor: Reservation selling costs a contract and a human per customer. With self-service, someone who wants four accelerators for nine days buys them at 2 a.m. and costs you nothing to serve.
Arrow
Where your demand comes from: Your existing colocation and connectivity customers, national research and public-sector buyers in your jurisdiction, and the long tail that never speaks to a salesperson.
Svg

Questions operators ask first

Does this run on the hardware I already have?
Substrack Icon

Yes, within a list you can read before you call us: the Genesis Grid hardware compatibility list names the accelerators, hosts, fabric and subnet manager we support, and the storage releases we test against. Mixed vendors and mixed generations serve tenants from the same cluster. What mixing does not do is put two accelerator generations inside one training job, and we would rather write that down than let a tenant discover it.

Do I need a Tier 3 facility?
Plus Icon

Uptime Institute Tiers I to IV describe topology — redundancy, and whether you can maintain power and cooling while the site runs. They are not availability percentages, and Uptime has not published percentages for years. Software resilience changes how much redundancy the building has to carry for a given class of tenant; it does not replace a second power path. A deployment is one cluster in one building: lose the power there and the cluster goes with it. If you sell availability that depends on maintenance while the site runs, you need a facility that delivers it, and we will say so before you write it into an SLA.

Do you operate it for me?
Plus Icon

No. You operate it, and we license you the software and ship releases. We ran it ourselves for eight years — our own sites were tenant zero and we carried the on-call rota — which is why the operational knowledge sits in the product. Volta is the licensee we may name, and Volta runs its own clusters itself.

Who holds the SLA towards my tenants?
Plus Icon

You do. It stays your brand, your contract and your tenant relationship, and you get the isolation, the headroom and the telemetry to stand behind it. Underneath, we commit 99.9 % for the control plane, the portal and the meter, and a 30-minute severity-one response, 24/7, for support. Availability of a tenant environment is a product of your building, your hardware and this software, so it is agreed per deployment.

What happens during an upgrade?
Plus Icon

A platform upgrade costs your control plane, portal and meter under 15 minutes; instances already running are untouched. Node maintenance is a different thing: you cordon the node, the tenant is notified 72 hours in advance, and the workload either migrates live or resumes from its last checkpoint.

Next step
Bring us your fleet, we will show what it sells
Arrow