One cluster, one control plane, many tenants

You license Genesis Grid and run it on the racks, fabric and storage you already own. It adopts your inventory as it is and turns one compute cluster into a general-purpose multi-tenant IaaS and PaaS product, with API, identity, quotas and metering. Nothing is ripped out and replaced.

The stack, bottom up

From bare metal to billed usage

Arrow
Hardware: You keep your building, your power envelope and your servers. Genesis Grid inventories accelerators, CPU, memory and storage per rack and manages the cluster as one unit, mixed vendors and generations included. New hardware is limited to control-plane and management nodes and out-of-band reachability.
Arrow
Fabric: Your tenants share one InfiniBand fabric without sharing a network. Partition keys separate them in the fabric itself, and the Ethernet side is segmented per tenant. The scope of all of it is the cluster: several sites means several clusters, each with its own control plane, failure domain and tenants. We do not claim that separate buildings behave as one elastic pool, because they do not.
Arrow
Orchestration: You run Kubernetes stripped back to the one job it does well here — orchestrating virtual machines instead of containers. Placement, capacity and node lifecycle are decided above that layer. What the PaaS layer on top covers is managed Kubernetes, Slurm, notebooks and model serving.
Arrow
Storage: Your tenants reach three services behind one API and one meter: Ceph RBD for block, Ceph RGW for S3-compatible object, VAST for POSIX file. Genesis integrated around 5 PiB of VAST behind this interface on the estate it ran itself.
Arrow
Identity and metering: You get an IAM and organisation model in the shape AWS established — accounts, organisations, roles, policies — and consumption metered at the point of use, from the moment the first node joins.
Engineering

Five decisions a systems engineer recognises

Arrow
Virtual machines, not containers: The control loop is Kubernetes, cut back to the scheduler and the reconciliation loop and pointed at virtual machines. If your engineers are about to ask how that relates to KubeVirt, the answer is direct: we run upstream KubeVirt unmodified, with our own scheduler, network and metering operators on top — better to name the upstream than imply we invented one. Ours is what sits around it: placement against fabric topology, the tenancy model, quota enforcement, accelerator lifecycle.
Arrow
Partition keys on InfiniBand: Tenant separation runs on partition keys in the fabric, enforced by the subnet manager — an unusual place to put a tenant boundary, unusual enough that it surprised NVIDIA's own engineers. Where that enforcement sits is set out on the multi-tenancy page.
Arrow
SONiC on white-box switches: Your leaf and spine run SONiC on white-box Broadcom hardware, so your network layer is not bound to one vendor's licence terms. What we contribute upstream — platform and SAI fixes for the Broadcom white boxes we run — is the same build your fabric runs, not a private fork.
Arrow
A user-space data plane: Packets run through VPP and block I/O through SPDK, both in user space, so the tenant data path does not spend the cores your accelerators need. The honest form of that number is a rate against a core budget and a feature set: 100 Gbit/s per node at 1500-byte frames on four cores, with per-tenant ACLs, NAT and policing, measured on your fabric before any of it enters a service level.
Arrow
The exit from OpenStack: We ran OpenStack as an operator, hit its limits with real tenants on it, and rewrote the stack twice over eight years. What you license is that result, not somebody else's distribution.
Operating envelope

What the architecture actually changes

Three numbers your operations team feels first: how quickly a further cluster contributes, how long a signed contract takes to reach the first invoice, and what an upgrade costs the control plane.

2 days

to stand up a further cluster on installed hardware

2 weeks

from signed contract to first invoice, onboarding included

under 15 minutes

control plane, portal and metering downtime during a platform upgrade

Svg

Questions your infrastructure team will ask

Do we have to replace what we already run?
Substrack Icon

No. Genesis Grid adopts the hardware, the network and the facility you have: NVIDIA and AMD accelerators, x86 and Arm hosts, Ethernet and InfiniBand fabrics. The supported list is a published hardware compatibility list, revised with every release, and you get it before the contract, not during rollout.

Do my tenants get virtual machines or bare metal?
Plus Icon

Tenants get virtual machines or whole bare-metal nodes; both are placed by the same scheduler and metered by the same meter. Whichever shape a tenant contract demands is provisioned, isolated and metered by the same control plane. What that costs a multi-node all-reduce against a run on the metal we quote from a benchmark on your own hardware, not from a slide.

Can software resilience make up for a lower facility class?
Plus Icon

Not in the way the question hopes. Uptime Institute's Tier I to IV are topology classes about redundancy and maintainability while the load runs, not availability percentages — Uptime has not published percentages for years. Software resilience shifts how much redundancy the building has to carry for a given tenant class: work is rebuilt against remaining capacity, so a node, rack or switch failure stops being a customer-visible event. It does not replace redundancy. One deployment is one cluster in one building, and if that building loses power the cluster goes with it. If you sell availability that depends on maintaining plant while tenants keep running, you need a facility that delivers it. The room is still yours.

What happens when a node fails?
Plus Icon

Affected slices are rebuilt against remaining capacity in the same cluster, and the headroom that takes is a fraction of the installed base. Planned maintenance is a different event: the node is cordoned, the tenant is notified 72 hours in advance, and the workload either migrates live or resumes from its last checkpoint.

Who runs the platform once it is installed?
Plus Icon

You do. Genesis Grid is licensed software, and Genesis carries escalation when the fault is in the stack — response times and coverage are 30 minutes on severity one, 24/7, with escalation to the engineers who build the stack. What changes is the shape of your team: smaller and less deeply specialised than a platform engineering organisation. You still need people, just fewer.

Your topology, our solutions architect
Walk the architecture through, layer by layer
Arrow