You license Genesis Grid and run it on the racks, fabric and storage you already own. It adopts your inventory as it is and turns one compute cluster into a general-purpose multi-tenant IaaS and PaaS product, with API, identity, quotas and metering. Nothing is ripped out and replaced.
Three numbers your operations team feels first: how quickly a further cluster contributes, how long a signed contract takes to reach the first invoice, and what an upgrade costs the control plane.
to stand up a further cluster on installed hardware
from signed contract to first invoice, onboarding included
control plane, portal and metering downtime during a platform upgrade
You mix vendors, generations and densities in one cluster. Scheduling goes against what a node can actually deliver, so older accelerators keep earning next to the newest ones.
Enforced by the subnet manager on the fabric and by the storage services themselves, one layer below anything a tenant workload can reach. The mechanism is set out in full on the multi-tenancy page.
Software resilience shifts how much redundancy a building has to carry for a given tenant class. It never manufactures availability the building cannot deliver, and the FAQ below says where it stops.
You serve long-term contracts and the short-term on-demand market from the same cluster, which pushes idle capacity towards a minimum.
No. Genesis Grid adopts the hardware, the network and the facility you have: NVIDIA and AMD accelerators, x86 and Arm hosts, Ethernet and InfiniBand fabrics. The supported list is a published hardware compatibility list, revised with every release, and you get it before the contract, not during rollout.
Tenants get virtual machines or whole bare-metal nodes; both are placed by the same scheduler and metered by the same meter. Whichever shape a tenant contract demands is provisioned, isolated and metered by the same control plane. What that costs a multi-node all-reduce against a run on the metal we quote from a benchmark on your own hardware, not from a slide.
Not in the way the question hopes. Uptime Institute's Tier I to IV are topology classes about redundancy and maintainability while the load runs, not availability percentages — Uptime has not published percentages for years. Software resilience shifts how much redundancy the building has to carry for a given tenant class: work is rebuilt against remaining capacity, so a node, rack or switch failure stops being a customer-visible event. It does not replace redundancy. One deployment is one cluster in one building, and if that building loses power the cluster goes with it. If you sell availability that depends on maintaining plant while tenants keep running, you need a facility that delivers it. The room is still yours.
Affected slices are rebuilt against remaining capacity in the same cluster, and the headroom that takes is a fraction of the installed base. Planned maintenance is a different event: the node is cordoned, the tenant is notified 72 hours in advance, and the workload either migrates live or resumes from its last checkpoint.
You do. Genesis Grid is licensed software, and Genesis carries escalation when the fault is in the stack — response times and coverage are 30 minutes on severity one, 24/7, with escalation to the engineers who build the stack. What changes is the shape of your team: smaller and less deeply specialised than a platform engineering organisation. You still need people, just fewer.