Expanse

Compute resource prediction platform for HPC and GPU clusters that right-sizes job submissions and diagnoses failures before they waste GPU time.

expanse.sh

Community-submitted · content unverified. This organization has not verified ownership or the information in this profile.

About

Expanse integrates with SLURM, Kubernetes, or Nomad via a single install command, collects telemetry, and uses pre-trained models continuously adapted to the customer's cluster to predict memory, runtime, GPU allocation, and failure risk before each job runs. The `expanse analyse` command delivers predictions as confidence distributions without changing existing submission scripts; `expanse diagnose` performs root-cause analysis on failed jobs with code-diff fix suggestions. In benchmarks against two national HPC systems, Expanse predicted runtime within 10% and memory within 5% median error versus ~85% and ~60% errors for frontier LLMs. Expanse is a product of Expanse.

Website
expanse.sh
Added via
web
Ownership
unclaimed

Where to go

Product details

Product kind
Developer Tool
Open source
No
cliapi
GPU resource right-sizingHPC job failure predictioncompute cost optimizationML training workload planningjob root-cause analysis
SLURMKubernetesNomad