Compute resource prediction platform for HPC and GPU clusters that right-sizes job submissions and diagnoses failures before they waste GPU time.
Expanse integrates with SLURM, Kubernetes, or Nomad via a single install command, collects telemetry, and uses pre-trained models continuously adapted to the customer's cluster to predict memory, runtime, GPU allocation, and failure risk before each job runs. The `expanse analyse` command delivers predictions as confidence distributions without changing existing submission scripts; `expanse diagnose` performs root-cause analysis on failed jobs with code-diff fix suggestions. In benchmarks against two national HPC systems, Expanse predicted runtime within 10% and memory within 5% median error versus ~85% and ~60% errors for frontier LLMs. Expanse is a product of Expanse.