IntermediateQuestion 91 of 112Source: Synopsys PrimeTime User Guide: HyperScale Distributed Analysis

What is HyperScale, and why use it instead of one flat run for a large chip?

From PDVerse STA Mentor Guide ยท pdVerse Mentor Guide

Short Answer

HyperScale is a PrimeTime capability for analyzing a very large design as a set of smaller blocks with lightweight timing models standing in for each block's internals, distributed across multiple machines, instead of loading the entire flattened netlist into one process. It exists because a full-chip flat run on a large SoC can outgrow the memory and runtime budget of a single machine long before it outgrows the design itself.

Technical Reference DiagramWhat is HyperScale, and why use it instead of one flat run for a large chip?
A large SoC split into a dozen blocks, each replaced at the top level by a small lightweight boundary-model box instead of its full internal netlist, with the blocks distributed across separate compute-farm machines that report back to a merged top-level result.

Technical Explanation

A flat, full-chip run is conceptually the simplest approach โ€” read in everything, time everything together โ€” but it does not scale forever.

  • A flat run loads every gate and every net into one process. For a design with hundreds of millions of instances, that means one machine needs enough memory to hold the whole netlist, all its parasitics, and every scenario's library data simultaneously.
  • HyperScale splits the design into blocks with lightweight boundary models. Each block is analyzed using a compact representation of its neighbors' timing behavior at the boundary, rather than the neighbors' full internal netlists, which is a similar idea to an extracted timing model but built into a distributed analysis flow.
  • Work is distributed across multiple machines. Instead of one process holding everything, separate blocks (or groups of scenarios) are analyzed on separate hosts and the results are merged, which is why HyperScale is described as a distributed analysis approach rather than just a modeling trick.
  • It still produces full-chip-accurate results for the paths that matter, because the boundary models are built to preserve the timing behavior a neighboring block actually sees, not a rough approximation โ€” the tradeoff is in how the work is organized and distributed, not in accuracy thrown away carelessly.
  • Both a top-down and a bottom-up usage flow exist. A top-down flow partitions the full design; a bottom-up flow builds and verifies each block's model independently first, and a team picks based on how their blocks are owned and scheduled.
  • Why it matters: without this, teams either wait too long for flat signoff turnaround, or under-verify by skipping scenarios just to fit available compute โ€” HyperScale removes that tradeoff.

Common Mistake

The Trap: assuming HyperScale is only a performance optimization that trades away accuracy for speed, the way a rough estimate would.

  • A designer treats HyperScale results as a quick, approximate pre-check and insists on a full flat run for the actual signoff decision, doubling the compute cost for no real accuracy gain on the paths HyperScale's boundary models already represent faithfully.
  • The more common real risk is the opposite mistake: building a block's boundary model once and never re-verifying it after later changes to that block's internals, letting a stale model quietly diverge from the block's true boundary timing.

Follow-up Question & Model Response

If one block's internal design changes significantly after its boundary model was built, what has to happen before top-level HyperScale results can be trusted again?

Candidate Model Response: That block's model has to be rebuilt from its updated netlist before the top-level analysis is meaningful again, since the boundary model is a snapshot of the block's timing at the time it was generated, not a live link to its current state. A stale model risks false confidence if the change made the block slower, or wasted margin if it made it faster. I would tie model regeneration to the same change-tracking discipline used for any other timing-relevant constraint file.

Practical Example

A 200-million-instance SoC is partitioned into 12 blocks for a HyperScale bottom-up flow, each analyzed across a compute farm rather than in one flat process. Full-chip flat signoff previously took 30 hours on a single high-memory machine and required cutting the active scenario set from 18 to 9 just to fit in memory. The HyperScale flow completes the full 18-scenario signoff in under 6 hours, and each block owner re-verifies their own boundary model whenever their block's netlist changes.

Complete STA Handbook

Get the complete 10-chapter STA handbook covering setup/hold margins, clock modeling, OCV/POCV, crosstalk noise, and PrimeTime closure.

VLSI Physical Design Planning Handbook โ€” fourteen chaptersDesign PlanningFourteen chapters, floorplanning through timing budgets. โ†’