Place & Route (PnR) for VLSI Physical Design Mentor Guide, 60 questions

Clock Tree Synthesis: PnR Interview Questions and Answers

Sum vs Pi topology, buffer insertion RC delay, skew/latency/jitter/uncertainty, CPP/CPPR, clock_opt, CCD, and ICG merging.

Intermediate Clock Tree Synthesis #97

What's the structural difference between a Sum and a Pi clock tree, and why do most modern designs default to Sum?

A Sum tree's total buffer count is the sum of buffers per level (n_level0 + n_level1 + ... ), producing an unbalanced tree where skew is minimized by delay matching along each path -- which makes it more process-corner-dependent. A Pi tree's total buffer count is a product across levels (n_level0 x n_level1 + n_level1 x n_level2 + ...), producing a balanced, symmetrical tree where skew depends on process uniformity rather than corner. Pi uses more buffers for the same structure. With dynamic power now a first-order concern, Sum configurations are used far more often, regardless of clock domain count, simply because they need fewer total buffers.

Intermediate Clock Tree Synthesis #98

Why does inserting buffers along a long clock wire actually reduce delay instead of just adding more delay?

A long wire's propagation delay is dominated by RC and grows with the square of length: t = rcL^2/2. Splitting that wire into N equal segments with a buffer between each one reduces the wire-delay term quadratically -- t = rc(L/N)^2 + (N-1)*t_b -- so even though you're adding N-1 buffer delays, the L^2-to-(L/N)^2 reduction more than pays for it once N is large enough. Setting the derivative to zero gives the optimal buffer count: N = L*sqrt(rc/t_b).

Beginner Clock Tree Synthesis #72

How do clock skew, latency, jitter, and uncertainty actually differ, and how do they combine?

Skew is a spatial difference: the gap in clock arrival time at two different leaf registers at the same moment. Jitter is a temporal difference: how much a clock edge's timing varies at one single point across cycles. Latency is how long the clock takes to arrive at all, split into source latency (before the clock enters the network) and network latency (through the distribution tree to the sink). Uncertainty is the combination CTS actually has to budget for -- skew plus jitter together -- because both eat into the same setup/hold margin even though they come from different physical causes.

Expert Clock Tree Synthesis #62

What is Common Path Pessimism, and how does CPPR remove it?

A good CTS algorithm branches clock paths as late as possible -- close to the leaf cells, not near the source -- so that under on-chip variation, delay differences along each path stay local and the shared (common) portion of the path doesn't contribute to skew. But that same shared common path then gets analyzed twice with opposite variation assumptions (worst-case for the launch side, best-case for the capture side), which introduces artificial pessimism into the reported slack -- that's Common Path Pessimism (CPP). CPPR (also called CRPR, Clock Reconvergence Pessimism Removal) is the analysis step that identifies the actual shared common path and removes that double-counted pessimism from the timing report.

Intermediate Clock Tree Synthesis #99

What does clock_opt actually do, stage by stage, in ICC2?

clock_opt runs CTS end to end through named stages you can target with -from/-to: build_clock builds the tree structure, route_clock routes the built tree, and final_opto is a post-build optimization pass. `clock_opt` alone runs the full flow; `clock_opt -from build_clock -to route_clock` builds and routes without the final optimization pass; `clock_opt -from route_clock -to route_clock` only routes clocks that are already built; `clock_opt -from final_opto` runs only post-build optimization on an already-built, already-routed tree.

Expert Clock Tree Synthesis #63

What is CCD (Concurrent Clock and Data optimization), and what does it actually change in the clock tree?

CCD applies useful-skew techniques during datapath optimization -- deliberately adjusting individual registers' clock arrival times to spend positive slack on paths that need it, instead of treating every register's clock latency as fixed. It's enabled by default for place_opt and clock_opt, but must be explicitly enabled for route_opt. The adjustments are stored as offsets -- set_clock_latency -offset and set_clock_balance_point -offset -- visible via write_script but NOT captured by write_sdc, which matters if you're trying to hand off timing intent downstream.

Expert Clock Tree Synthesis #64

How does CTS merge clock-gating cells, and when should you disable it?

By default, CTS merges ICGs (integrated clock-gating cells) only within the same clock tree level -- cts.icg.merge_cross_level is false by default, meaning merging across levels is an explicit opt-in, not automatic. Separately, place_opt.flow.merge_clock_gates controls whether ICG merging happens at all during the placement-stage flow (merge_clock_gates must run before create_placement). Disabling merging matters when two ICGs that look mergeable actually gate logic with different enable timing requirements -- merging them would force one clock structure to serve two enable conditions that shouldn't share one.

Beginner Clock Tree Synthesis #73

How does CTS group leaf registers before it starts inserting buffers?

CTS first identifies leaf sink points (non-clock pins of standard cells not defined as clock ports) and groups nearby ones into a virtual cluster; leaf cells far from any cluster get moved to the nearest one. Once clusters and locations are set, buffer insertion proceeds so propagation delay is equal to each cluster and skew inside each cluster is minimized. Smaller clusters mean less skew but more buffering levels, which raises total insertion delay -- a real tradeoff, not a free choice.

Beginner Clock Tree Synthesis #74

What design rule constraints does CTS actually have to meet, separate from timing?

CTS has its own DRC set, independent of setup/hold: max transition, a skew requirement, max capacitance, and max fanout. These aren't optional nice-to-haves -- a clock tree that meets setup and hold perfectly but violates max transition on a clock net is still a broken tree, because electrical DRC violations are checked and reported separately from timing violations.

Beginner Clock Tree Synthesis #75

Why does clock distribution get so much attention in power closure specifically?

The clock distribution network accounts for 30% or more of total dynamic power in modern ASICs -- some sources put it at 30-40% of total chip power -- because clock nets switch at the highest frequency in the design and carry considerable capacitive loading. An optimum clock tree matters for both performance and power, and clock gating substantially reduces this cost.

Beginner Clock Tree Synthesis #76

What is a multisource clock tree, and why would a design need one instead of a traditional CTS-built tree?

A multisource clock tree (MSCTS) is a custom clock structure with more on-chip-variation tolerance and better cross-corner performance than a traditional tree. It consists of a global clock structure (root, global tree -- usually an H-tree -- mesh drivers, and the clock mesh) plus local subtrees driven either by predefined tap drivers (a regular MSCTS) or directly from multiple mesh points (a structural MSCTS).

Beginner Clock Tree Synthesis #77

What does report_clock_qor actually show you, and what do its -type variants each report?

report_clock_qor is the summary command for clock tree quality: latency, skew, DRC violations, area, and buffer count in its default form. The -type option narrows it: -type latency for longest/shortest path, -type drc_violators for max transition/capacitance violators, -type robustness for cross-corner ratio, -type balance_groups for interclock balance QoR, -type local_skew for worst local skew per group, and -type power for per-clock power breakdown.

Beginner Clock Tree Synthesis #78

Why does CTS have to wait until after placement instead of running earlier in the flow?

CTS needs real physical locations for every leaf register to build a tree that actually balances delay to each one -- before placement, those locations don't exist yet. Running CTS on an unplaced or partially-placed design would mean building a tree around positions that are about to change, making the whole tree invalid the moment placement moves anything.

Beginner Clock Tree Synthesis #79

How are clock tree DRC constraints different from clock tree timing constraints, and why does CTS need both?

Timing constraints (setup/hold via skew and latency) ask whether the clock arrives at the right time at every register. DRC constraints (max transition, max capacitance, max fanout) ask whether every individual clock net is electrically sound, independent of when it arrives. A tree can be timing-clean and DRC-broken, or DRC-clean and timing-broken -- they're genuinely separate axes CTS has to satisfy simultaneously.

Beginner Clock Tree Synthesis #80

When does CTS actually split a clock cell, and how can you tell a split cell apart from the original?

split_clock_cells splits a clock cell to fix a DRC violation by dividing its load between two cells. A cell is NOT split if it has no DRC violations, has dont_touch/size_only/fixed placement, or would require punching a new port on a frozen power domain or block boundary. New cells are named <original_cell_name>_split_<integer>, and settings/constraints are copied from the original.

Beginner Clock Tree Synthesis #81

Are "clock-gating cell" and "ICG" the same thing, or is there a real distinction?

In practice they refer to the same physical cell -- a clock-gating cell (CGC) is what gets called an integrated clock gate (ICG) once it's placed in the tree and merging/optimization options (merge_clock_gates, cts.icg.merge_cross_level, optimize_icgs) act on it. The terms aren't describing two different components; CGC is the general term and ICG is the specific instance type the flow's options operate on.

Beginner Clock Tree Synthesis #82

What is a clock mesh driver, and how is it different from a tap driver?

A mesh driver feeds the clock mesh itself -- the 2D grid of horizontal and vertical straps joined by vias at intersections -- and is inserted via create_clock_drivers with -short_outputs to tie all driver outputs together for the mesh. A tap driver is different: it's what feeds a local subtree off the mesh in a regular multisource clock tree, connected to the mesh but driving downstream logic, not the mesh grid itself.

Intermediate Clock Tree Synthesis #101

What is a boundary cell in CTS, and what are you not allowed to do to it?

A boundary cell is a fixed buffer inserted immediately after a block or module's boundary clock pin, to preserve the boundary conditions of that pin for hierarchical CTS. Two hard rules: a boundary cell cannot be moved or resized, and no cells may be inserted between a clock pin and its boundary cell -- both rules exist to keep the hierarchical boundary's timing contract intact.

Intermediate Clock Tree Synthesis #102

What does a "don't-touch subtree" constraint actually protect during CTS, and when would you use it?

The don't-touch subtree constraint selectively preserves a portion of the clock tree at a particular clock pin -- useful, for example, at a mux between two clocks feeding different downstream paths, where you want CTS free to optimize everywhere except that specific already-correct structure. It's a scoped protection, not a whole-tree freeze.

Intermediate Clock Tree Synthesis #103

What's the practical difference between positive and negative clock skew, and why would you deliberately choose negative?

Positive clock skew routes the clock in the same direction as data flow, improving performance with tighter setup margin -- most CTS algorithms default to this. Negative clock skew routes the clock opposite to data flow, virtually eliminating the setup skew requirement but tending to degrade hold time. To get negative skew deliberately, you set the clock delay/latency to a negative number, which makes leaf registers connect to lower levels of the tree (closer to the source).

Intermediate Clock Tree Synthesis #104

What does report_ccd_timing actually show you by default, and how do you dig deeper into one specific decision?

By default, report_ccd_timing reports setup/hold slack of the worst capture (D-slack) and launch (Q-slack) paths for the 5 most critical endpoint registers. -type stage shows previous/current/next stage info for a specific pin; -type chain shows the full previous/current/next chain (multiple stages back and forward); -prepone/-postpone with -pins lets you analyze the effect of shifting clock arrival at a specific endpoint before committing to it.

Intermediate Clock Tree Synthesis #105

Why would you need copy_useful_skew instead of just trusting CCD ran the same everywhere?

copy_useful_skew copies CCD offsets to new scenarios for consistent timing/QoR reporting when CTS balance-point offsets were derived in only a subset of active scenarios. It also ensures set_clock_latency constraints are set on all pins/scenarios by computing the average offset across active scenarios and scaling it back per corner -- without it, scenarios that never got their own CCD run would report inconsistent, incomparable timing.

Intermediate Clock Tree Synthesis #106

How do you actually invoke split_clock_cells, and what are the two ways to specify what gets split?

You can split by cell (-cells [get_cells U1/ICG*]) to target specific clock cells directly, or by load (-loads {list1 list2}) to specify which load groups should end up on which resulting split cell -- two different entry points into the same underlying operation, chosen based on whether you're thinking in terms of cells or in terms of load distribution.

Intermediate Clock Tree Synthesis #107

What does mark_clock_trees actually let you control beyond just "this tree is done"?

mark_clock_trees can scope to specific clocks (-clocks, default is all clocks/modes/active scenarios), mark a tree synthesized (-synthesized, the default) or clear that (-clear), apply or remove dont_touch, propagate non-default clock routing rules or clock cell spacing rules, mark clock sinks fixed, and freeze routing of clock nets. It's a multi-purpose flagging command, not just a single boolean toggle.

Intermediate Clock Tree Synthesis #108

If you need to rebuild a clock tree from scratch, what does remove_clock_trees actually preserve versus remove?

remove_clock_trees traverses root to sinks, removing buffers and inverters except dont_touch/size_only ones -- but several object types survive by design: don't-touch cells, fixed cells, generated clocks defined on a buf/inv pin, ICGs (traversal continues past them), block abstraction models, isolation cells, and level shifters. Inverters are only removed in pairs, and a three-state buffer stops traversal entirely rather than being removed.

Intermediate Clock Tree Synthesis #109

How do you constrain which metal layers a clock tree is allowed to route on, and can you be more specific than "the whole tree"?

set_clock_routing_rules -min_routing_layer M4 -max_routing_layer M7 applies to all clock trees by default, but you can scope it to specific clocks (-clocks), or to specific net types within the tree using -net_type root|sink|internal -- root nets run from the clock root to the first branch, internal nets run from there toward sinks, and sink nets connect directly to leaf sinks.

Intermediate Clock Tree Synthesis #110

What's the actual structural difference between a regular and a structural multisource clock tree?

A regular MSCTS has local subtrees driven by predefined tap drivers connected to the mesh, built with the normal synthesize_clock_trees/clock_opt flow. A structural MSCTS has subtrees driven directly from multiple mesh points, preserving a user-defined structure that's optimized by merging/splitting/sizing rather than built fresh by the standard CTS algorithm -- the choice determines whether you're letting the tool build subtrees or preserving a structure you already designed.

Intermediate Clock Tree Synthesis #111

How do a clock mesh, an H-tree, and a spine differ as clock distribution structures?

A clock mesh is a 2D grid of horizontal and vertical straps joined by vias at intersections -- the global structure in a multisource clock tree. An H-tree is the typical topology for the global tree feeding that mesh. A spine is a 1D or 2D set of straps; a 2D spine connects multiple 1D spines to multiple orthogonal stripes, with the minimum distance between different spines' stripes called the backoff.

Intermediate Clock Tree Synthesis #112

What does create_clock_drivers actually do, and what step do you always have to run afterward?

create_clock_drivers inserts clock driver cells at specified loads, using either -boxes (a grid over the core or a bounded area) or -location (exact locations) for placement, with -lib_cells specifying the driver library cells. Critically, the command does NOT legalize -- you must run legalize_placement after every create_clock_drivers call, since the inserted drivers are marked fixed/dont-touch but not automatically legalized into the design.

Intermediate Clock Tree Synthesis #113

Where does voltage optimization actually sit relative to place_opt and clock_opt, and why does order matter here?

The voltage_opt flow runs place_opt and clock_opt first at the original voltage, then defines a scaling library group and a voltage range/target (set_vopt_range, set_vopt_target) before running voltage_opt itself -- routing and postroute optimization then happen at the newly optimized voltage. Running voltage optimization before placement/CTS are stable would mean optimizing against a design that's still going to change, wasting the analysis.

Intermediate Clock Tree Synthesis #114

In report_clock_power, what's the actual difference between a "segment" and a "subtree"?

A segment is the clock buffer tree from one ICG cell to the next ICG cells or sinks in its fanout -- a local slice. A subtree is the clock tree from an ICG cell all the way to the sinks -- the full downstream structure. report_clock_power -type per_segment gives you the local slice view; -type per_subtree gives you the full downstream view from that ICG onward.

Intermediate Clock Tree Synthesis #115

Why do clock buffers specifically need equal rise and fall delay, and what's the practical fallback when that's hard to guarantee?

Buffers used to taper clock paths need equal rise and fall delay time to maintain the original duty cycle and prevent clock signal overlap from differing propagation delays -- critical for very high-speed designs. Because perfect rise/fall balance is genuinely hard to achieve in a real buffer, a common remedy is to use inverters instead of buffers for clock tapering; incorrectly selected clock buffers with unequal rise/fall cause clock pulse width degradation as the signal propagates.

Intermediate Clock Tree Synthesis #116

How does drive strength actually scale across levels of a tapered clock buffer chain?

Tapering buffers don't have to be identical size -- they can increase drive strength monotonically by a factor alpha per clock tree level: alpha^0*d, alpha^1*d, alpha^2*d, and so on. This matches drive strength to the growing load as the tree fans out level by level, rather than using one fixed buffer size everywhere regardless of how much load it's actually driving at that point.

Intermediate Clock Tree Synthesis #118

What's the practical difference between symmetric and user-controlled tap driver configuration in an H-tree?

cts.multisource.tap_selection defaults to user, meaning tap positions follow what you specify. Setting it to symmetric produces a symmetric tap configuration automatically instead -- useful when you want the tool to enforce structural symmetry across the H-tree rather than trust manually-specified tap boxes to be symmetric themselves.

Intermediate Clock Tree Synthesis #119

How do dont_touch cells, fixed cells, and boundary cells differ in how remove_clock_trees treats them?

All three are preserved by remove_clock_trees, but for different underlying reasons: a dont_touch cell is preserved because optimization is explicitly excluded from touching it; a fixed cell is preserved because its placement is locked, independent of dont_touch; a boundary cell is preserved specifically because it protects a hierarchical timing contract, and additionally cannot be moved or resized even outside the context of tree removal.

Expert Clock Tree Synthesis #66

Why is multimode clock synthesis genuinely harder than single-mode CTS, not just "more of the same work"?

Clock distribution must be balanced for both functional and scan mode simultaneously -- and this is made harder by multilevel clock gating, clock dividing, mode-switching circuits, and a scan clock all being present in the same tree. Balancing for one mode doesn't guarantee balance for another; the tree that's optimal for functional-mode skew may not be optimal for scan-shift-mode skew, and CTS has to satisfy both from one physical structure.

Expert Clock Tree Synthesis #67

What's the real difference between integrated and user-specified clock-gate latency estimation, and which one is the default?

Integrated latency estimation is the default and more accurate method -- the tool estimates and updates clock-gate latency throughout the place_opt flow automatically. User-specified latency uses set_clock_latency directly, where lat_reg is the estimated network latency to ungated registers' clock pins, and lat_cgtoreg is the estimated delay from a CGC's clock pin to the gated register's clock pin -- for CGC clock pins specifically, you use lat_reg minus lat_cgtoreg, not lat_reg alone.

Expert Clock Tree Synthesis #68

If both set_clock_gate_latency and set_clock_latency are applied to the same clock-gating cell, which one wins?

set_clock_latency has higher precedence and is not overwritten. set_clock_gate_latency specifies clock network latency BEFORE clock gates are inserted, used at compile_fusion insertion time, as a function of clock domain, gating stage, and CGC fanout -- but if set_clock_latency has also been applied to the same object, that value governs, not the gate-latency-by-stage estimate.

Expert Clock Tree Synthesis #69

What exactly counts as a "boundary path" for CCD, and what's the strongest way to protect them from useful-skew changes?

Boundary paths are paths connected to boundary registers -- the transitive fanout of input ports or fanin of output ports. ccd.optimize_boundary_timing (false) excludes them from CCD entirely; ccd.optimize_boundary_timing_upstream (false) goes further and additionally prevents the tool from changing the clock-tree fanin cone of boundary registers, a heavier restriction on CCD's scope near I/O.

Expert Clock Tree Synthesis #70

What's the actual difference between targeting CCD at the preroute stage versus the postroute stage?

At the preroute stage (place_opt/clock_opt), ccd.targeted_ccd_path_groups plus ccd.targeted_ccd_end_points_file target specific path groups and endpoints, with ccd.enable_top_wns_optimization targeting the worst 300 paths specifically. At the postroute stage (route_opt), route_opt.flow.enable_targeted_ccd_wns_optimization is a separate flag, and ccd.targeted_ccd_select_optimization_moves controls which move types are allowed (auto includes buffering; size_only/equal_or_smaller_sizing/footprint_sizing restrict to sizing only).

Expert Clock Tree Synthesis #71

How does ICD actually make CCD faster, and what's the real mechanism behind the speedup?

Pre-CTS CCD is about 10% of place_opt runtime, mostly from medium/high-effort CUS (compute useful skew) calls. ICD (place_opt.flow.enable_fast_cus_in_final_opto) reduces one global CUS call to two LOCAL CUS calls focused only on critical endpoints, in the final_opto stage -- trading exhaustive global recomputation for a narrower, targeted recomputation that still catches the paths that matter most.

Expert Clock Tree Synthesis #72

What specific failure mode does the nworst skew framework in multithreaded CTS fix, and why does it happen in the first place?

MTCTO (multithreaded CTS) minimizes global skew -- the gap between the longest and shortest clock path. If optimization stalls on the shortest path specifically, other paths remain under-optimized, because the algorithm's attention is consumed trying to fix the one extreme rather than distributing improvement across the whole distribution. cts.optimize.enable_nworst_skew_optimization addresses this by considering more than just the single worst-case pair.

Expert Clock Tree Synthesis #73

What is a clock balance group actually for, and why can't the tool balance skew between a generated clock and its master automatically?

A clock balance group is a set of clocks considered together for delay balancing, created with create_clock_balance_group and optionally explicit per-clock offset_latencies. derive_clock_balance_constraints auto-identifies clocks with interclock timing paths worse than a threshold, and balance_clock_groups actually performs the balancing. The one hard limit: the tool cannot balance skew between a generated clock and other clocks -- a generated clock's relationship to its master is already fully determined by the generation relationship itself, not something interclock balancing can independently adjust.

Expert Clock Tree Synthesis #74

How does the tool actually decide what counts as a "root net" for clock NDR purposes, and what's the edge case that can silently change that?

Root nets are single-fanout nets from the clock root to the first branch; internal nets run from that branch point toward sinks; sink nets connect to leaf sinks. -root_ndr_fanout_limit sets the transitive fanout limit for identifying root nets -- a smaller value means MORE nets get root-net constraints, raising routing congestion risk. The edge case: if an identified root net is shorter than 10 microns, the tool uses internal-net constraints instead, regardless of the fanout-based classification.

Expert Clock Tree Synthesis #75

What real problem does topological_ndr solve, and what's the actual sort order it enforces?

Without the level-based rule, post-CTS NDR application at different tree levels can be discontinuous -- an internal net at one level might get a heavier NDR than the root net feeding it, which is physically backwards. With cts.compile.topological_ndr enabled, NDRs are topologically sorted: Root NDR > Internal NDR > Sink NDR, ensuring the NDR weight actually decreases as you move from source toward leaves, matching real current/EM demand.

Expert Clock Tree Synthesis #76

What problem do via ladders actually solve on a clock tree, and where does insert_via_ladders fit in the clock_opt flow?

Via ladders address electromigration and high-current-path concerns on critical clock nets by inserting redundant/stacked via structures. The sequence matters: define via ladder rules and constraints first, enable opt.common.enable_via_ladder_insertion for HP (high-performance) plus EM ladders on critical paths, run clock_opt -to build_clock to get the tree structure in place, THEN insert_via_ladders, THEN clock_opt -from route_clock to complete routing with the ladders already placed.

Expert Clock Tree Synthesis #77

How does create_clock_straps decide whether it's building a mesh or a spine, and what does -margins actually control?

The -type option per direction decides the topology: stripe in both directions builds a mesh; one direction as user_route (with the orthogonal direction as detect for spine detection) builds a spine instead. -margins sets the margin within which a strap may move -- default 0, which means an exact position or no strap at all, a real constraint if you need flexible strap placement rather than fixed coordinates.

Expert Clock Tree Synthesis #78

What's the actual tradeoff between fishbone, comb, and sub_strap topologies when routing to clock straps?

Fishbone (the default) connects each driver pin individually to the nearest stripe with comb routing for loads onto a single finger, controlled by -fishbone_fanout, -fishbone_span/-fishbone_sub_span, -fishbone_layers. Comb routes driver and load pins directly to the nearest stripe, falling back to Steiner topology beyond the comb distance (default 2 global routing cells) -- good for many pins directly under stripes, but creates many stacked vias, a real physical DRC risk. Sub_strap adds parallel straps on intermediate layers specifically to reduce those stacked vias versus comb.

Expert Clock Tree Synthesis #79

Why does a clock mesh specifically need SPICE-level analysis instead of the normal STA delay-calculation flow, and what are the actual prerequisites?

A clock mesh has multiple drivers feeding the same net, which normal STA delay calculation isn't built to resolve accurately -- analyze_subcircuit performs transistor-level circuit simulation to back-annotate accurate timing instead. Prerequisites: a detail-routed clock mesh net, a circuit-level model per gate and transistor model per transistor, and access to a SPICE simulator (NanoSim, FineSim, or HSPICE) -- this is a heavier, more accurate analysis specifically because mesh topology genuinely needs it.

Expert Clock Tree Synthesis #80

What's the real tradeoff behind htree_explore_all_repeater_solutions, and why would you NOT just always enable it?

htree_explore_all_repeater_solutions evaluates all repeaters per node for DRC-compliant solutions and selects the least-latency top-level solution -- genuinely better results, but runtime degrades as the library grows, since it's evaluating every candidate repeater at every node rather than a fast heuristic choice. htree_single_repeater_at_node avoids placing two repeaters at the same node (avoiding routing DRC, at the cost of possibly needing more levels) -- a different, narrower tradeoff.

Expert Clock Tree Synthesis #81

What does a multisource clock sink group actually guarantee, and why would you need one beyond normal tap assignment?

Normal tap assignment (set_multisource_clock_tap_options) assigns each sink to its closest tap driver independently. A multisource clock sink group (create_multisource_clock_sink_group) overrides that independence for specific skew-critical objects, keeping them under the same tap even if that's not each individual object's closest driver -- because for a genuinely skew-critical group, being on the same tap matters more than each member individually being closest to its own nearest driver.

Expert Clock Tree Synthesis #82

What does the Buffer Adjust step in the early CCD-iSMSCTS flow actually do, and why is it needed specifically for SMSCTS-built clocks?

In the early CCD-iSMSCTS flow, CCD runs during initial_opto on iSMSCTS-synthesized clocks, generating level-friendly useful-skew offsets. The Buffer Adjust step (buffer_adjust), introduced in final_place before placement, adds or removes repeaters to physically match those useful-skew offsets -- because on an SMSCTS structure, an abstract latency offset has to be realized as actual repeater insertion/removal, not just a number applied post-hoc the way it might be on a simpler traditional tree.

Expert Clock Tree Synthesis #83

When would you actually need irregular global tree synthesis instead of a standard H-tree, and what does "irregular" specifically mean here?

Irregular global tree synthesis handles non-symmetric, unevenly inserted tap drivers -- floorplan or global-tree complexities that require manual global tree insertion rather than an automated symmetric H-tree. It constructs multilevel, DRC-clean global trees without tight OCV or common-path requirements, using multilevel buffer trees in a non-H-tree style specifically for minimum latency to all tap drivers -- usable at block level or for top-level channel distribution.

Expert Clock Tree Synthesis #84

How does -bias_to_nets actually provide shielding on a clock mesh, and why would you want that specifically for clock straps?

-bias, -bias_to_nets, and -bias_margins on create_clock_straps bias strap position relative to PG nets or specific nets -- effectively using those nets as a shield against coupling noise on the clock strap. This matters for clock straps specifically because clock nets are high-switching-activity aggressors and victims both; positioning a strap with deliberate bias relative to a PG net gives it a stable, low-impedance neighbor rather than leaving its exact position purely up to the routing grid.

Expert Clock Tree Synthesis #85

How does the ML-based global-route optimization flow actually improve preroute/postroute timing correlation for clock trees, and what has to stay constant across iterations?

For designs with poor preroute/postroute timing correlation, ML collects Features (inputs) and Labels (predicted outputs, e.g. real post-route delay) from detail routing to train a Model relating them. Iteration N trains the model (est_delay.ml_delay_gre_mode = feature, then estimate_delay -train_model after detail routing); iteration N+1 uses it (est_delay.ml_delay_gre_mode = enable). The hard requirement: the same active scenarios must be used at each step and iteration, and the model must be re-created if the design, environment, or flow changes -- a stale model silently mispredicts.

Expert Clock Tree Synthesis #86

What does report_clock_qor -type robustness actually measure, and why is -robustness_corner a required option, not optional?

The robustness metric is the ratio of a sink's latency at the reporting corner to its latency at a required -robustness_corner -- a direct cross-corner comparison, not an absolute number. -robustness_corner is required specifically because the metric is meaningless without a reference corner to compare against; a ratio needs two points, and the tool won't guess which corner should be the baseline.

Expert Clock Tree Synthesis #87

How do you build more than one global H-tree in different parts of the same floorplan, and how are the sections actually defined?

set_regular_multisource_clock_tree_options -htree_sections takes a list of section definitions, each with -section_name, -prefix, and either -tap_boundary (a bounding box) plus -tap_boxes (a symmetric grid within it) or explicit -tap_locations -- letting you build genuinely separate H-trees for genuinely separate floorplan regions instead of forcing one global tree to span an entire irregular floorplan.

Expert Clock Tree Synthesis #88

What extra steps does multi-level physical hierarchy (MLPH) clock driver insertion need beyond the normal create_clock_drivers flow?

MLPH clock driver insertion needs set_editability -blocks {...} -value true first (to allow the insertion to touch multiple physical hierarchy blocks), then cts.multisource.enable_mlph_flow set true, then create_clock_drivers as usual, then synthesize_multisource_global_clock_trees with -roots and -leaves specified plus -use_zroute_for_pin_connections. Ports are reused if they already exist, and new ports are only created if set_freeze_ports allows it -- a real constraint on hierarchical boundaries that ordinary single-block flows don't have to think about.