Before CTS, STA times against ideal clocks -- user-specified set_clock_latency and set_clock_uncertainty values standing in for a tree that doesn't exist yet. After CTS, the tool switches to propagated clocks: real, computed insertion delay and real skew from the actual built tree, not an estimate. This is why a design that looked clean pre-CTS can show new setup or hold violations post-CTS -- it's not that the design got worse, it's that STA is finally measuring the real thing instead of a placeholder.
For a normal post-CTS hold fix, add_buffer_on_route inserts a buffer directly on the existing route topology, or size_cell resizes an existing cell -- both are logic/metal-only changes to an already-placed, already-routed design. Freeze-silicon ECO is a different tier: it maps the fix onto pre-placed programmable spare cells (matched by psc_type_id) so the change requires only a metal-mask respin, not a full new mask set -- used specifically when the design has already taped out or is close enough that touching silicon layers is unacceptable.
STA runs several distinct check types at sequential boundaries: setup and hold (data stability before/after the clock edge), removal (minimum time between a clock edge and releasing an async control signal like reset), recovery (minimum time between releasing that signal and the next clock edge), period (minimum time for one full clock cycle), and minimum pulse width high/low (MPH/MPL). Each protects a different failure mode -- removal and recovery specifically guard asynchronous reset/set behavior, not the synchronous data path.
MPH (Minimum Pulse Width High) verifies a clock's positive pulse stays high long enough; MPL (Minimum Pulse Width Low) verifies the negative pulse stays low long enough. The same underlying MPL-style logic is also used for transparent-latch setup/hold slack adjustments -- a latch's transparent window has to stay open long enough for data to actually pass through it, which is structurally the same "minimum duration" problem as a clock pulse.
Violating setup or hold can drive a flip-flop into a metastable state, where the output settles to an intermediate value and may never resolve to a clean 0 or 1 in time. The standard mitigation is double synchronization -- passing the signal through two or more flip-flops in series, giving the metastable state extra clock cycles to resolve before it's used by downstream logic.
The four path types are input-to-register (I2R), register-to-register (R2R), register-to-output (R2O), and input-to-output (I2O) -- the standard 4-way STA path group classification. This classification matters because different path types often need different constraint treatment (I/O delays for boundary paths, clock relationships for R2R paths) and different optimization priority during timing closure.
A zero wire-load model means pre-layout, zero net delays -- cell propagation delays only, useful early but unrealistic. A real wire-load model estimates loading/delay effect of fanout from the technology library (an area-based default) -- better than zero, still an estimate. Post-layout extracted-parasitics STA, using actual LEF/routing data, is far more accurate and meaningfully reduces timing closure surprises later in the flow.
The rule of thumb: max fanout is typically 10 -- any net should drive a load equivalent to no more than 10 input cells. This is used to map the correct drive-strength cell against the stated fanout -- a driver sized for 10 loads won't perform correctly (transition, delay) if the design actually connects 20 loads to it.
No -- this is a specific, important exception: max fanout DRC is NOT honored by ICC2 as a hard design rule constraint. Instead, opt.common.max_fanout is a soft optimization constraint. Min capacitance, max capacitance, and max transition DO have technology-specific defaults in the logic libraries and can be overridden as real DRC -- max fanout is handled differently.
The tool fixes the Worst Negative Slack (WNS) path first; once that path meets timing it moves to the next-worst, stopping when it hits a path it can't fix. To avoid the tool exhausting all effort chasing one dominant path group, paths are grouped into cost groups -- sets of critical paths with an assigned priority/weight -- so effort gets balanced across groups instead of consumed entirely by one group's single worst path.
Yes -- by default STA assumes single-cycle for ALL paths. A path that's actually allowed multiple cycles to reach its capture flip-flop (e.g. a configuration-register-driven enable stable across several clocks) still gets checked as if it needed to close in one cycle unless explicitly declared with set_multicycle_path. Without that declaration, the tool over-constrains the path and misreports it as violating, when it actually has legitimate extra cycles to work with.
For setup checks: worst PVT for the clock and data (launch) paths, best PVT for the reference clock (capture) path. For hold checks: best PVT for the clock and data paths, worst PVT for the reference clock (capture) path -- the launch/capture PVT assignment flips entirely between the two check types, which is exactly what makes setup and hold pull a design in opposite directions.
Flat OCV derating applies one derate factor to every cell/net delay in the early direction (hold) and one in the late direction (setup), regardless of path depth or distance -- simple but pessimistic, since it assumes worst-case variation compounds identically everywhere. AOCV (Advanced OCV) instead makes the derate factor a function of logic depth and/or physical distance, since variation statistically partially averages out over more stages/more distance -- less pessimistic, but it requires real AOCV characterization data (depth/distance-vs-derate tables) the library has to actually supply.
SoCs operate in multiple modes (active, sleep, test) sharing the same logic, each needing its own constraints -- sleep mode may use a different supply voltage or clock frequency, for instance. Fixing timing in one mode can genuinely reopen violations in another, because the same physical cells and paths are shared across modes -- a fix that helps mode A's timing can change delay in a way that hurts mode B's, since they're not independent designs, they're the same design analyzed under different constraint sets.
For a violating R2R path, the two canonical fixes are: swap to faster cells along the path (if faster library variants exist), or split the path by inserting a register partway through, creating a new, shorter R2R endpoint. Either fix requires re-running LEC between the modified netlist and the golden RTL to confirm functional equivalence -- this same swap-vs-split methodology is the core approach used in post-PD ECO timing closure generally, not just this one worked example.
route_opt is the standard post-route optimization command -- up/down-sizing drive strength along failing paths without disturbing other placement. hyper_route_opt is a distinct, more intensive variant. Targeted Endpoint Optimization (set_route_opt_target_endpoints) lets you work on a specific SUBSET of endpoints rather than optimizing the whole design -- useful when only a known set of endpoints are actually violating and you want to avoid disturbing everything else's already-closed timing.
add_spare_cells takes -cell_name (naming prefix) plus either -lib_cell with -num_instances (same count per cell type) or -num_cells {CELL count ...} (different counts per type). -repetitive_window {w h} repeats a window of spare cells throughout the placement area at that pitch. Distribution control includes -hier_cell, -boundary, -voltage_areas, -random_distribution (ignore cell density), and -density_aware_ratio (default 100% density-based placement).
Programmable spare cells (gate array filler cells) are matched to real standard cells via a shared psc_type_id attribute set on both the fill library cell and the standard cell library cells it can represent -- set_attribute [get_lib_cells ...] psc_type_id <N> on both sides. During the ECO flow, the tool swaps an ECO cell with a programmable spare cell based on matching psc_type_id, cell width, and voltage area -- all three have to match, not just the type id alone.
add_buffer inserts a buffer on a net by specifying the net and library cell directly (-object_list, -lib_cell), letting the tool determine placement -- appropriate for a net that isn't already routed, or where preserving exact existing routing topology doesn't matter. add_buffer_on_route (ABOR) specifically adds buffers based on the EXISTING routing topology, minimizing disturbance to routes that are already in place -- the right choice post-route specifically, when you don't want to risk re-routing everything.
split_fanout optimizes net fanout during a timing ECO by inserting buffers to break up a high-fanout net -- specified either by -net/-driver plus -max_fanout (a numeric limit) or -load (an explicit list of pins/ports, optionally -hierarchy-aware). -on_route adds the buffers on the existing route topology (like ABOR); -max_distance_for_incomplete_route handles splitting fanout of a net that isn't fully routed, within a max distance, and errors if 0 or more than 1 routed segment is found in that distance.
create_virtual_connection -pins {pin1 pin2} defines a connectivity relationship for placement guidance purposes only, with an optional -weight to control how strongly it pulls cells together -- it does NOT create a real electrical connection. This is useful when you want an ECO cell placed near a specific pin for a reason (e.g. anticipated future connectivity, or minimizing eventual routing distance) without actually wiring them together yet.
A gate is sensitized if a transition can propagate through it from a particular input to the output while other inputs hold a NON-controlling value (logic 1 at an AND input, logic 0 at an OR input) -- a controlling value (logic 0 at AND, logic 1 at OR) would force the output regardless, blocking propagation. Sensitization criteria are static (a set of input vectors exists giving non-controlling values at all side inputs along the path) or dynamic (vectors applied at different TIMES produce non-controlling values when the propagating transition actually arrives -- considered very complex and still an active research area).
For an interconnect with N nodes, Elmore delay at node i is D_i = sum over k=1..N of (R_ki * C_k), where R_ki is the resistance of the path segment common to the input-to-node-i and input-to-node-k paths, and C_k is the capacitance at node k. It's popular for its algebraic simplicity and is accurate for nodes FAR from the driving point -- but can be off by orders of magnitude for nodes NEAR the driving point, because of resistive shielding: when wire resistance is comparable to or larger than the driver's output resistance, the metal resistance shields the wire's capacitance from the driver, and Elmore's simple summation doesn't capture that shielding effect.
Given driver resistance R_d, wire resistance R_w, and loads C1/C2, wire delay is D_w = R_d*(C1+C2) + R_w*C2. If R_d >> R_w, driver delay is accurately a function of total load (C1+C2). If R_w is comparable to or exceeds R_d, driver delay decreases and part of C2 is shielded, so delay is characterized as a function of (C1 + k*C2), where k (the effective capacitance factor) ranges 0 to 1. Because k itself depends on driver resistance -- which depends on the delay being computed -- it has to be computed iteratively, independent of any specific driver model, rather than solved in one closed-form step.
AWE (Asymptotic Waveform Evaluation) constructs a pole-residue transfer function H(s) = sum(i=1..q) of k_i/(s - p_i), with time-domain impulse response h(t) = sum of k_i * e^(p_i*t) -- using higher-order moments than Elmore's single first moment. AWE matches the first 2q moments of the network's transfer function to h(t)'s moments to uniquely specify the poles and residues. q=2 or 3 is typically sufficient for reasonable accuracy at reasonable computational cost -- going higher captures more of the waveform's true shape but with rapidly diminishing returns against a real computational cost increase.
ANDR (Auto Non-Default Rules) are non-default rules the tool creates AUTOMATICALLY during preroute and global-route optimization for timing/power -- not user-specified NDRs, tool-generated ones. ERI (enable_runtime_improvements) is a runtime-focused post-route optimization enhancement. EWI (enable_wireopt_improvements) specifically enables ANDR wire optimization and layer promotion, mostly during global-route-based optimization in clock_opt -- meaning EWI and ANDR are directly linked (EWI is what turns on ANDR's wire-optimization behavior), while ERI is a separate, runtime-oriented improvement.
eco_netlist compares the current design against a golden reference -- either a golden Verilog netlist (-by_verilog_file) or a golden block (-block, assumed in the current design library unless -golden_lib is given). -write_changes <file> is REQUIRED, not optional, and produces Tcl netlist-editing commands representing the diff. By default it ignores differences in physical-only cells, timing ECO changes, and power/ground objects -- -compare_physical_only_cells and -extract_timing_eco_changes opt into including those.
record_layout_editing -start begins recording, and every layout-editing operation performed afterward (create/remove of layout objects via Tcl, set_attribute, and GUI move/resize) gets captured. record_layout_editing -stop -output <file> writes those recorded operations as a Tcl script. Multiple engineers can each record their own independent session this way, producing separate Tcl files that can ALL be applied to the original, unmodified layout -- rather than each engineer's changes needing to be applied sequentially to an already-modified copy.
free_site_only (the default) legalizes ECO cells only on free sites without moving pre-existing cells -- can cause large displacement if no free site is nearby, since the cell may have to travel far to find one. allow_move_other_cells legalizes to the nearest legal location by moving pre-existing cells instead -- no free-site search, but disturbs the existing layout. minimum_physical_impact tries free sites first, and only moves pre-existing cells for cells that still have no nearby free site -- a hybrid that minimizes disturbance while still bounding displacement.
-check_prerequisites validates a real checklist before placement is attempted: site rows present, each net has one driver and can have multiple loads, pin/port locations assigned, standard cells carrying the pins are placed, supernets routed, route ends within 5 microns of driver/load pins in the search region, and more. Real error examples: ECO-329 ("the net net_wo_load does not have loads") and ECO-331 ("the cell driver1 of the pin driver1/X is not placed") -- both point to a genuinely missing prerequisite, not a placement algorithm failure.
place_group_repeaters offers several mutually exclusive location-specification methods: -repeater_distance (distance between groups, with -first_distance and -min_distance_repeater_to_load defaulting to half the repeater distance), -relative_distance (driver-to-first, first-to-second, etc.), -location (exact center of each group), -number_of_repeater_groups (total count inserted evenly along the route), and -cutlines (location pairs defining cutlines perpendicular to the route, with groups centered around them). Each represents a genuinely different way of thinking about placement -- distance-based, count-based, or geometry-based.
Variant cells differ only in mask color (or small geometry) from a master variant, with identical timing data. In track_pattern mode, the variant type id is derived from the library cell's first track offset plus pitch value, starting from 0 (e.g. type 0 = master/mask_one at 1/2 site width, type 1 = mask_two at 1/6 site width). In siteId_based mode, mapping info is defined explicitly in the cell group -- with two sub-flows: 2-Variants (traditional, max 2 master/variant per group, only a mapping id is set, flipped id derived from mapping id plus cell width) and N-Variants (both mapping id and flipped mapping id are set explicitly, supporting more than 2 variants per group).