Timing closure is the loop of implementing, extracting, analyzing and fixing until every timing check passes in every scenario that matters. A design is closed when setup and hold both show WNS of zero or better and TNS of zero in every active scenario, the design rule checks are clean, and signal integrity has been included. That result has to come from PrimeTime on extracted parasitics, with analysis coverage checked so you know nothing was left untested.
WNS is the worst negative slack, the single worst endpoint. TNS is the total negative slack, the sum of the negative slacks of all failing endpoints. NVE is the number of violating endpoints. Together they tell you how bad the worst path is, how much total work is left, and whether the problem is a few paths or a wide spread.
Setup slack is the required time minus the arrival time. Arrival time is when data actually reaches the capture flop, starting from the launch clock edge. Required time is the capture clock edge minus the setup time and uncertainty. Positive slack means the data arrives with time to spare; negative slack means it arrives too late.
Read a timing report top to bottom in four blocks: the header, the launch side, the capture side and the slack line. The header names the startpoint, endpoint, path group and path type. The launch side builds data arrival time and the capture side builds data required time. The last line subtracts one from the other.
A setup violation means data reaches the capture flop too late for the next clock edge. The usual causes are too much logic between flops, weak drivers, long or detoured wires, a capture clock that arrives earlier than the launch clock, derates, and crosstalk that slows the signal. Most real failures are two or three of these adding together.
A hold violation means new data reaches the capture flop too soon after the clock edge, before the flop has safely stored the old value. It is caused by short data paths, a capture clock that arrives later than the launch clock, and fast process corners. Slowing the clock does not help because the hold check compares launch and capture on the same clock edge, so the period is not part of the equation.
Setup wants data to arrive earlier and hold wants it to arrive later, so any change to a data path helps one check and costs margin on the other. Clock skew has the same effect: a later capture clock helps setup and hurts hold on the same flop. The goal is to fix each on the paths where the other has spare margin.
DRVs are design rule violations: a net whose transition time, load capacitance or fanout exceeds the limit set by the library or the constraints. They are fixed before timing because DRV fixing changes cells and buffers, which changes timing, and because timing on a violating net comes from outside the library characterized range. In PrimeTime, DRC fixing has the highest priority and can degrade setup or hold, so it runs first.
A max-transition violation is fixed by making the driver stronger or the load it sees smaller. The three standard moves are upsizing the driver, inserting a buffer to break a long net, and splitting a large fanout across several buffers. PrimeTime can do this automatically with `fix_eco_drc` (PT), and ICC2 has `size_cell` (ICC2) and `add_buffer` (ICC2) for manual fixes.
A max-capacitance violation means a driver sees more total load than its library limit. It matters beyond timing because the library was only characterized up to that load, and heavy load means higher current through the driver and its output wire, which is an electromigration and power concern. The fixes are to upsize the driver or split the load across buffers.
Upsizing replaces a cell with a stronger version of the same function, which drives its load faster and cuts delay. The cost is a larger input capacitance, which slows the stage before it, plus more area and leakage. It is the first fix PrimeTime tries for setup: by default `fix_eco_timing -type setup` (PT) uses cell sizing alone.
A Vt swap replaces a cell with the same function, size and footprint but a different threshold voltage. Lower Vt cells switch faster and leak more, higher Vt cells are slower and leak less. You swap to LVT only on paths that need the speed, and swap non-critical cells to HVT to recover leakage.
A buffer speeds up a path when the wire it breaks is long enough that its RC delay is larger than the buffer delay added. Wire delay grows roughly with the square of length, because both resistance and capacitance grow with length. Splitting the wire into shorter segments makes total delay grow roughly linearly. On a short wire, a buffer only adds delay.
Hold violations are fixed by adding delay to the data path so new data arrives after the hold window closes. The delay cell goes near the capture flop, on the branch that only the hold-failing endpoint uses. That keeps the delay off shared logic that may feed setup-critical endpoints. Always check the setup margin of the fixed path at the slow corner.
Useful skew is clock skew added on purpose to fix timing. Delaying the capture clock of a failing flop gives its incoming path more time, and takes the same amount from the path leaving that flop. It works when the next stage has spare slack to lend. ICC2 does this automatically with concurrent clock and data optimization.
Path groups split a block's timing into buckets so one bad bucket cannot hide the others. By default each clock gets a group, and you add your own with `group_path` (SDC), commonly reg2reg, in2reg, reg2out and in2out. Read WNS, TNS and the violating-path count per group, because each group points at a different owner and fix.
The worst path is only the first line of a long list. Behind it there are usually many paths within a few picoseconds, and they often share cells with it, so fixing one path leaves the rest failing or moves the problem to the next endpoint. You need the slack distribution and a report that covers every violating endpoint, not only the top path.
An ECO, or engineering change order, is a small, controlled change to a design that is already placed and routed, made instead of running the flow again. You edit the netlist, place only the changed cells, reroute only the affected nets and re-check. Timing ECOs fix timing or design rule violations, while functional ECOs change the logic itself.
A timing ECO changes how fast the logic is without changing what it computes: sizing, Vt swaps, buffers and hold delay cells. A functional ECO changes what the logic computes, usually to match a corrected RTL. That difference decides where the change comes from, what equivalence is checked against and how much risk the change carries.
The ECO loop splits the job between the tool that signs off timing and the tool that owns the layout. PrimeTime finds and fixes violations with signoff parasitics and writes the edits with `write_changes` (PT); ICC2 sources that script, places the new or resized cells and reroutes the touched nets. After fresh extraction PrimeTime re-times the design, and the loop repeats until nothing is left to fix.
ICC2 is an optimizer: its timer runs inside placement, CTS and routing loops and is tuned for speed because optimization calls it over and over. PrimeTime is the dedicated analysis engine that the team agrees to trust, run on signoff parasitics with crosstalk, path-based recalculation and every scenario at once. Signing off in PT means one reference engine and one data set decide whether the chip meets timing.
Graph-based analysis times every node once and keeps the worst arrival and the worst slew at each pin, even when they come from different inputs, so it is fast and pessimistic. Path-based analysis takes one path, drops the side inputs and recalculates the delays with the slews that belong to that path, so it is more accurate and slower. You find violations with GBA and use PBA to see how much of the remaining violation is real.
A timing derate is a multiplier applied to calculated delays to cover on-chip variation: late delays are scaled up and early delays scaled down. In a PrimeTime report you only see it if you ask: `report_timing -derate` (PT) adds a Derate column next to each incremental delay. Reading that column tells you which factor was applied to which cell or net, and whether it is the one you meant.
CRPR appears as a line called clock reconvergence pessimism on the capture side of a PrimeTime timing report, just after the clock network delay. It adds back the pessimism created when the shared part of the launch and capture clock paths was timed late for one and early for the other. Because it only removes pessimism, it can only improve slack.
With crosstalk analysis on, PrimeTime already includes crosstalk in every path delay, but you only see it separately if you ask. `report_timing -crosstalk_delta` (PT) adds a Delta column that shows the delay change on each victim net arc caused by switching neighbours. To find the nets that cause the most trouble across many paths, `report_si_bottleneck` (PT) ranks them.
Clock uncertainty is a margin you add for clock effects you do not model, such as jitter or skew not yet known. For setup, PrimeTime subtracts it from the data required time, so data must arrive earlier; for hold, it adds it to the required time, so data must stay stable longer. Every picosecond of uncertainty is a picosecond of slack taken away.
Every ECO edits the netlist, and any edit can change function by mistake: a wrong library cell, a buffer on the wrong pin, a dropped inverter or a patch that does not match the new RTL. Timing and physical checks do not look at logic, so only equivalence checking proves the post-ECO netlist still computes what it should. It is quick for small ECOs and it is the only check that catches these bugs before silicon.
`report_qor` (PT) gives the one-page state of a design, organized by path group. Read it group by group, checking worst slack, TNS and number of violating paths, then the hold and design rule summaries; in PT, `report_constraint` (PT) lists the individual design rule violators behind those counts. The PT and ICC2 versions have the same idea and different numbers, because the engines, parasitics and settings differ.
Timing is closed only when PrimeTime shows no violations in every signoff scenario and you can prove the analysis covered everything it should. That means clean constraints, no unexplained untested checks, no setup, hold or DRC violators, SI enabled, and any path saved by PBA confirmed with exhaustive analysis. A clean WNS number alone is not evidence.
`write_changes` (PT) writes every netlist change made during the PrimeTime session as a change list, in a format chosen with `-format`. For ICC2 the format is icctcl, a Tcl script of netlist edits; ICC2 sources it, then places the changed cells with `place_eco_cells -eco_changed_cells` (ICC2) and reconnects them with `route_eco` (ICC2). The file is only a list of edits, so ICC2 still has to make them physically legal.
Start with the fix that costs least and disturbs the layout least, and climb only when the cheaper rung runs out. The usual order is Vt swap, cell sizing, buffering or fanout split, logic restructuring or cloning, useful skew, and finally placement or floorplan changes. Each step up moves more cells, touches more nets and puts more already-closed timing at risk, so you stop at the first rung that clears the path.
Add delay where the data path has setup margin to spare, not simply where the hold violation shows up. Delay added at a pin helps hold on every path through that pin and costs setup on the same paths, so the right pin is the one whose worst setup path still passes after the delay goes in. For very small violations, around 5 ps or less, a load cell adds just enough delay without the overshoot of a whole buffer.
`fix_eco_timing` (PT) will not run without `-type setup` or `-type hold`, because the two problems need opposite changes. By default, setup fixing uses cell sizing alone to cut data path delay, and hold fixing uses sizing plus buffer insertion to add delay. Both work only on data paths, not clock networks. Setup fixing avoids new DRC violations but may create hold violations, while hold fixing avoids creating either setup or DRC violations.
`fix_eco_drc` (PT) fixes what `report_constraint` (PT) reports for max capacitance, max transition and max fanout, noise violations from `report_noise` (PT), crosstalk delta delay found by `report_si_bottleneck` (PT), and cell electromigration from `report_cell_em_violation` (PT). It sizes cells and inserts buffers or inverter pairs while keeping area low, and by default it ignores timing. Delta delay needs a threshold because no constraint limits it, so there is no delta delay slack to fix against.
`fix_eco_power` (PT) swaps cells for cheaper ones on paths with positive setup slack, or removes buffers on paths with positive hold slack, and backs out any change that creates or worsens a timing or DRC violation. What counts as cheaper depends on the mode: smallest area by default, a numeric library attribute with `-power_attribute`, a name or string priority list with `-pattern_priority`, or PrimePower data with `-power_mode`. You choose one mode per run and run the command again for another.
The order follows which step is allowed to hurt which. Power recovery goes first because it never creates a violation and frees room; DRC and noise come next because DRC fixing has the highest priority and may move setup and hold slack. Setup follows because it honours DRC but may spend hold margin, then hold, which honours both, and finally a leakage-only Vt swap that changes no layout.
The `-physical_mode` setting decides where PrimeTime may put a new or resized cell. occupied_site places or resizes even with no free site, so it fixes more but leaves ICC2 legalization to push neighbours; open_site uses only free sites, so it fixes less but moves almost nothing; freeze_silicon places new cells only on spare cells so the silicon layers stay the same. All three belong to the physically aware ECO flow, which requires a PrimeTime-ADV license.
Distributed multi-scenario analysis runs one PrimeTime manager and several workers, each worker analyzing one scenario, where a scenario is one combination of operating condition and mode. The manager sets up, hands out tasks and merges results, but does no timing analysis itself. ECOs are fixed across all scenarios together because a change that fixes setup in the slow corner can break hold in the fast corner or timing in test mode, and only a run that sees every scenario can refuse that change.
Merged reporting lets you run a report once at the manager and get one answer across every scenario in command focus, with duplicates removed, results sorted by slack and each line labelled with its scenario. It works with `report_timing` (PT), `get_timing_paths` (PT), `report_constraint` (PT), `report_analysis_coverage` (PT), `report_si_bottleneck` (PT), `report_clock_timing` (PT) and `report_min_pulse_width` (PT). The main exception is cover-design reporting: `report_timing -cover_design` (PT) runs at the workers but cannot be merged at the manager.
TNS depends on how an endpoint that sits in more than one path group is counted. With `timing_report_union_tns` (PT) at its default of true, each violating endpoint adds its worst slack once; with false, it adds once for every path group it violates in. Two runs with different settings, or with different path groups, report different TNS for exactly the same slacks.
`report_bottleneck` (PT) ranks leaf cells by how many violating paths pass through them, so a cell with a cost of 20,000 sits on 20,000 failing paths. Fixing that one cell can improve many endpoints in one change, which is often better than grinding on the few worst paths. It needs `timing_save_pin_arrival_and_slack` (PT) set to true before the first timing update.
Run exhaustive PBA when graph-based violations are close enough that removing pessimism could clear them, and always for final signoff numbers. Keep it affordable by scoping it: search only failing paths with `-slack_lesser_than`, check with `-pba_mode path` first, and narrow to the groups or endpoints you need. Never waive a violation on `ml_exhaustive` results, because ML-PBA can miss worse paths and extra failing endpoints.
Run `check_timing` (PT) first to fix missing clocks, missing I/O constraints and loops, then run `report_analysis_coverage` (PT) and account for every untested check. With `-status_details {untested}` the report lists each untested check and the reason, such as no_clock, constant_disabled or no_paths. You have proof when every untested check is either tested in another scenario or has a reason you deliberately created.
With `design.eco_freeze_silicon_mode` (ICC2) set to true, `size_cell` (ICC2) first looks for a compatible spare cell within five times the unit site height of the cell being sized, and if none exists it refuses to size. When it finds one, the original cell becomes a spare renamed with a _Spare suffix, and a new cell with the original name, connections and new library cell is placed over the matched spare. `-max_distance_to_spare_cell` changes the search radius and `-not_spare_cell_aware` switches the check off.
PrimeTime lets you edit its in-memory netlist, see the timing effect, and only then write the change out for ICC2. Use `estimate_eco` (PT) to rank the options quickly, commit the one you like with `size_cell` (PT) or `insert_buffer` (PT), and check it with `report_timing` (PT). The timing update stays incremental, and therefore fast, only while you stick to the edit commands PrimeTime supports for what-if analysis.
After an ECO change list is applied and legalized, run `report_eco_physical_changes` (ICC2) to see how far each sized or added cell moved and how much its nets grew. Cells that landed far from their connections usually cost more delay than they saved. Undo them with `revert_eco_changes -cells` (ICC2), which removes the added buffer or inverter pair, or restores the original cell, location and orientation.
`legalize_eco_cells` (ICC2) is a short form of careful minimum-impact legalization: you give it the ECO cells and a tolerance level, low or high, instead of exact displacement numbers. Use it when the ECO cells already have good locations and you only need them made legal with as little disturbance as possible. Hand-tuned `place_eco_cells` (ICC2) is still the tool when you need control over placement itself, exact rejection thresholds or filler removal.
Run `report_cell_feasible_space` (ICC2) on the region around your violations before PrimeTime adds any buffers. It reports free placeable sites, grouped by width, after subtracting existing cells, placement blockages and macros. If there are no gaps wide enough for the cells you plan to add, fix with sizing, use PrimeTime's open-site mode, or make room first.
By default PrimeTime ECO only touches data paths. With physical clock data enabled through `set_eco_options -physical_enable_clock_data` (PT) and `-cell_type clock_network` on `fix_eco_timing` (PT), it can size and insert buffers in the clock tree to shift arrival times. Two options limit it: `-clock_fixes_per_change` sets how many violations each change must fix, and `-clock_max_level_from_reg` sets how far from the register clock pin a change may go.
TNS-driven fixing lets `fix_eco_timing` (PT) make a clock network change that helps several endpoints even though it makes one worse, as long as total negative slack goes down and WNS stays within a limit. By default no clock change may worsen any violating endpoint, so one shared clock buffer that would fix many paths is rejected. You accept a few worse endpoints because less TNS means less total fixing left, and the WNS cap keeps the trade bounded.
Start with `report_si_bottleneck -cost_type delta_delay` (PT) to rank victim nets by the crosstalk delay they add on failing paths. Then fix them with `fix_eco_drc -type delta_delay -delta_delay_threshold <v>` (PT), which sizes drivers and inserts buffers on victims whose delta delay exceeds the threshold. A threshold is needed because there is no delta delay constraint, and so no delta delay slack to drive the fixing.
When a path has `set_multicycle_path 2 -setup` (SDC) but no matching hold exception, the hold check moves forward with the setup check and PrimeTime reports a hold violation close to one clock period. That is a constraint bug, not a silicon problem. Fix the SDC with `set_multicycle_path 1 -hold` (SDC); buffering it wastes area, spends setup margin and leaves the wrong constraint in every later run.
Before you spend an ECO on a violation, walk a short checklist. Confirm the constraints are complete, the clocks and their relationship are right, no exception is missing or ignored, the derates are the intended ones, the path still fails under exhaustive PBA, and it fails in a scenario you sign off. Only a violation that survives every step earns a fix.
Tag the LVT library cells as a threshold-voltage group, declare it low-Vt with `set_threshold_voltage_group_type -type low_vt LVT` (ICC2), and cap it with `set_multi_vth_constraint -low_vt_percentage <p>` (ICC2), by cell count or by area. `place_opt` (ICC2), `clock_opt` (ICC2) and `route_opt` (ICC2) then honour the cap on data-path cells. Set it early, because `route_opt` (ICC2) deliberately limits this optimization to avoid disturbing QoR.
Fix the RTL, produce a golden ECO netlist, and prove it matches the new RTL in Formality. In ICC2, `eco_netlist -by_verilog_file` (ICC2) compares that netlist with the layout netlist and writes the edits as Tcl, which you source, place with `place_eco_cells -eco_changed_cells` (ICC2) and route with `route_eco` (ICC2). Then prove the layout netlist again in Formality and close timing in PrimeTime.
Each changed cell carries the `eco_change_status` (ICC2) attribute, and its value says what kind of edit touched it. `place_eco_cells -eco_changed_cells` (ICC2) works on cells whose status is create_cell, change_link, add_buffer, size_cell or add_buffer_on_route. After legalization the status becomes eco_legalized, so the same cells are not picked up again.
Each corner scales cell and wire delays differently, so hold slack, which depends on the difference between data delay and clock skew, can be worst in more than one corner. A delay cell that fixes hold in the fast corner adds even more delay in the slow corner and can break setup there. Run the fix in DMSA with every signoff scenario in focus, `current_scenario -all` (PT) then `fix_eco_timing -type hold` (PT), so each change is checked against all of them.
`route_opt` (ICC2) does the bulk of postroute fixing: setup, hold, area and logical DRC on data paths, with legalization and ECO routing in the same step. PrimeTime ECO handles what is left at signoff: violations only the signoff timer sees, with full extraction, SI, exhaustive PBA and every scenario. If ICC2 can see a violation, `route_opt` (ICC2) should fix it; PrimeTime should get a short list.
An ECO buffer on a net that crosses domains can land in a voltage area whose power domain does not match, get the wrong supply, sit on the wrong side of an isolation cell, or use a single-rail cell where a dual-rail cell is needed. Repair it with `fix_mv_design -buffer` (ICC2) for buffer trees and `fix_mv_design -diode` (ICC2) for diodes, and connect the supplies of new ECO cells with `eco_update_supply_net` (ICC2). Then check again, since the fixes are not guaranteed placement or timing legal.
Do not start fixing. Five thousand violations on a first routed run is a diagnosis problem: first prove the constraints and the analysis setup are right, then group the violations until they collapse into a handful of causes. Only after that do you decide whether each cluster is ECO work in PrimeTime or a reason to go back to placement or floorplan.
The loop exists because each fixing step is allowed to hurt something else. PrimeTime defines precedence rules for exactly this reason: DRC fixing can degrade setup and hold, setup fixing honors DRC but may degrade hold, and hold fixing honors both setup and DRC. Run the steps in that order, give each step margins, and find the paths where no legal answer exists instead of looping over them.
Chase TNS when there are many shallow violations and the aim is to reduce the amount of fixing work left. Chase WNS when a few deep paths set the frequency or block signoff. Early in closure TNS tells you how much work remains; at signoff every violating endpoint must be fixed, so the order is only a question of which method moves the design fastest.
Set `eco_enable_more_scenarios_than_hosts` (PT) to true so DMSA ECO can run with fewer hosts than scenarios. The hosts then swap scenarios in and out, which is slow, so the script has to minimise swaps: fewer merged reports, one batched `remote_execute` (PT) block, and change lists written from one scenario. Also check whether 40 scenarios are all needed.
HyperTrace accelerates path-based analysis, and PrimeTime can use it to speed up ECO fixing when `fix_eco_timing` (PT) runs with a PBA mode. It changes runtime, not the fixing goal: the same PBA slack is targeted, reached faster. It needs a PrimeECO license, and it pays off little when fixing is graph-based or when the PBA work is small to begin with.
In a freeze silicon ECO only metal and via layers may change, so every new cell must land on a spare cell that is already on the silicon. PrimeTime generates fixes restricted to spare cells with `-physical_mode freeze_silicon`, and ICC2 implements them in freeze silicon mode, checks feasibility, maps each ECO cell to a spare cell and reroutes only what changed. Then you extract and sign off again.
By default ICC2 maps an ECO cell only to a spare with the same library cell name, so a type with no matching spare stays unmapped. Find those cells with `check_freeze_silicon` (ICC2), then let `create_freeze_silicon_leq_change_list` (ICC2) write a script that replaces each one with a logically equivalent spare or a combination of up to two spares. Review the script, source it, and map again.
A clock change moves an edge for every register below it, so one fix can shift slack on many paths at once. Delaying a capture clock helps setup into that register but hurts hold into it and setup out of it, and the same change lands in every scenario and can alter CRPR. PrimeTime can do clock network fixing, but it should come after data-path fixing and with limits on how far up the tree it reaches.
Fix hold once, in a DMSA session that contains all 12 scenarios, so every change is judged against setup in the tight corners and hold in the failing ones at the same time. Ping-pong comes from fixing the hold corners in one run and checking setup in another. Where no data-path change satisfies both, the endpoint is a real conflict and needs a clock or constraint answer.
Crosstalk depends on the neighbours and on timing windows, and an ECO can change both without touching the victim. ECO routes fill gaps next to existing nets, upsized drivers switch faster, and shifted arrival times make aggressor and victim windows overlap. Find the victims with `report_si_bottleneck` (PT), fix delta delay with `fix_eco_drc -type delta_delay` (PT), and implement with `route_eco` (ICC2).
Standard corner timing assumes every cell sees the full rail voltage. With IR drop, cells in hot spots see less, switch slower, and paths through them lose slack that no corner shows. Voltage-aware timing brings the drop into STA so those paths are fixed where they are, and in ICC2 the power integrity flow reduces the drop during placement and CTS so fewer paths need fixing later.
Yes, a GBA violation can be waived when exhaustive path-based analysis proves the endpoint passes, because GBA is deliberately pessimistic. The evidence must be regular exhaustive PBA to that endpoint, in every scenario where it failed, with no path limit cutting the search. Path-mode PBA, a single recalculated path, or ML-PBA do not qualify.
Stop when the fixes you need are not the kind ECO can make, or when each iteration stops paying for itself. PrimeTime ECO sizes, swaps and buffers existing logic in the space that exists. When the problem is logic depth, wire length set by placement, congestion, or no room left for changes, you go back to placement or floorplan. The decision is cheaper the earlier it is made, so track the evidence from the first ECO round.
A block ECO changes the block interface timing the top sees, so every abstraction of that block used at the top is now out of date. Regenerate the block model, whether an extracted timing model from `extract_model` (PT) or a HyperScale block model, and rerun top-level timing. If the top changes as a result, regenerate the block context with `characterize_context` (PT) and `write_context` (PT) so the block is fixed against the real environment.
Mode merging creates a superset mode so fewer scenarios need analysis, but the merged constraints are only as good as the merge. If a clock, case value or exception from one mode masks a path that another mode checks, the merged scenario never tests it. Catch it by reviewing the merge report, comparing analysis coverage against the original modes, and running the individual modes at defined checkpoints.
When data reaches a transparent latch after its opening edge, the path borrows time and the next stage starts later by the same amount. Closure looks easy on the first stage and the deficit shows up downstream, possibly several stages later. Capping borrowing with `set_max_time_borrow` (SDC) keeps each stage honest and leaves margin in the pulse width, which you want in signoff.
Normalized slack divides a path's slack by the propagation delay it is allowed, so paths with different cycle counts can be compared on their effect on frequency. With the same launch and capture clock, the worst normalized slack gives the period change directly: ΔPeriod = −(worst normalized slack) × period. Enable it with `set_app_var timing_enable_normalized_slack true` (PT) before the timing update.
PrimeTime uniquifies the edited instance automatically. If you `size_cell` (PT) a cell inside one of several instances of the same block, that instance, and every block above it up to the first singly instanced block, becomes a new unique design. The timing is right for what you did, but the physical side now has two different versions of a block that was built once, so decide first whether the change belongs to every instance.
After the last ECO, nothing incremental counts. The gate is a full extraction of the final layout, a full PrimeTime run in every scenario with SI and exhaustive PBA, a coverage and constraint check, equivalence checking of the final netlist, and clean physical checks. Every one must pass on the same final database, or the block does not go.
Buffering each bit of a bus at the same interval puts all the buffers in one column, where they fight for sites and pin access and push cells aside. ICC2 has bus buffer patterns for this: define a stagger with `create_eco_bus_buffer_pattern` (ICC2), then place the buffers with `add_buffer_on_route -user_specified_bus_buffers` (ICC2), giving the pattern, the library cell and the first buffer location.