Placement is the step that turns an abstract netlist ("this gate connects to that gate") into real geometry โ every standard cell gets an exact (X, Y) coordinate and a legal row orientation. It's optimizing three things at once, and they pull against each other: total wirelength (measured cheaply via HPWL), timing slack under an idealized zero-skew clock, and routability โ keeping local congestion from spiking anywhere.
Production placement isn't one monolithic step โ it runs as four distinct, purpose-built phases, and knowing which phase you're in tells you what kind of problem you're actually debugging. **Global Placement** comes first: an analytical engine treats the whole core as continuous space and spreads cells around to minimize estimated wirelength (typically half-perimeter wirelength) while avoiding density spikes. Cell coordinates here are still non-integer floats, and cells can still overlap โ this stage is about rough positioning, not legality.
Tie cells give a gate a clean, high-impedance logic-1 or logic-0 reference instead of letting a designer wire an input straight to the VDD or VSS rail โ that direct wiring is actually dangerous, not just sloppy. Power rails aren't perfectly quiet: switching transients create inductive voltage spikes (Lยทdi/dt), and a direct tie exposes the transistor's ultra-thin gate oxide to those spikes, which can literally punch a hole through it.
Every bulk-CMOS standard cell secretly contains a parasitic four-layer PNPN structure โ basically an accidental thyristor sitting between VDD, the N-well, the P-substrate, and VSS. If substrate current builds up (from noise, I/O undershoot, or a fast transient) and the resulting IยทR drop across the well/substrate resistance crosses about 0.7 V, that parasitic thyristor turns on and regeneratively shorts VDD straight to VSS โ that's latchup, and it can physically destroy the die.
Standard cells depend on continuous N-well, substrate, and poly geometry running the length of a row โ an abrupt dead-end at the row's edge isn't just untidy, it's a manufacturing hazard. At an open row end, lithography sees a sudden transition from patterned silicon to empty space, and that optical discontinuity distorts nearby polysilicon gate shapes during exposure โ endcap cells absorb that distortion so real logic cells don't take the hit.
HPWL is the cheapest useful stand-in for "how long will this net's wiring be": draw the smallest bounding box around all the net's pins, and HPWL is just (width + height) of that box. It's fast for a reason โ global placement has to evaluate millions of candidate cell moves per second, and there's no time to run real routing on each trial, so HPWL gives an O(N) estimate instead.
Core utilization is a single chip-wide number โ total cell area (standard cells plus macros) divided by total core area โ and it's typically targeted around 65โ75% for modern digital blocks as a static planning goal. Local placement density is a much more local, moving-window measurement: cell area inside one G-cell tile divided by that tile's area โ and it can spike to 90โ98% in one pocket even while 40% of the die sits completely empty.
Complex standard cells โ AOI22s, OAI33s, wide MUXes, scan flops โ pack a lot of pins into a small footprint, and when several of them abut directly, their combined pin demand can exceed what the local metal layers can route, causing shorts. Padding is a placement-time trick: tell the tool to treat a cell as wider than it physically is (e.g., a 4-site cell padded to look 6 sites wide), reserving empty "halo" sites on either side that no other cell may occupy.
Magnet placement is exactly what it sounds like: you designate a fixed object โ an I/O port, a macro, or even one macro pin โ as a "magnet," and the placer applies an attractive pull to every standard cell directly wired to it. Without this, interface registers connected to a fixed off-block port can end up scattered randomly across the core by unconstrained global placement, which then makes closing the interface's input-delay timing painful.
Spare cells are unconnected, uncommitted logic gates dropped into the die before CTS, purely as insurance โ if a functional bug turns up after tapeout, you may be able to fix it without a full new mask set. A brand-new mask set costs millions and takes months; if spare gates already exist on silicon nearby, the fix can sometimes be done with just a metal-and-via change, which is roughly 5โ10x cheaper and 3x faster.
Picture every connected pair of cells joined by a rubber band โ that's the wirelength force. The tighter the connectivity, the harder cells get pulled toward each other, because minimizing total quadratic wirelength directly reduces interconnect delay: `ฮฆ = ยฝ ฮฃ c_ij[(xiโxj)ยฒ + (yiโyj)ยฒ]`. Pulled that hard on its own, every cell would collapse into one overlapping point at the center of the die โ obviously unplaceable โ so there needs to be an opposing force.
HFNS builds balanced buffer/inverter trees for high-fanout non-clock control signals (asynchronous resets, scan-enables, chip enables) during placement to fix max-transition and max-capacitance DRC violations. Unlike CTS, HFNS focuses on slew compliance rather than skew balancing.
During placement, registers are still moving around the floorplan constantly โ you simply can't have a real, physically-routed clock tree yet, because the endpoints it would feed haven't settled into their final locations. Because of that, the timing engine models clocks as "ideal": every register's clock pin is assumed to transition at the same instant (often modeled as t=0 or a flat virtual-source latency), with clock uncertainty standing in as a placeholder margin (commonly around 100 ps) for the skew CTS will eventually introduce.
Synthesis stitches scan flip-flops together purely by RTL naming order (reg_a[0] โ reg_a[1] โ ... โ reg_b[0]) with zero awareness of where those flops will eventually sit on the die. Once placement happens geographically, that naming-order chain can become a routing nightmare โ reg_a[0] might land top-left while reg_a[1] lands bottom-right, forcing the scan chain to criss-cross the entire die.
Legalization snaps floating continuous cells onto discrete row site grids, eliminates all instance overlaps, enforces row orientation flipping (N/FS) for power rail abuttal, aligns multi-height cell tracks, and ensures DRC pin-access clearance.
Decap (decoupling capacitor) cells act like tiny local batteries sitting right next to your power-hungry logic โ when thousands of cells switch at once and draw a sudden current surge, the decap discharges its stored charge locally instead of making that current travel all the way back through the power grid, which is what causes dynamic IR drop. Physically, a decap cell is usually just an empty inverter shell wired backwards as a capacitor โ PMOS gate tied to VSS, NMOS gate tied to VDD โ so it behaves as a parallel-plate capacitor rather than an active logic gate.
Multi-Bit Flip-Flop (MBFF) banking merges multiple independent single-bit registers (e.g., 2-bit, 4-bit, or 8-bit) into a single standard cell that shares an internal clock inverter and power structure, reducing clock tree pin capacitance by 40% to 50% and cutting total CTS dynamic power.
Analytical placement is great at scattering random control logic to minimize wirelength, but it's actually the wrong tool for a highly regular datapath โ a 64-bit ALU's bit-0 through bit-63 all follow the identical structure, and letting the placer treat each bit independently scatters them unevenly, creating uneven wire delays and real clock skew. Relative Placement fixes this by letting the designer lock the datapath into an explicit matrix grid โ say, 64 rows by 4 columns โ with each instance assigned an exact relative row/column offset to its neighbors.
Congestion debugging starts with a fast global-routing trial over a grid of G-cells (think of it as running a rough traffic simulation before building real roads) to estimate routing demand versus available capacity. The core metric is overflow: `Overflow = Demand โ Supply` per G-cell, where supply is how many tracks that tile actually has (layers ร tile height รท (width+spacing)) and demand is how many nets need to cross it.
Pre-CTS optimization (`place_opt`) closes WNS and TNS setup slack by executing gate sizing (upsizing drive strength), buffer insertion (splitting long RC nets), logic restructuring (pin swapping, cloning/de-cloning), and multi-threshold voltage (Multi-Vt) swapping under ideal clock assumptions.
Global placement is a somewhat idealized floating-point world โ cells can overlap slightly and sit at fractional coordinates. Legalization is where reality hits: every cell has to snap onto a real, non-overlapping row site, and that snapping can shove cells noticeably far from where the timing-optimized global placement wanted them. In dense regions (say, 90% local utilization), several timing-critical cells are all competing for the same handful of legal sites, forcing the legalizer to displace some of them outward โ sometimes 5 to 20 ยตm โ just to resolve the overlap.
Below about 7nm, the lower metal layers (M1, M2, M3) are typically built with multi-patterning lithography (LELE or SADP) and strict unidirectional routing tracks โ which means simply landing a via on a cell pin is no longer trivial the way it used to be. The core pin-access problem: connecting a via down to a tiny M1 pin stub has to simultaneously satisfy on-grid via center landing, minimum end-of-line enclosure on M1, and minimum cut-to-cut spacing across different lithography masks โ three constraints that can easily conflict with each other on a dense cell.
Think of a multi-voltage chip as separate mini-countries with their own currency (voltage) โ a cell that belongs to the 0.75V CPU domain physically cannot live outside the CPU's voltage-area polygon, because the rows there are wired to VDD_CPU, not the top-level 0.95V rail. The voltage area is a real floorplan object the placer treats as an exclusive move bound โ `create_voltage_area -power_domains {PD_CPU} -region {...}` โ and cells mapped to that domain are hard-constrained inside it.
Buffer bloat is what happens when `place_opt` inserts so many buffers to fix electrical or timing problems that your standard-cell count balloons 20โ40% โ and the instinct to just throw more area at it is usually treating the symptom, not the disease. Root cause #1 โ unrealistic input transitions: if `set_input_transition 5.0` is declared on a port feeding a 0.5 ns clock domain, the tool has to build a chain of 10+ cascaded buffers just to slew that edge down to something usable. Check your SDC input transitions against your actual clock period before blaming the placer.
Scan flip-flops sit in long shift chains connected through dedicated scan-in/scan-out pins, and during scan-shift mode the clock runs much slower than functional speed โ hold time between adjacent scan flops in that mode depends purely on local clock skew. The catch: during placement, clocks are still modeled as ideal (zero latency, zero skew, per the earlier "why clocks must remain ideal" question) โ so under that ideal model, literally every back-to-back scan flop pair looks like a hold violation, even though none of them actually are.
This is the classic "tool divergence" tapeout risk: your implementation tool (ICC2/Innovus) reports clean, comfortable slack pre-CTS, and then signoff PrimeTime โ often run with different assumptions โ reports massive negative slack on the same paths. Closing that gap is what correlation work is about. **Align the SDC first, always.** Both environments need identical clock definitions, generated clocks, multicycle paths, and case analysis โ literally run `check_timing` in PrimeTime and confirm it matches what the implementation tool has active. A single missing generated-clock relationship or mismatched case-analysis setting can produce a huge, misleading divergence that has nothing to do with the physical implementation at all.
Narrow corridors between adjacent SRAMs or IP blocks are the single most common source of unroutable placement layouts โ the global placer naturally wants to pack logic there to minimize wirelength to nearby macro pins, without knowing the channel's metal tracks are already spoken for by macro power rings, PG straps, and wide data buses. Remedy 1 โ partial placement blockages: cap placement density in the channel (say, 30-50%), which leaves the router room to use the rest of the channel's tracks for macro feedthroughs and bus routing while still allowing some minimal buffer placement.
When a combinational logic cone between two pipeline stages has badly imbalanced delay, sizing and buffering alone often can't close setup timing โ retiming actually moves the register boundary itself to rebalance the pipeline. Mechanically: forward retiming slides a flip-flop from a logic gate's input pins to its output pin, and backward retiming does the reverse โ moving a register from a gate's output back across to all of its input pins โ reshaping where pipeline stage boundaries physically sit.
Optimizing a chip for a single corner/mode combination inevitably hurts it in the other operating views it also has to survive โ this is the core tension MMMC placement optimization exists to manage. Concrete example of the conflict: Scenario A (functional mode, slow-slow corner, setup-critical) wants high-drive LVT/ULVT cells and wide buffers to overcome slow transistor switching and high RC delay.
Proceeding into CTS on an unverified or illegal placement is a guaranteed way to build an unroutable clock tree and blow timing closure downstream โ this gate exists precisely to catch that before it becomes an expensive problem. Placement legality must be exactly zero across the board: zero unplaced instances, zero cell overlaps, zero off-grid instances, zero illegal orientation flips, and zero pin-access violations โ "mostly clean" doesn't count here.
place_opt is not one operation -- it's five named stages run in sequence: initial_place (merge clock-gating logic, coarse-place, scan chain optimization if SCANDEF is present), initial_drc (remove existing buffer trees, high-fanout-net synthesis, electrical DRC fixing), initial_opto (timing/area/congestion/leakage-power optimization), final_place (incremental placement to improve timing/congestion, then legalize), and final_opto (further optimization and legalization). You can run any contiguous slice with -from/-to -- the stage names are the actual debug handles.
Before CTS: no outstanding timing violations and no DRV violations from placement, derived clocks correctly expanded and understood, critical clock transitions/capacitance/fanout/congestion areas identified, high-fanout nets already driven with correct drive strength, and any logic areas needing shielding or max-capacitance limits already flagged. CTS is not a fresh start -- it inherits every unresolved placement problem and makes debugging them mixed up with genuine clock-tree problems.
By default, place_opt.flow.trial_clock_tree is false and place_opt uses ideal (zero-delay) clocks throughout placement -- fast, but blind to real clock-tree insertion delay. Setting place_opt.flow.trial_clock_tree = true makes the tool build a temporary clock tree and use propagated clocks during placement instead, which is exactly what CCD (Concurrent Clock and Data optimization) needs to compute useful skew accurately during place_opt. The tradeoff is real: more accurate pre-CTS timing, at the cost of building and discarding a throwaway tree during every placement iteration.
A cluster is a group of cells placed near each other, with location undefined until placement completes -- rarely used now that interconnect-driven placement matured. A region is similar but its location is defined BEFORE placement -- a soft region is a physical constraint with a boundary that may still change during placement; a hard region has boundaries cells cannot cross. Regions are further exclusive (only assigned cells allowed) or non-exclusive.
High-fanout nets (like reset or chip-enable) have one source driving many cells across the core -- not timing-critical individually, but strongly impacting routing area. The reasonable target: reduce fanout to between 40 and 50 connections per driving cell, via buffer insertion (high-fanout net synthesis). Without this, one driver trying to reach hundreds of loads directly would create a routing and drive-strength problem the tool has to solve some other way.
A long wire (small fanout, but driver far from receiver) is often the result of the receiver having stronger connectivity to other instances than to its actual driver. Being highly resistive, it causes a large input transition at the receiver, which increases receiver propagation delay. The fix: segment long wires with buffers -- the same underlying principle as clock buffer insertion, applied to a data-path net.
set_dont_touch excludes a cell from optimization entirely -- nothing about it changes. set_size_only allows sizing (swapping to a different drive-strength variant) but nothing else -- the cell stays in place, keeps its function, but can still be resized to help timing. dont_touch is the stronger, more restrictive constraint; size_only is a middle ground that still gives the optimizer one lever to use.
set_auto_disable_drc_nets controls DRC checking on specific net categories: -none re-enables DRC for all nets, -all disables DRC on all clock/constant/scan-enable/scan-clock nets, and individual flags (-constant, -on_clock_network, -scan) let you target just one category. This exists because some net categories (e.g. constant-tied nets) genuinely don't need the same electrical DRC scrutiny as functional data nets, and checking them anyway wastes analysis effort or produces noise.
By default, set_dont_touch_network [get_clocks CLK] protects BOTH clock paths AND clock-as-data paths -- meaning a clock signal used somewhere as an ordinary data input is also protected. -clock_only narrows that to just the clock paths, leaving clock-as-data usage open to optimization. -clear removes the protection entirely.
legalize_placement removes cell overlap and fits the placement into the row structure, correcting illegal positions to legal ones -- a correctness operation. refine_placement is incremental placement specifically to minimize congestion (with an -effort level, default medium) -- a quality-improvement operation on an already-legal placement. They're not interchangeable: legalization fixes illegality, refinement improves an already-legal result.
Physical synthesis tools combine several primary logic functions into few standard cells, decompose functional gates into equivalent primary gates, and/or duplicate combinational logic (cloning) -- all focused on reconstructing critical paths that are missing timing. Cloning specifically means duplicating a piece of combinational logic so different downstream paths can each get their own copy, rather than sharing one instance whose fanout is spreading the load (and the delay cost) across all of them.
Congestion-driven placement relaxes cell density at the cost of slightly higher interconnect length and silicon area -- it prioritizes routability. Timing-driven placement chases the best timing, possibly leaving congestion issues unresolved. Neither is strictly better; the choice depends on which risk (a routing-incomplete design, or a timing-failing design) is the bigger concern for a given block.
During gain-based optimization, the algorithm computes the gain of each cell along the critical path and tries to maintain equal gain for each stage -- equal gain per stage is optimal timing, a real result from logical effort theory, not an arbitrary heuristic. If an instance needs more gain because of increased output load, its input capacitance is increased to maintain the original gain, rather than just accepting a slower stage.
Dynamic power: Pd = V^2 * sum(fi * Ci) -- summed over nodes, frequency times loading capacitance. Reduce it by lowering supply voltage and/or reducing nodal loading capacitance, which in placement practice means limiting max allowable load capacitance. The real cost: limiting load capacitance has a negative area impact from the excess buffering that becomes necessary to keep individual net loads under that limit.
Static (leakage) power: Ps = V * sum(Ij) -- supply voltage times the sum of per-component leakage currents, characterized per cell in the library. The practical optimization: replace low-Vt cells on non-critical paths with high-Vt cells, since high-Vt cells leak less. Excess static power is a limiting factor in high-performance deep-submicron CMOS, which is exactly why placement algorithms need to be leakage-aware, not just timing- and congestion-aware.
create_placement (plain) does coarse placement and scan chain optimization if SCANDEF is present. -timing_driven adds timing awareness to the coarse pass. -buffering_aware_timing_driven additionally considers high-fanout/long nets and buffers to be added during placement itself. -congestion with -congestion_effort high biases toward routability. -congestion_driven_restructuring runs several iterations of placement plus restructuring together. -floorplan is used specifically in the RP (relative placement) plus hard-macro flow.
place.coarse.cong_restruct_effort (low|medium|high|ultra, default medium) sets how aggressively the tool restructures nets to reduce congestion during coarse placement. place.coarse.cong_restruct_depth_aware, when true, limits path depth growth to 3 logic levels -- a safeguard against restructuring so aggressively that it adds excessive logic depth to a path, potentially trading a congestion win for a new timing problem.
place.legalize.optimize_orientations, when true, flips cells to reduce displacement -- sometimes a flipped orientation lands a cell on a legal site much closer than any unflipped position would. place.legalize.stream_place, when true, moves MANY cells a small amount each rather than one cell a long distance -- tunable via stream_effort and stream_effort_limit. Both target the same underlying goal (minimize disruptive displacement) through genuinely different mechanisms.
GRLB (Global-Route-Layer-Based) improves preroute/postroute correlation, controlled by opt.common.use_route_aware_estimation -- auto (enabled only when per-unit resistance varies across layers), true (always), or a command to remove all global-route-based estimation before routing. RDE (Route-Driven Estimation) performs actual global routing plus extraction from those routes, auto-enabled for technologies below 16nm (otherwise opt.common.enable_rde must be set). Critically, when RDE is enabled, the GRLB setting is IGNORED -- they aren't both active at once.
set_isolate_ports inserts a buffer/inverter pair at specified ports for model accuracy -- {in3 out1} isolates those two ports; -driver LIBCELL_BUFF picks a specific driver cell; -type inverter picks an inverter pair instead. It is NOT applied to bidirectional ports, ports defined as clock sources or power pins, or ports connected to dont_touch nets -- these are deliberate, documented exclusions, not gaps to work around.
place.coarse.icg_auto_bound, when true, auto-generates group bounds for ICGs and the sequential cells they drive -- created at the START of placement and removed at the END, excluding cells already in another group bound. place.coarse.icg_auto_bound_fanout_limit caps the fanout considered for an automatic bound, defaulting to 40 -- an ICG driving more than that many sequential cells doesn't get this automatic grouping treatment.
Quadrature alternately partitions the core into equal instance counts vertically and horizontally, minimizing cut sizes in each direction, starting from the center -- producing an equilibrium between horizontal and vertical routing with no congestion area, which is exactly why most P&R tools default to it. Bisection repeatedly bisects with cut lines until bins hold one or two rows, without necessarily minimizing cut sizes between partitions. Slice-and-bisection divides cells so a smaller partition N/k is assigned to a row by horizontal slicing, then uses recursive vertical bisecting.
Quadratic placement minimizes total squared wire length: Phi(x,y) = 1/2 * sum(c_ij) * [(xi-xj)^2 + (yi-yj)^2]. With a symmetric connectivity matrix C and modified matrix B=D-C (D diagonal, d_ii = sum of c_ij), this reduces to Phi(x,y) = x^T*B*x + y^T*B*y -- because x and y are symmetric and independent, only a 1-D problem needs solving in each dimension. The main problem: it creates very high cost for long wires and very low cost for short wires, so a highly-connected cluster can spread out over the core, increasing congestion and reducing routing-resource flexibility.
Simulated annealing's acceptance rule: P=1 if the cost change (delta C) is <= 0 (always accept an improvement); P = exp(-delta C / T) if delta C > 0 (probabilistically accept a WORSE move, with probability shrinking as temperature T falls). The algorithm starts at a very high temperature and cools per an annealing schedule, so cost-increasing moves become progressively less likely as T falls -- eventually only cost-reducing moves are accepted. A higher initial temperature means more trials and longer runtime, since more of the early search space gets explored via probabilistically-accepted worse moves.
Load-based optimization approximates interconnect capacitance with a wire capacitance model, choosing cell drive strength from the estimated load -- fine when intrinsic delay dominated and wire resistance was low. As designs grew, synthesis using an estimated wire-load model could cause a NON-TERMINATING iterative process of resizing during timing closure, because the wire-load model no longer predicted actual wire lengths until physical design was complete. Gain-based optimization (logical effort) became preferred specifically because of load-independent cell delay -- it doesn't have this circular estimate-then-re-estimate problem.
Total CMOS gate delay: d = p + f (p = parasitic/intrinsic delay, f = effort/extrinsic delay), in units of tau (process-characterized delay through the smallest inverter). f = g*h (g = logical effort, h = electrical effort/gain = Cl/Ci). Logical effort g = T_gate / T_inverter -- the cell's ability to produce output current based on topology, independent of transistor size. A typical inverter (2 PMOS + 1 NMOS transistor units, equal rise/fall) has T=3, giving g=1 -- inverters are the BASELINE; more complex gates are slower (g>1) purely from their topology, before size is even considered.
Several targeted options exist beyond the general effort knob: congestion_layer_aware (per-layer congestion instead of combined layers), increased_cell_expansion (expand virtual cell area per local routing need), congestion_expansion_direction=both (default is horizontal only), ndr_area_aware (accounts for clock NDR area specifically), seq_array_icg_aware (reduces congestion from sequential-array clock nets), spread_repeater_paths (spreads repeaters orthogonally, avoiding clumping at macro/blockage edges), and wide_cell_use_model (wide-cell density modeling for advanced nodes).
set_technology -node {7|7+|5|s5|s4} sets the tech-specific pin-cost model -- a prerequisite. place.coarse.pin_cost_aware (default false) and place.coarse.pin_density_aware (default false, applicable to all tech nodes) are two independent toggles. The COMBINATION of all three -- which node was set, and which of the two booleans are true -- determines the actual resulting behavior: pin-cost-aware placement, pin-density-aware placement, both, or neither.
place.legalize.enable_variant_aware, when true, makes the legalizer fix legalization DRC violations by SWAPPING a cell with an equivalent variant (same function, different mask color or minor geometry, identical timing) rather than moving it. If no variant exists for that cell, the legalizer falls back to the normal behavior: moving the cell. This is a genuinely different repair strategy -- fix-in-place via substitution, versus fix-by-relocation -- and it only works where equivalent variants actually exist in the library.
The flow: place and legalize the block first (or place_opt -to final_place), then source a RedHawk setup script, run analyze_rail -voltage_drop static -nets {VDD VSS}, then enable place.coarse.ir_drop_aware, then optionally set additional IR-drop settings and re-run placement (or place_opt -from final_place). It requires Digital-AF and SNPS_INDESIGN_RH_RAIL license keys. The tool creates three cell categories by total cell count: Upper (top 1%, most spread), Middle (next 5%, less spread), Lower (the rest, not spread) -- percentages changeable via dedicated app-options.
place.coarse.max_density controls max density during non-congestion-driven placement; place.coarse.congestion_driven_max_util (default 0.93) controls max utilization in less-congested areas surrounding highly congested ones. place.coarse.auto_density_control defaults to -enhanced, following a preset schedule unless disabled -- and the enhanced default improves total power and wire length. Critically, you NEVER need to disable auto_density_control to override it -- explicit user settings for max_density/congestion_driven_max_util always take precedence, and settings apply INDEPENDENTLY to each placeable area (voltage area, exclusive move bound), not averaged over the whole block.
Transition (short-circuit) power occurs when the input transition is slow enough that NMOS and PMOS conduct simultaneously, creating a direct supply-to-ground path that contributes nothing to gate operation: Pt = I^2 * (Rp + Rn). This is a genuinely different mechanism from dynamic (switching) power -- it's wasted current from both transistors briefly conducting together, not useful charge/discharge of a load capacitance. Reduce it by controlling max input transitions during placement, or specifying max allowable transition per cell in the library.
magnet_placement declares a fixed object (a fixed macro cell, a pin of a fixed macro cell, or an I/O port) a magnet so connected standard cells get placed close to it -- best performed BEFORE standard cell placement, to improve congestion in a complex floorplan or improve timing. The key rule: cells are pulled ONLY if they form a CONTIGUOUS data path from the magnet. magnet_placement C0 -cells {C3 C4 C5} pulls NOTHING if C3/C4/C5 aren't contiguous with C0; magnet_placement C0 -cells {C3 C6 C8} also pulls nothing if that specific set isn't a contiguous data path, even though C3, C6, and C8 might each individually be reachable from C0 through some path.