pdVerse Notes

Electromigration: What It Is, Why It Happens, and How Physical Design Fixes It

What electromigration actually is, why it happens, what it ends up costing a chip out in the field, and the layout techniques and sign off commands I use to keep it from ever getting that far.

8
sections
7
figures
9
commands referenced
1
equation that runs the whole topic

A Where this sits

power planning → routing → electromigration → signoff

I have noticed most people treat electromigration as just another signoff checkbox, something the tool reports clean or dirty at the end of the flow, and then you move on. But that is backwards, because by the time signoff runs, the decisions that actually decide whether a design has an EM problem, strap width, via count, whether a jog is filleted or not, those were already made way back during power planning and routing. So in this guide I am going in the order that actually matters, what electromigration is, why it happens, what it ends up costing, and where in the flow each fix actually belongs.

This is not meant to replace a full physical design course. I am assuming you already know what a power strap and a via are, and I am going deep on this one reliability mechanism instead of covering the whole signoff stage.

B What electromigration is

Electromigration (EM) is basically the slow displacement of metal ions inside an interconnect, caused by electrons transferring momentum as they keep flowing through it. Give it months or years of operation and this displacement can open up a wire, or short it to whatever is sitting next to it. This is what gets called a reliability failure, and it is a different thing from a functional failure caught at first silicon, because here the chip works fine when it ships and only fails later, out in the field.

There is one distinction I see people skip past a lot, current flowing through a wire and electrons flowing through a wire are not moving in the same direction. Conventional current, the arrow you draw in a schematic, is defined as going from the positive terminal through the circuit to the negative terminal. Electrons, which are the actual charge carriers in a metal, physically move the opposite way.

Current Source + conventional current, I electron flow, e- (the actual carriers) cathode end electrons enter - ions deplete → void anode end electrons exit - ions pile up → hillock Earth Ground

Figure 1. One wire, two directions. Conventional current I is the direction we draw and reason about; electron flow is the actual, physical direction the charge carriers move, the opposite way. The wire's cathode end is where electrons enter, its anode end is where they leave.

Read it: a current source drives a wire to ground. The current arrow points from the source's plus terminal to ground. The electrons actually making that current happen travel the opposite way, from ground, through the wire, back to the source. Electromigration happens because of the electrons' direction, not the current's, and that is the part people mix up.
Why this distinction matters

Electromigration is driven by electron wind, electrons physically colliding into metal ions and pushing them along as they travel. Since electrons flow from cathode to anode, that is also the direction the ions get pushed. So the cathode end of a stressed wire is where metal keeps depleting, and the anode end is where it piles up. Get the direction backwards in your head and you will end up looking for the void on the wrong end of the wire.

C Why it happens

Aluminum and its alloys are especially prone to this electron wind effect, because the bond between aluminum ions and the surrounding lattice is comparatively weak, and grain boundaries basically give the ions an easy path to diffuse through. A sustained, one direction current at normal on chip current densities is enough to move ions along that path over the design's operating lifetime. Copper, which is what almost everyone uses now, is a lot less prone to this at the same current density, which is part of why it replaced aluminum as feature sizes shrank and current densities went up, but it is not immune, and EM rules still apply to copper interconnect too.

❌ Trap

"Electromigration is a DC power rail problem, so signal nets do not need to worry about it." I hear this a lot, and it is not quite right. There are three recognized EM rule types, DC, time varying unidirectional, and bidirectional AC. Power and ground rails are the classic DC case, because they carry current in one direction basically all the time, but any net with a strongly asymmetric duty cycle, a clock net, a high fanout enable, can build up meaningful unidirectional stress too. ERC, electrical rule checking, is what enforces EM rules in the flow, and it is not scoped to power nets alone.

What actually happens at the metal level, over the design's operating life, is this. Electrons enter at the cathode end and exit at the anode end. Every collision between an electron and a metal ion transfers a tiny bit of momentum, and over billions of collisions a second that adds up to a real ion drift in the direction the electrons are flowing. Ions leave the cathode faster than diffusion can replace them, and that opens up a void. Ions pile up at the anode faster than the lattice can absorb them, and that grows into a hillock. Leave it long enough and a void can grow until it severs the wire, that is an open, and a hillock can grow until it touches the line sitting next to it, that is a short.

t₀ - fresh interconnect no current history yet cathode anode electron flow resistance 1.0x (nominal) void depth 0% of width hillock height 0% of width t₁ - months under stress electron wind has moved ions cathode anode electron flow resistance ≈1.4x (rising) void depth ≈35% of width hillock height ≈30% of width t₂ - years under stress: open cathode void has severed the line open neighbor line above cathode anode electron flow resistance ∞ (open) void depth 100% - line severed hillock height touching neighbor - short risk Legend interconnect metal void (ion-depleted) hillock (ion pile-up) reliability-critical marker Cost ledger - Black's equation (illustrative) MTTF = A / J^2 · exp(Ea / kT) A = metal/process constant, J = current density, Ea = activation energy, k = Boltzmann's constant, T = temperature Current density term (1/J^2) doubling J cuts MTTF to roughly 1/4 - the strongest lever a designer controls directly Temperature term (exp(Ea/kT)) ≈10°C of extra junction temperature can roughly halve MTTF, for typical Ea Sign-off target MTTF ≥ 10 years at rated current and worst-case operating temperature, per ERC/EM rules

Figure 2. One interconnect segment, three points in its life under sustained current. Same wire, same coordinate frame in every panel, only the cathode void and anode hillock change. Resistance and void/hillock figures are illustrative, not measured silicon data.

Read it: at t₀ the wire is fresh, full cross section, nominal resistance. By t₁, months into operation, a void has started at the cathode and a hillock at the anode, and resistance is already rising, which is often the earliest symptom a test program actually catches. By t₂, years in, the void has fully severed the wire, that is an open, while the hillock has grown tall enough to risk shorting the line above it.

MTTF = A / J2 · exp(Ea / kT)
Black's equation · A: metal/process constant · J: current density · Ea: activation energy · k: Boltzmann's constant · T: temperature

This is Black's equation, and it is basically the empirical model this whole topic rests on. It says Mean Time To Failure comes down to exactly two things a design actually controls, current density and temperature. Of the two, current density is the one a physical designer sets directly, through strap width, via count, routing decisions, and because it is squared in the denominator it is also the more leveraged one. Double the current density through a wire and you are not doubling the failure rate, you are roughly quartering the time to failure. I will use this directly in Section F.

A side note

Black's equation is old, James Black published it back in 1969 studying aluminum thin films, and it is an empirical fit, not something derived from first principles. Every foundry recalibrates A and Ea for its own process, and those constants are not something a designer chooses or even usually sees. What a designer does see, and does control, is J. So I treat the equation as telling me which lever matters, not as something I am ever going to plug numbers into by hand.

D Consequences

An electromigration failure is a reliability failure, and reliability failures are expensive in a very specific way, they do not show up while you can still fix them cheaply. A design with a marginal EM via array is not slower or wrong at first silicon bring up, it passes every functional test, ships, and then fails later out in the field, sometimes slowly as a resistance drift before it fully opens, sometimes as a hard short once a hillock finally bridges to the line next to it. Both failure modes trace back to the same electron wind mechanism, just at opposite ends of the same wire.

The industry convention is to require a Mean Time To Failure of at least ten years, at the design's rated operating current and worst case junction temperature. That number is not random, it is set to comfortably outlast the product's expected service life, given that Black's equation describes a statistical wear out process and not a hard cutoff, some fraction of parts will always fail earlier than the mean, and ten years is chosen so that fraction stays small enough across a shipped volume.

What if it fails after tapeout?

Then it is not a bug fix anymore, it is a respin, a new mask set, a new fabrication run, and months added to the schedule, on top of whatever field returns and warranty cost already piled up before anyone even traced it back to a via array. It is the same reasoning as any other physical verification escape, the cost of debugging something rises roughly an order of magnitude at every stage it survives to, wafer, then chip, then field, and that is exactly why EM gets checked as ERC during physical design instead of being left for characterization to find later.

Physical Design & Planning Handbook

Dive into 14 comprehensive chapters covering netlist sanity, FinFET grids, macro placement, power grids, CTS, and timing budgeting.

VLSI Physical Design Planning Handbook — fourteen chaptersDesign PlanningFourteen chapters, floorplanning through timing budgets.

E Why it must be fixed here, not later

Static power analysis, which is usually where EM first gets checked in the ICC2 flow, runs after CTS, and it uses the power distribution network's own resistance to estimate voltage drop, ground bounce, and current density with Ohm's and Kirchhoff's laws, instead of running a full dynamic simulation, which would be way too slow to do at this stage. The tool builds a resistor network out of the extracted P/G routing, treats each power net as a constant voltage source, and spreads the average current every transistor draws across that network to work out node voltages and branch current densities.

This approximation assumes there is enough decoupling capacitance to smooth out the sharpest localized transients, and if that assumption does not hold, the analysis can end up being too optimistic, because it is not capturing localized dynamic effects on its own. So the same power planning decision, enough decap, placed close enough to where switching current actually spikes, ends up supporting both the IR drop analysis and the EM analysis built on top of it.

❌ Trap

"IR drop passed, so EM is fine too." I see this assumption a lot as well. They are computed from the same resistor network, but they are different checks against different limits, one on voltage and one on current density, and a design can pass one while failing the other. A wide, well connected mesh can hold voltage drop comfortably in budget while one narrow tap, or an unredundant via, still runs hot enough to fail its own EM limit. Check both, do not assume one from the other.

Because electromigration is a wear out mechanism and not an instant failure, the fix has to happen where the current density is actually set, power planning and routing, not patched on afterward at signoff. Reworking a via array or a strap width after chip finishing is far more disruptive than sizing it right the first time, and it is the same just right tension that shows up everywhere else in physical design, getting a decision made once, correctly, early, costs a lot less than any amount of iteration later.

F How to fix it

Every fix below comes down to one of two levers in Black's equation, lower J, or build in enough redundancy that one weak spot does not turn into a single point of failure. None of them need you to touch the EM checker itself, the checker is only reporting what the layout already is. Let me walk through the five techniques one at a time, roughly in the order I would actually reach for them, each with a picture, so the mechanism is something you can see instead of a bullet point I am asking you to just take on faith.

Before - minimum width, single via default strap, one jog, one via sharp notch - current crowds here 1 via, full J no reservoir past the via strap width = 1.0 W After - widened, redundant vias mitigated strap, filleted jog, via pair, reservoir wide, filleted - no notch 2 vias, J/2 each reservoir segment strap width = 1.4 W Legend power strap metal via enclosure via cut current-crowding risk Cost ledger (illustrative) Strap width 1.0W → 1.4W: J scales as 1/width, so current density drops ≈ 29% Via count at the tap 1 → 2 vias: current per via drops ≈ 50%, and via EM is checked per-via Combined effect on MTTF (Black's equation, 1/J^2 term) ≈(1/0.71)^2 × (1/0.5)^2 ≈ 8x - illustrative, not a measured number Routing cost wider strap + second via + reservoir consume extra track and via-layer area - budget it at power-planning time

Figure 3. Same strap to via connection, before and after mitigation, at identical scale. Left: minimum width, a sharp notch at the jog, one via, no reservoir. Right: widened strap, the jog filleted instead of notched, a redundant via pair, and a reservoir segment left past the via.

Read it: the strap on the right is 1.4× the width of the one on the left, and that alone drops current density by about 29%, since current density scales inversely with width for a fixed current. Splitting the tap across two vias instead of one roughly halves the current each via carries. Combine both, and using the 1/J² term from Black's equation, that works out to a big multiplier on MTTF, the exact number in the ledger is illustrative and not measured, but the direction and the order of magnitude are real. The four figures that follow pull each of these levers apart on their own, so we can reason about what each one buys by itself.

F.1  Widen the strap wherever current density is the binding constraint

This is the most direct lever you have, because J enters Black's equation squared, so even a modest width increase buys a disproportionate MTTF gain, which Figure 3 already showed, 29% less current density from a 1.4× width increase, which roughly quarters the failure rate contribution from that term alone. The trade is routing track, a wider strap on M5 leaves less room for the M5 tracks running next to it, so this is fundamentally a power planning decision, made when the grid's pitch is first drawn, not something you retrofit after the mesh has already been routed around a narrower assumption.

A side note

I have seen new designers reach for via redundancy first, because it feels like the more surgical fix, touch two vias instead of the whole strap. But in practice width is usually cheaper to plan for and harder to add back later, since it has to be reserved in the floorplan from the start. If you only get to fix one thing early, fix the width, via redundancy and the rest of this section are what you layer on top once the width budget is already set.

F.2  Add via redundancy at every tap, and via ladders through the stack

A single via carries a tap's entire local current through one small cut, so that one cut is also the tap's single point of failure, if it starts degrading there is no second path for the current to take. A redundant via pair splits the same current across two cuts, and because each cut now carries roughly half the current, the 1/J² term in Black's equation works in your favor at that specific spot. A via ladder takes the same idea and repeats it at every layer transition between the pin and the strap, instead of only at the last hop, the tool builds this by marking a via rule as EM motivated and letting the router stack redundant cuts as it climbs the metal stack.

A -- single via all current through one cut M4 (pin / tap) M5 (strap) single cut -- full I current, I cross-section B -- redundant via pair current splits across two cuts M4 (pin / tap) M5 (strap) two cuts -- I/2 each current, I cross-section C -- via ladder (stacked) redundancy at every layer transition M2 M3 M4 M5 redundant pair redundant pair redundant pair same redundancy repeated at every layer transition Legend M4 (pin/tap layer) M5 (strap layer) via cut Cost ledger -- via redundancy (illustrative) Single via (A) carries 100% of the tap's local current through one cut -- the whole cut is the failure point Redundant pair (B) current splits roughly evenly -- each cut sees ~I/2, and 1/J^2 in Black's equation makes MTTF rise sharply Via ladder (C) the same split repeated at every layer transition from pin to strap -- no single weak layer transition left in the stack Escalation order postroute insertion first, then concurrent soft-rule insertion, then near-100% hard-rule only where genuinely required

Figure 4. Panel A: one via carries the tap's entire current. Panel B: a redundant pair splits it roughly in half at that layer. Panel C: the same redundant pair repeated at every transition from M2 up to M5, so no single layer hop is left as the weak point.

Read it: the cross section insets under panels A and B show the physical difference, one cut versus two, both with enclosure metal visible on the layer above and below, per the technology's via rule. Panel C is the same trick stacked, a tap is only as strong as its weakest single layer transition, so a via ladder removes that weak point at every hop instead of just the one closest to the strap.

F.3  Eliminate notches and sharp jogs on major current carrying lines

A right angle turn in a wire does not change its nominal width, and a DRC deck will pass it without complaint. But electrically, current does not turn corners evenly, it crowds toward the inside edge of a sharp bend, so the effective conducting width at that corner ends up less than the drawn width, even though nothing in the layout actually looks narrower. This is the notch problem in its most common form, not a literal manufacturing defect, just a geometry induced current density spike hiding at a perfectly legal jog.

A -- sharp notch at the jog current crowds into the inside corner sharp inside corner -- current crowds to this edge W I B -- filleted jog, same nominal width corner current spreads across full width chamfer holds width constant through the turn W I Legend M4 strap metal current-crowding marker Cost ledger -- notch elimination (illustrative) Nominal strap width identical in both panels (W) -- the fix costs no extra routing track Current path at the turn panel A: current lines bunch toward the inside corner, so the effective conducting width is less than W; panel B: the chamfer lets current spread across the full width Current density at the turn panel A: J spikes locally at the inside corner -- a hidden EM hot spot even though nominal width never changed; panel B: J stays close to uniform MTTF impact removing one notch removes one 1/J^2 penalty at that specific site -- no benefit anywhere else on the strap

Figure 5. Panel A: a plain right angle jog. Current lines bunch toward the inside corner, so J spikes there even though the strap's nominal width W never changes. Panel B: the same jog with the inside corner chamfered instead of cut square, letting current spread across the full width through the turn.

Read it: both panels carry the same nominal width W, dimensioned identically on the right of each panel, so this fix costs basically no extra routing track, unlike widening the whole strap. Chamfering, or on some routers rounding, the inside corner is usually a router setting or a post route cleanup pass, not a change to the strap's declared width anywhere else on the net.

Why this one is easy to miss

A notch or a sharp jog will not show up in a static width report, because that report is checking drawn geometry, not current distribution. It only shows up in a current density aware check, the same analyze_rail -electromigration style of analysis I talk about in Section G. So a design can look completely clean by width and still be carrying a hidden EM hot spot at every unfilleted jog on a heavily loaded strap.

F.4  Leave a reservoir segment past the last via on a tap

If you remember from Section C, a void grows by consuming metal ions at the cathode end of a stressed wire. If a tap's metal ends exactly at the via, the void has nowhere else to draw from but the electrically active connection itself, and it starts eating into that connection almost right away. A reservoir segment is just a short length of extra metal left past the via, wired to nothing, its only job is giving the void something to consume first, before it can reach the connection that actually matters.

A -- no reservoir void reaches the via almost immediately cathode end via to M3/pin electron flow → cathode void grows leftward same tap, months later: void already closing on the via B -- reservoir segment left past the via void must consume the reservoir first reservoir, L_r extra metal, unconnected cathode end via to M3/pin electron flow → cathode void grows leftward same tap, months later: via still fully protected Legend M4 tap metal void (ion-depleted) reliability-critical marker Cost ledger -- reservoir segments (illustrative) Routing cost one short stub of otherwise-unused metal beyond the last via -- a few extra square microns per tap What it buys the void must fully consume the reservoir length before it can start eating into the electrically active via connection Where it matters most taps and jogs that already run close to their EM limit, where a small MTTF margin is the difference between pass and fail Not a substitute for adequate strap width or via redundancy -- a reservoir delays the failure, it does not lower the current density that causes it

Figure 6. Panel A: no reservoir, the via sits at the very end of the tap, so months of stress bring the void right up against it. Panel B: the same tap with a reservoir segment Lr left past the via, at the same elapsed time the void is still consuming the reservoir and the via is untouched.

Read it: both "months later" insets are drawn at the same elapsed stress time, in the same coordinate frame, so the comparison is fair, the only difference between the two taps is whether that extra stub of metal exists past the via. It is one of the cheapest fixes in this whole section, in routing track terms, but it only buys time, it does not lower the current density that is actually driving the void.

F.5  Distribute clock buffers rather than clumping them

Everything up to this point has been about one tap or one jog. This technique is basically the same idea at the block level. Static power analysis, like I talked about in Section E, is built on the block's average current draw spread across the power grid's resistor network. If clock buffers, or any other high activity cells, end up clustered together during placement, the tap closest to that cluster sees a lot more than its average share of current, even while the block wide average still looks comfortably within budget.

A -- clock buffers clumped one strap tap carries the whole block's clock current 0x 0x 8x 0x 0x demand at each strap tap, in buffer-units: B -- clock buffers distributed the same total load, spread across five taps 3x 1x 1x 2x 1x demand at each strap tap, in buffer-units: Legend clock buffer instance M5 power strap local current hot spot (>=5x) Cost ledger -- buffer distribution (illustrative) Block average current identical in both panels -- eight buffers, same total switching load, same average across the whole block Panel A -- clumped 8 buffers land on one strap tap -- that tap alone carries roughly 8x a single buffer's current, even though the block average looks fine Panel B -- distributed the same 8 buffers spread across five taps -- no tap carries more than 1-2 buffer-units, well inside its EM limit Why average-based checks miss this static IR/EM analysis reads current density per shape, not per block -- a clumped tap fails locally while the block-level average stays comfortably in budget

Figure 7. Same eight clock buffers, same total switching load, placed two different ways. Panel A: clumped near one strap tap, which then carries roughly eight buffer units of current. Panel B: spread across five taps, none of which carries more than one or two buffer units.

Read it: the numbers under each strap tap are just a straight tally of which buffers connect to which tap, capacity and demand drawn as things you can count, not a shaded heatmap. Panel A's marked tap is exactly the kind of hot spot a block average IR/EM summary can miss, because the tool's average based static analysis is reading current density per shape and not per block, one overloaded tap can fail while the reported block average stays perfectly clean.

In practice

The most common source of EM, and IR, trouble is not exotic, it is undersized width right where power and ground pads connect into the core P/G bus, because that junction carries the highest local current density on the whole grid. Widening the connections at pad to bus junctions specifically (F.1), and keeping those runs notch free (F.3), closes most of the gap before you even need to think about via level redundancy (F.2).

Note

Via redundancy is not free of trade offs elsewhere in the flow though. Pushing the redundant via conversion rate toward 100% makes routing DRC convergence harder and can affect signal integrity, and very high rates can force a lower floorplan utilization just to leave room. Start with postroute insertion, escalate to concurrent soft rule insertion only if you actually need more, and save near 100%, hard rule insertion for blocks that genuinely require it.

G Related commands

These span two tools, ICC2, where the layout decisions actually get made, and RedHawk / RedHawk-SC Fusion, where PG electromigration actually gets analyzed. None of them replace the layout decisions from Section F, they are just how those decisions get inserted, verified, and signed off.

Command / optionWhat it doesWhere it runs
create_via_rule ... for_electro_migration trueMarks a via ladder rule as EM motivated rather than performance motivated, so ladder insertion prioritizes redundancy over minimum delay at that pin.ICC2 (Zroute)
set_via_ladder_candidate / insert_via_laddersAssigns EM via ladder rules to specific pins or nets and inserts the stacked, redundant via structure before global routing.ICC2 (Zroute)
add_redundant_vias -effort highConverts single vias into redundant pairs after detail routing, on signal and PG nets, directly lowering per via current.ICC2 (Zroute)
create_routing_rule -widths {...}Defines a non default routing rule that widens specific nets, power straps in particular, beyond the technology default, lowering current density.ICC2
spread_wires / widen_wiresPost route DFM commands that widen wires or spread them apart, reducing both critical area risk and current density hot spots.ICC2 (Zroute)
analyze_rail -electromigrationRuns PG electromigration analysis after chip finishing, reporting current density violations per shape against the technology's EM limits.ICC2 + RedHawk / RedHawk-SC Fusion
rail.em_only_tech_filePoints RedHawk-SC at an EM only technology rule file, for a run scoped to electromigration rather than full rail analysis.RedHawk-SC Fusion
signoff_check_drc (foundry runset)Foundry DRC decks encode EM related width, spacing, and notch rules for the power grid as ordinary signoff DRC, not a separate electromigration checker.ICC2 + IC Validator
Black's equationNot a command, it is the physical model every EM current density limit, via count rule, and width rule in the technology file is ultimately derived from.Physics, not a tool
🍀 Takeaway

Every command in this table is either changing current density, adding redundancy, or checking the result, because those are the only two levers Black's equation actually gives a designer. If a proposed fix does neither, it is not really an EM fix, whatever else it might be good for.

H Check yourself

1

A wire's cathode end is losing metal and its anode end is gaining metal. Which direction are the electrons flowing, and which direction is conventional current flowing?

2

A block passes its IR drop signoff comfortably. Explain, in one sentence, why that does not by itself tell you the power grid is EM clean.

3

Using Black's equation, roughly what happens to a wire's MTTF if a floorplan change increases the current density through it by 50%?

4

Name two layout level fixes for electromigration risk that do not require widening a wire.

5

Why does an EM failure typically show up months or years after tapeout rather than at first silicon bring up?

Six stages, one flow

Electromigration is just one detail inside power planning and routing, two out of six stages between an empty floorplan and a chip that actually boots. I put together a free 14 chapter Design Planning handbook that covers all six, with the mechanisms and the tradeoffs.

Get the handbook

Free interview practice

Practise interview questions on this topic

Premium libraries

Go deeper with the pdVerse libraries

Complete pdVerse LibraryAll eight collections: every library above plus the Low Power Tool Guide, in one purchase.₹1,237