A Where this sits
I have noticed most people treat electromigration as just another signoff checkbox, something the tool reports clean or dirty at the end of the flow, and then you move on. But that is backwards, because by the time signoff runs, the decisions that actually decide whether a design has an EM problem, strap width, via count, whether a jog is filleted or not, those were already made way back during power planning and routing. So in this guide I am going in the order that actually matters, what electromigration is, why it happens, what it ends up costing, and where in the flow each fix actually belongs.
This is not meant to replace a full physical design course. I am assuming you already know what a power strap and a via are, and I am going deep on this one reliability mechanism instead of covering the whole signoff stage.
B What electromigration is
Electromigration (EM) is basically the slow displacement of metal ions inside an interconnect, caused by electrons transferring momentum as they keep flowing through it. Give it months or years of operation and this displacement can open up a wire, or short it to whatever is sitting next to it. This is what gets called a reliability failure, and it is a different thing from a functional failure caught at first silicon, because here the chip works fine when it ships and only fails later, out in the field.
There is one distinction I see people skip past a lot, current flowing through a wire and electrons flowing through a wire are not moving in the same direction. Conventional current, the arrow you draw in a schematic, is defined as going from the positive terminal through the circuit to the negative terminal. Electrons, which are the actual charge carriers in a metal, physically move the opposite way.
Figure 1. One wire, two directions. Conventional current I is the direction we draw and reason about; electron flow is the actual, physical direction the charge carriers move, the opposite way. The wire's cathode end is where electrons enter, its anode end is where they leave.
Electromigration is driven by electron wind, electrons physically colliding into metal ions and pushing them along as they travel. Since electrons flow from cathode to anode, that is also the direction the ions get pushed. So the cathode end of a stressed wire is where metal keeps depleting, and the anode end is where it piles up. Get the direction backwards in your head and you will end up looking for the void on the wrong end of the wire.
C Why it happens
Aluminum and its alloys are especially prone to this electron wind effect, because the bond between aluminum ions and the surrounding lattice is comparatively weak, and grain boundaries basically give the ions an easy path to diffuse through. A sustained, one direction current at normal on chip current densities is enough to move ions along that path over the design's operating lifetime. Copper, which is what almost everyone uses now, is a lot less prone to this at the same current density, which is part of why it replaced aluminum as feature sizes shrank and current densities went up, but it is not immune, and EM rules still apply to copper interconnect too.
"Electromigration is a DC power rail problem, so signal nets do not need to worry about it." I hear this a lot, and it is not quite right. There are three recognized EM rule types, DC, time varying unidirectional, and bidirectional AC. Power and ground rails are the classic DC case, because they carry current in one direction basically all the time, but any net with a strongly asymmetric duty cycle, a clock net, a high fanout enable, can build up meaningful unidirectional stress too. ERC, electrical rule checking, is what enforces EM rules in the flow, and it is not scoped to power nets alone.
What actually happens at the metal level, over the design's operating life, is this. Electrons enter at the cathode end and exit at the anode end. Every collision between an electron and a metal ion transfers a tiny bit of momentum, and over billions of collisions a second that adds up to a real ion drift in the direction the electrons are flowing. Ions leave the cathode faster than diffusion can replace them, and that opens up a void. Ions pile up at the anode faster than the lattice can absorb them, and that grows into a hillock. Leave it long enough and a void can grow until it severs the wire, that is an open, and a hillock can grow until it touches the line sitting next to it, that is a short.
Figure 2. One interconnect segment, three points in its life under sustained current. Same wire, same coordinate frame in every panel, only the cathode void and anode hillock change. Resistance and void/hillock figures are illustrative, not measured silicon data.
This is Black's equation, and it is basically the empirical model this whole topic rests on. It says Mean Time To Failure comes down to exactly two things a design actually controls, current density and temperature. Of the two, current density is the one a physical designer sets directly, through strap width, via count, routing decisions, and because it is squared in the denominator it is also the more leveraged one. Double the current density through a wire and you are not doubling the failure rate, you are roughly quartering the time to failure. I will use this directly in Section F.
Black's equation is old, James Black published it back in 1969 studying aluminum thin films, and it is an empirical fit, not something derived from first principles. Every foundry recalibrates A and Ea for its own process, and those constants are not something a designer chooses or even usually sees. What a designer does see, and does control, is J. So I treat the equation as telling me which lever matters, not as something I am ever going to plug numbers into by hand.
D Consequences
An electromigration failure is a reliability failure, and reliability failures are expensive in a very specific way, they do not show up while you can still fix them cheaply. A design with a marginal EM via array is not slower or wrong at first silicon bring up, it passes every functional test, ships, and then fails later out in the field, sometimes slowly as a resistance drift before it fully opens, sometimes as a hard short once a hillock finally bridges to the line next to it. Both failure modes trace back to the same electron wind mechanism, just at opposite ends of the same wire.
The industry convention is to require a Mean Time To Failure of at least ten years, at the design's rated operating current and worst case junction temperature. That number is not random, it is set to comfortably outlast the product's expected service life, given that Black's equation describes a statistical wear out process and not a hard cutoff, some fraction of parts will always fail earlier than the mean, and ten years is chosen so that fraction stays small enough across a shipped volume.
Then it is not a bug fix anymore, it is a respin, a new mask set, a new fabrication run, and months added to the schedule, on top of whatever field returns and warranty cost already piled up before anyone even traced it back to a via array. It is the same reasoning as any other physical verification escape, the cost of debugging something rises roughly an order of magnitude at every stage it survives to, wafer, then chip, then field, and that is exactly why EM gets checked as ERC during physical design instead of being left for characterization to find later.
Physical Design & Planning Handbook
Master ASIC Physical Design Planning & Floorplanning
Dive into 14 comprehensive chapters covering netlist sanity, FinFET grids, macro placement, power grids, CTS, and timing budgeting.

E Why it must be fixed here, not later
Static power analysis, which is usually where EM first gets checked in the ICC2 flow, runs after CTS, and it uses the power distribution network's own resistance to estimate voltage drop, ground bounce, and current density with Ohm's and Kirchhoff's laws, instead of running a full dynamic simulation, which would be way too slow to do at this stage. The tool builds a resistor network out of the extracted P/G routing, treats each power net as a constant voltage source, and spreads the average current every transistor draws across that network to work out node voltages and branch current densities.
This approximation assumes there is enough decoupling capacitance to smooth out the sharpest localized transients, and if that assumption does not hold, the analysis can end up being too optimistic, because it is not capturing localized dynamic effects on its own. So the same power planning decision, enough decap, placed close enough to where switching current actually spikes, ends up supporting both the IR drop analysis and the EM analysis built on top of it.
"IR drop passed, so EM is fine too." I see this assumption a lot as well. They are computed from the same resistor network, but they are different checks against different limits, one on voltage and one on current density, and a design can pass one while failing the other. A wide, well connected mesh can hold voltage drop comfortably in budget while one narrow tap, or an unredundant via, still runs hot enough to fail its own EM limit. Check both, do not assume one from the other.
Because electromigration is a wear out mechanism and not an instant failure, the fix has to happen where the current density is actually set, power planning and routing, not patched on afterward at signoff. Reworking a via array or a strap width after chip finishing is far more disruptive than sizing it right the first time, and it is the same just right tension that shows up everywhere else in physical design, getting a decision made once, correctly, early, costs a lot less than any amount of iteration later.
F How to fix it
Every fix below comes down to one of two levers in Black's equation, lower J, or build in enough redundancy that one weak spot does not turn into a single point of failure. None of them need you to touch the EM checker itself, the checker is only reporting what the layout already is. Let me walk through the five techniques one at a time, roughly in the order I would actually reach for them, each with a picture, so the mechanism is something you can see instead of a bullet point I am asking you to just take on faith.
Figure 3. Same strap to via connection, before and after mitigation, at identical scale. Left: minimum width, a sharp notch at the jog, one via, no reservoir. Right: widened strap, the jog filleted instead of notched, a redundant via pair, and a reservoir segment left past the via.
F.1 Widen the strap wherever current density is the binding constraint
This is the most direct lever you have, because J enters Black's equation squared, so even a modest width increase buys a disproportionate MTTF gain, which Figure 3 already showed, 29% less current density from a 1.4× width increase, which roughly quarters the failure rate contribution from that term alone. The trade is routing track, a wider strap on M5 leaves less room for the M5 tracks running next to it, so this is fundamentally a power planning decision, made when the grid's pitch is first drawn, not something you retrofit after the mesh has already been routed around a narrower assumption.
I have seen new designers reach for via redundancy first, because it feels like the more surgical fix, touch two vias instead of the whole strap. But in practice width is usually cheaper to plan for and harder to add back later, since it has to be reserved in the floorplan from the start. If you only get to fix one thing early, fix the width, via redundancy and the rest of this section are what you layer on top once the width budget is already set.
F.2 Add via redundancy at every tap, and via ladders through the stack
A single via carries a tap's entire local current through one small cut, so that one cut is also the tap's single point of failure, if it starts degrading there is no second path for the current to take. A redundant via pair splits the same current across two cuts, and because each cut now carries roughly half the current, the 1/J² term in Black's equation works in your favor at that specific spot. A via ladder takes the same idea and repeats it at every layer transition between the pin and the strap, instead of only at the last hop, the tool builds this by marking a via rule as EM motivated and letting the router stack redundant cuts as it climbs the metal stack.
Figure 4. Panel A: one via carries the tap's entire current. Panel B: a redundant pair splits it roughly in half at that layer. Panel C: the same redundant pair repeated at every transition from M2 up to M5, so no single layer hop is left as the weak point.
F.3 Eliminate notches and sharp jogs on major current carrying lines
A right angle turn in a wire does not change its nominal width, and a DRC deck will pass it without complaint. But electrically, current does not turn corners evenly, it crowds toward the inside edge of a sharp bend, so the effective conducting width at that corner ends up less than the drawn width, even though nothing in the layout actually looks narrower. This is the notch problem in its most common form, not a literal manufacturing defect, just a geometry induced current density spike hiding at a perfectly legal jog.
Figure 5. Panel A: a plain right angle jog. Current lines bunch toward the inside corner, so J spikes there even though the strap's nominal width W never changes. Panel B: the same jog with the inside corner chamfered instead of cut square, letting current spread across the full width through the turn.
A notch or a sharp jog will not show up in a static width report, because that report is checking drawn geometry, not current distribution. It only shows up in a current density aware check, the same analyze_rail -electromigration style of analysis I talk about in Section G. So a design can look completely clean by width and still be carrying a hidden EM hot spot at every unfilleted jog on a heavily loaded strap.
F.4 Leave a reservoir segment past the last via on a tap
If you remember from Section C, a void grows by consuming metal ions at the cathode end of a stressed wire. If a tap's metal ends exactly at the via, the void has nowhere else to draw from but the electrically active connection itself, and it starts eating into that connection almost right away. A reservoir segment is just a short length of extra metal left past the via, wired to nothing, its only job is giving the void something to consume first, before it can reach the connection that actually matters.
Figure 6. Panel A: no reservoir, the via sits at the very end of the tap, so months of stress bring the void right up against it. Panel B: the same tap with a reservoir segment Lr left past the via, at the same elapsed time the void is still consuming the reservoir and the via is untouched.
F.5 Distribute clock buffers rather than clumping them
Everything up to this point has been about one tap or one jog. This technique is basically the same idea at the block level. Static power analysis, like I talked about in Section E, is built on the block's average current draw spread across the power grid's resistor network. If clock buffers, or any other high activity cells, end up clustered together during placement, the tap closest to that cluster sees a lot more than its average share of current, even while the block wide average still looks comfortably within budget.
Figure 7. Same eight clock buffers, same total switching load, placed two different ways. Panel A: clumped near one strap tap, which then carries roughly eight buffer units of current. Panel B: spread across five taps, none of which carries more than one or two buffer units.
The most common source of EM, and IR, trouble is not exotic, it is undersized width right where power and ground pads connect into the core P/G bus, because that junction carries the highest local current density on the whole grid. Widening the connections at pad to bus junctions specifically (F.1), and keeping those runs notch free (F.3), closes most of the gap before you even need to think about via level redundancy (F.2).
Via redundancy is not free of trade offs elsewhere in the flow though. Pushing the redundant via conversion rate toward 100% makes routing DRC convergence harder and can affect signal integrity, and very high rates can force a lower floorplan utilization just to leave room. Start with postroute insertion, escalate to concurrent soft rule insertion only if you actually need more, and save near 100%, hard rule insertion for blocks that genuinely require it.
G Related commands
These span two tools, ICC2, where the layout decisions actually get made, and RedHawk / RedHawk-SC Fusion, where PG electromigration actually gets analyzed. None of them replace the layout decisions from Section F, they are just how those decisions get inserted, verified, and signed off.
| Command / option | What it does | Where it runs |
|---|---|---|
create_via_rule ... for_electro_migration true | Marks a via ladder rule as EM motivated rather than performance motivated, so ladder insertion prioritizes redundancy over minimum delay at that pin. | ICC2 (Zroute) |
set_via_ladder_candidate / insert_via_ladders | Assigns EM via ladder rules to specific pins or nets and inserts the stacked, redundant via structure before global routing. | ICC2 (Zroute) |
add_redundant_vias -effort high | Converts single vias into redundant pairs after detail routing, on signal and PG nets, directly lowering per via current. | ICC2 (Zroute) |
create_routing_rule -widths {...} | Defines a non default routing rule that widens specific nets, power straps in particular, beyond the technology default, lowering current density. | ICC2 |
spread_wires / widen_wires | Post route DFM commands that widen wires or spread them apart, reducing both critical area risk and current density hot spots. | ICC2 (Zroute) |
analyze_rail -electromigration | Runs PG electromigration analysis after chip finishing, reporting current density violations per shape against the technology's EM limits. | ICC2 + RedHawk / RedHawk-SC Fusion |
rail.em_only_tech_file | Points RedHawk-SC at an EM only technology rule file, for a run scoped to electromigration rather than full rail analysis. | RedHawk-SC Fusion |
signoff_check_drc (foundry runset) | Foundry DRC decks encode EM related width, spacing, and notch rules for the power grid as ordinary signoff DRC, not a separate electromigration checker. | ICC2 + IC Validator |
| Black's equation | Not a command, it is the physical model every EM current density limit, via count rule, and width rule in the technology file is ultimately derived from. | Physics, not a tool |
Every command in this table is either changing current density, adding redundancy, or checking the result, because those are the only two levers Black's equation actually gives a designer. If a proposed fix does neither, it is not really an EM fix, whatever else it might be good for.
H Check yourself
A wire's cathode end is losing metal and its anode end is gaining metal. Which direction are the electrons flowing, and which direction is conventional current flowing?
A block passes its IR drop signoff comfortably. Explain, in one sentence, why that does not by itself tell you the power grid is EM clean.
Using Black's equation, roughly what happens to a wire's MTTF if a floorplan change increases the current density through it by 50%?
Name two layout level fixes for electromigration risk that do not require widening a wire.
Why does an EM failure typically show up months or years after tapeout rather than at first silicon bring up?
Six stages, one flow
Electromigration is just one detail inside power planning and routing, two out of six stages between an empty floorplan and a chip that actually boots. I put together a free 14 chapter Design Planning handbook that covers all six, with the mechanisms and the tradeoffs.
Get the handbookFree interview practice
Practise interview questions on this topic
- What does EM signoff check: average, RMS or peak current? Beginner
- Postroute, concurrent or near-100% redundant vias: how do you escalate? Intermediate
- How does electromigration constrain routing, and what does Black's equation tell you? Intermediate
- What does chasing near-100% redundant vias cost, and how do you protect timing? Expert
Premium libraries
Go deeper with the pdVerse libraries
Master VLSI physical design planning. Read 14 chapters free online or get the complete PDF bundle with floorplanning, power, CTS & timing budgeting.
Free to read · PDF ₹179 13 guidesSignoff Academy13 ASIC signoff guides: DRC, LVS, ERC, antenna, metal fill, extraction, STA/SI, CLP, EM/IR, LEC, and tapeout handoff evidence.
Full bundle ₹179 10 chaptersSTA Library10-chapter STA handbook for VLSI engineers: setup/hold, slack, Liberty, parasitics, OCV/AOCV, crosstalk, PrimeTime reports, and interview Q&A.
Full bundle ₹199 9 chaptersTiming Constraints (SDC) LibrarySynopsys Design Constraints handbook: create_clock, generated clocks, I/O delays, clock groups, false paths, multicycle paths, and SDC linting.
Full bundle ₹179 10 chaptersMMMC LibraryMMMC / MCMM timing signoff: modes vs corners, PVT, RC parasitics, analysis views, ICC2 vs PrimeTime correlation, and scenario explosion.
Full bundle ₹179 9 chaptersLow Power LibraryLearn low-power VLSI design through practical eBooks on power domains, isolation, level shifting, retention, UPF and multivoltage implementation.
Full bundle ₹179 8 chaptersPnR Flow Mentor GuideRead the 8-chapter ICC2 Implementation Mentor Guide free online, or get the PDF bundle. Covers placement, CTS, routing, and ECO flow.
Free to read · PDF ₹199