Skip to main content

Clock Gating

PPA Tradeoffs named switching-power reduction as one of the three pulls every synthesis run balances. Clock gating is the single most common technique for it — and it's easy to confuse with UPF's power switches, which solve a genuinely different problem despite both being called "saving power."

Finding what to gate​

A synthesis tool identifies groups of flip-flops that share a common enable signal — the same condition that already decides, in the RTL, whether those flops actually update this cycle or hold their current value. Rather than letting the clock keep toggling those flops every cycle regardless, the tool inserts a clock gating cell (an integrated clock gating cell, or ICG) — supplied directly by the standard cell library, the same library technology mapping draws every other cell from — that uses the enable signal to stop the clock itself from reaching those flops whenever the enable is false.

Why a plain AND gate isn't good enough​

The obvious first idea — AND the clock directly with the enable signal — has a real, specific flaw: if enable changes while the clock is high, the AND gate's output can glitch, producing a spurious extra clock edge nothing downstream expects. The standard fix is a latch-based gating cell: a level-sensitive latch captures the enable signal only while the clock is low, holding it stable for the entire high phase that follows — guaranteeing the gated clock's actual edges only ever happen where a real, intended edge belongs.

Timing diagram for signals: clk (ungated), enable, gated clkclk (ungated)enablegated clk

enable drops during a clk-low phase, so the latch inside the ICG captures it cleanly — the gated clock simply stops toggling for as long as enable stays low, then resumes the instant it goes high again, with no glitch and no missing or extra edge at either boundary.

While enable is low, the flops behind this gating cell see no clock edges whatsoever — no switching, and critically, zero dynamic power for exactly those cycles, on exactly those flops.

Not every enabled flop is worth gating: the minimum-bit-width threshold​

An ICG isn't free — it's real cell area, and its own internal latch and gating logic burn a small amount of switching power every cycle, gated or not. Applying one to a group of flip-flops only pays for itself once enough flops actually share that gate's clock — gate a single flop, or two, and the ICG's own overhead can cost more area and power than the switching it saves on that tiny group. Synthesis tools handle this with a minimum-bit-width threshold: a user-set option (commonly a value like 4 or 8) below which the tool won't insert clock gating at all, even where an enable condition would otherwise make a group a valid candidate. Setting the threshold too low wastes ICGs on groups too small to break even; setting it too high leaves real, gate-able switching power on the table. Getting this number right is itself part of the same PPA balancing act PPA Tradeoffs already named — clock gating's own insertion decision has a tradeoff buried inside it, not just a binary "gate it or don't."

Why the clock specifically is worth this much attention​

The clock signal isn't just one more net among many — it's the single most frequently toggling signal in the entire design, switching every cycle whether or not the logic behind it actually does anything useful that cycle. Because dynamic power scales directly with switching activity, the clock distribution network alone is commonly cited as accounting for something like a third to half of a chip's total dynamic power — a genuinely disproportionate share for a signal that carries no data of its own. That's exactly why clock gating, despite being conceptually simple, is worth this much engineering attention: cutting switching on the highest-activity signal in the design is a far bigger lever than the same effort spent gating any one data path.

The real distinction from UPF's power switches​

Both techniques are described as "saving power," and it's worth being precise about why they aren't the same tool for the same job:

Clock gatingUPF power switch
What it stopsThe clock toggling those specific flopsThe domain's actual supply
Power savedDynamic (switching) power onlyDynamic and static (leakage) power
Domain stateStays fully powered the whole timeGenuinely unpowered — real leakage current stops
Wake-up costNone — resumes toggling the instant enable goes high againReal latency — the power-up sequence UPF already covered, isolation/retention included
GranularityTypically per flip-flop group, applied automatically by synthesisTypically per power domain, explicitly authored in UPF

Clock gating is cheap and instantaneous specifically because it doesn't actually remove power — it just stops switching. A UPF power switch saves more (leakage too, not just switching activity) but costs real complexity and real wake-up latency to do it. Real chips use both, at different granularities, for exactly the reasons this table lays out — they aren't competing solutions to the same problem.

What's next​

Every optimization technique in this topic — technology-independent restructuring, mapping, constraint-driven optimization, and now clock gating — has been covered. The next page covers what synthesis actually hands off once all of it finishes: the real files that leave this process and feed directly into content this site already shipped.