A Global-Local Optimization Framework for Simultaneous Multi-Mode Multi-Corner Clock Skew Variation Reduction Kwangsoo Han, Andrew B. Kahng, Jongpil Lee, Jiajia Li and Siddhartha Nath VLSI CAD LABORATORY, UC San Diego Outline Motivation
"A Global-Local Optimization Framework for" is the property of its rightful owner. Permission is granted to
download and print the materials on this website for personal, non-commercial use only, and to display it
on your personal computer provided you do not modify the materials and that you retain all copyright
notices contained in the materials. By downloading content from our website, you accept the terms of this
agreement.
Presentation Transcript
01
A Global-Local Optimization Framework for Simultaneous Multi-Mode Multi-Corner Clock Skew Variation Reduction Kwangsoo Han, Andrew B. Kahng, Jongpil Lee, Jiajia Li and Siddhartha Nath
VLSI CAD LABORATORY, UC San Diego<br>
02
Outline Motivation
Related Work
Our Optimization Framework
Experimental Setup and Results
Conclusions<br>
03
Motivation Many signoff PVT corners in modern SoCs
Clock skew variation across corners “ping-pong” effect == fixing timing issues at one corner leads to timing violation at others
Our goal: Minimize clock skew variation Low voltage: gate delay dominates
High voltage: wire delay dominates
Skew reversal Power/area overheads<br>
04
Outline Motivation
Related Work
Our Optimization Framework
Experimental Setup and Results
Conclusions<br>
05
Related Work Skew minimization at multiple corners
[Cho05] perform temperature-aware skew reduction based on an improved DME
[Lung10] minimize the worst clock skew across corners with delay correlation factors
Skew variation minimization across corners
[Restle01] propose two-level non-tree structure, in which mesh is applied at bottom level
[Su01] use mesh for top-level of clock network
[Rajaram04] insert crosslinks in a clock tree to minimize skew variation
Our work: systematic optimization framework for minimization of clock skew variation in clock tree<br>
06
Skew Variation Reduction Problem r: root; i, j: sinks …<br>
07
Outline Motivation
Related Work
Our Optimization Framework
Experimental Setup and Results
Conclusions<br>
08
Our Optimization Framework Incremental optimization of a CTS solution
Perform both global and local optimization
Global optimization uses LP to determine delta delays on arcs
Local optimization performs iterative local moves<br>
09
Global Optimization: LP Formulate linear program to minimize skew variation Determine the delta delay on each arc at each corner Based on LUTs to insert/remove buffer and detour wires
Discreteness of buffer delays ECO feasibility is important (1) Minimize number of ECO changes
(2) Sweep U for solution with minimum skew variation
(3) Ensure no skew degradation
(4) Maximum clock latency constraint
(1, 5, 6) Improve ECO feasibility<br>
10
Our Optimization Framework Incremental optimization of a CTS solution
Perform both global and local optimization
Global optimization use LP to determine delta delays on arcs
Local optimization perform iterative local moves<br>
11
Local Optimization: Moves Iterative local moves to minimize skew variation
Tree types of local moves
Displacement {N, S, E, W, NE, NW, SE, SW} by 10μm x one-step sizing
Displacement by 10μm x one-step sizing on child buffer
Reassign to a new driver (i) at the same level, (ii) within bounding box of 50μm x 50μm Each move is expensive (= legalization, ECO routing, RC extraction, STA)
Each buffer has ~100 candidate moves
Which move is the best? Our solution: learning-based model<br>
12
Machine Learning-Based Model Predict driver-to-fanout latency change due to local moves Local move Each attempt is a local move
114 buffers
45 candidate moves for each buffer
Learning-based model identifies best moves for more buffers with less #attempts<br>
13
Outline Motivation
Related Work
Our Optimization Framework
Experimental Setup and Results
Conclusions<br>
14
Experimental Setup Technology: foundry 28nm LP
Initial clock tree from Synopsys IC Compiler
Testcases: (a) high-speed application processor, (b) memory controller
Corners<br>
15
Experimental Results (1) Up to 22% reduction on sum of skew variation over all sink pairs
No skew degradation at all corners
Negligible area and power overhead<br>
16
Experimental Results (2) Figure shows comparison of skew variation on (a)
Our optimization significantly reduces the large skew variation between corner pairs<br>
17
Outline Motivation
Related Work
Our Optimization Framework
Experimental Setup and Results
Conclusions<br>
18
Conclusion and Future Works First framework to minimize sum of skew variation over all sink pairs in a clock tree
Up to 22% reduction of the sum of skew variation
Future works
Study resultant power and area benefits
Model to predict a buffer location for minimum skew over a continuous range of possible locations Thank You!<br>
19
Backup Slides<br>
20
Experimental Results (3) Figure shows distribution of skew ratios between C0 and C1
Our optimization significantly reduces the variation of skew ratios between corner pairs<br>