Algorithmic Improvements for Fast Concurrent

Published  . 0 views
↓ Download
Algorithmic Improvements for Fast Concurrent
1 / 1
Algorithmic Improvements for Fast Concurrent - slide 1 of 43 Algorithmic Improvements for Fast Concurrent - slide 2 of 43 Algorithmic Improvements for Fast Concurrent - slide 3 of 43 Algorithmic Improvements for Fast Concurrent - slide 4 of 43 Algorithmic Improvements for Fast Concurrent - slide 5 of 43 Algorithmic Improvements for Fast Concurrent - slide 6 of 43 Algorithmic Improvements for Fast Concurrent - slide 7 of 43 Algorithmic Improvements for Fast Concurrent - slide 8 of 43 Algorithmic Improvements for Fast Concurrent - slide 9 of 43 Algorithmic Improvements for Fast Concurrent - slide 10 of 43 Algorithmic Improvements for Fast Concurrent - slide 11 of 43 Algorithmic Improvements for Fast Concurrent - slide 12 of 43 Algorithmic Improvements for Fast Concurrent - slide 13 of 43 Algorithmic Improvements for Fast Concurrent - slide 14 of 43 Algorithmic Improvements for Fast Concurrent - slide 15 of 43 Algorithmic Improvements for Fast Concurrent - slide 16 of 43 Algorithmic Improvements for Fast Concurrent - slide 17 of 43 Algorithmic Improvements for Fast Concurrent - slide 18 of 43 Algorithmic Improvements for Fast Concurrent - slide 19 of 43 Algorithmic Improvements for Fast Concurrent - slide 20 of 43 Algorithmic Improvements for Fast Concurrent - slide 21 of 43 Algorithmic Improvements for Fast Concurrent - slide 22 of 43 Algorithmic Improvements for Fast Concurrent - slide 23 of 43 Algorithmic Improvements for Fast Concurrent - slide 24 of 43 Algorithmic Improvements for Fast Concurrent - slide 25 of 43 Algorithmic Improvements for Fast Concurrent - slide 26 of 43 Algorithmic Improvements for Fast Concurrent - slide 27 of 43 Algorithmic Improvements for Fast Concurrent - slide 28 of 43 Algorithmic Improvements for Fast Concurrent - slide 29 of 43 Algorithmic Improvements for Fast Concurrent - slide 30 of 43 Algorithmic Improvements for Fast Concurrent - slide 31 of 43 Algorithmic Improvements for Fast Concurrent - slide 32 of 43 Algorithmic Improvements for Fast Concurrent - slide 33 of 43 Algorithmic Improvements for Fast Concurrent - slide 34 of 43 Algorithmic Improvements for Fast Concurrent - slide 35 of 43 Algorithmic Improvements for Fast Concurrent - slide 36 of 43 Algorithmic Improvements for Fast Concurrent - slide 37 of 43 Algorithmic Improvements for Fast Concurrent - slide 38 of 43 Algorithmic Improvements for Fast Concurrent - slide 39 of 43 Algorithmic Improvements for Fast Concurrent - slide 40 of 43 Algorithmic Improvements for Fast Concurrent - slide 41 of 43 Algorithmic Improvements for Fast Concurrent - slide 42 of 43 Algorithmic Improvements for Fast Concurrent - slide 43 of 43
Description: Algorithmic Improvements for Fast Concurrent Cuckoo Hashing Xiaozhou Li (Princeton) David G. Andersen (CMU) Michael Kaminsky (Intel Labs) Michael J. Freedman (Princeton) How to build a fast concurrent hash table algorithm and data structure

Related Topics

Download Presentation

"Algorithmic Improvements for Fast Concurrent" is the property of its rightful owner. Permission is granted to download and print the materials on this website for personal, non-commercial use only, and to display it on your personal computer provided you do not modify the materials and that you retain all copyright notices contained in the materials. By downloading content from our website, you accept the terms of this agreement.

Presentation Transcript

slide1. Algorithmic Improvements for Fast Concurrent Cuckoo Hashing Xiaozhou Li (Princeton)
David G. Andersen (CMU)
Michael Kaminsky (Intel Labs)
Michael J. Freedman (Princeton)<br>
slide2. How to build a fast concurrent hash table
algorithm and data structure engineering
Experience with hardware transactional memory
does NOT obviate the need for algorithmic optimizations In this talk<br>
slide3. Concurrent hash table Indexing key-value objects
Lookup(key)
Insert(key, value)
Delete(key)
Fundamental building block for modern systems
System applications (e.g., kernel caches)
Concurrent user-level applications
Targeted workloads: small objects, high rate<br>
slide4. Goal: memory-efficient and high-throughput Memory efficient (e.g., > 90% space utilized)
Fast concurrent reads (scale with # of cores)
Fast concurrent writes (scale with # of cores)<br>
slide5. Preview our results on a quad-core machine Throughput (million reqs per sec) 64-bit key and 64-bit value
120 million objects, 100% Insert ★ ★ cuckoo+ uses (less than) half of the memory compared to others<br>
slide6. Background: separate chaining hash table Chaining items hashed in same bucket lookup Good: simple
Bad: poor cache locality
Bad: pointers cost space
e.g., Intel TBB concurrent_hash_map<br>
slide7. Background: open addressing hash table Probing alternate locations for vacancy
e.g., linear/quadratic probing, double hashing lookup Good: cache friendly
Bad: poor memory efficiency performance dramatically degrades when the usage grows beyond 70% capacity or so
e.g., Google dense_hash_map wastes 50% memory by default.<br>
slide8. Our starting point Multi-reader single-writer cuckoo hashing [Fan, NSDI’13]
Open addressing
Memory efficient
Optimized for read-intensive workloads<br>
slide9. Each bucket has b slots for items (b-way set-associative)
Each key is mapped to two random buckets
stored in one of them buckets key x hash1(x) hash2(x) Cuckoo hashing<br>
slide10. Predictable and fast lookup Lookup: read 2 buckets in parallel
constant time in the worst case Lookup x<br>
slide11. move keys to alternate buckets Insert may need “cuckoo move” Both are full? Insert y a x b k r c s e n f x a b Insert: Write to an empty slot in one of the two buckets possible locations x a b possible locations possible locations<br>
slide12. Insert: move keys to alternate buckets
find a “cuckoo path” to an empty slot
move hole backwards Insert y x a b y Insert may need “cuckoo move” A technique in [Fan, NSDI’13]
No reader/writer false misses b a x<br>
slide13. Review our starting point [Fan, NSDI’13]: Multi-reader single-writer cuckoo hashing Benefits
support concurrent reads
memory efficient for small objects
over 90% space utilized when set-associativity ≥ 4
Limits
Inserts are serialized
poor performance for write-heavy workloads<br>
slide14. Improve write concurrency Algorithmic optimizations
Minimize critical sections
Exploit data locality
Explore two concurrency control mechanisms
Hardware transactional memory
Fine-grained locking<br>
slide15. Algorithmic optimizations Lock after discovering a cuckoo path
minimize critical sections
Breadth-first search for an empty slot
fewer items displaced
enable prefetching
Increase set-associativity (see paper)
fewer random memory reads<br>
slide16. Previous approach: writer locks the table during the whole insert process lock();
Search for a cuckoo path;
Cuckoo move and insert;
unlock(); // at most hundreds of bucket reads // at most hundreds of writes All Insert operations of other threads are blocked<br>
slide17. Lock after discovering a cuckoo path while(1) {
Search for a cuckoo path;
lock();
Cuckoo move and insert;
if(success)
unlock();
break;
unlock();
} // no locking required Multiple Insert threads can look for cuckoo paths concurrently ←collision Cuckoo move and insert while the path is valid; unlock();<br>
slide18. Cuckoo hash table ⟹ undirected cuckoo graph ⟹ x z y 0 a b c 3 1 7 6 9 a x y b z c bucket ⟶ vertex
key ⟶ edge<br>
slide19. Previous approach to search for an empty slot: random walk on the cuckoo graph One Insert may move at most hundreds of items when table occupancy > 90% cuckoo path:
a➝e➝s➝x➝k➝f➝d➝t➝∅
9 writes Insert y<br>
slide20. Insert y Breadth-first search for an empty slot Insert y<br>
slide21. Insert y Breadth-first search for an empty slot cuckoo path:
a➝z➝u➝∅ 4 writes Reduced to a logarithmic factor Same # of reads
Far fewer writes ⟶ unlocked
⟶ locked Prefetching: scan one bucket and load next bucket concurrently<br>
slide22. Concurrency control Fine-grained locking
spinlock and lock striping
Hardware transactional memory
Intel Transactional Synchronization Extensions (TSX)
Hardware support for lock elision<br>
slide23. Lock elision acquire acquire release release critical section critical section Thread 1 Thread 2 Time Hash Table Lock: Free No serialization if no data conflicts<br>
slide24. Implement lock elision with Intel TSX -- Abort reasons:
data conflicts
limited HW resources
unfriendly instructions execute success abort LOCK START TX Critical Section COMMIT UNLOCK fallback retry ? optimized to make better decisions<br>
slide25. Principles to reduce transactional aborts Minimize the size of transactional regions.
Algorithmic optimizations
lock later, BFS, increase set-associativity cuckoo search: 500 reads
cuckoo move: 250 writes —
cuckoo move: 5 writes/reads<br>
slide26. Principles to reduce transactional aborts Avoid unnecessary access to common data.
Make globals thread-local
Avoid TSX-unfriendly instructions in transactions
e.g., malloc()may cause problems
Optimize TSX lock elision implementation
Elide the lock more aggressively for short transactions<br>
slide27. Evaluation How does the performance scale?
throughput vs. # of cores
How much each technique improves performance?
algorithmic optimizations
lock elision with Intel TSX<br>
slide28. Experiment settings Platform
Intel Haswell i7-4770 @ 3.4GHz (with TSX support)
4 cores (8 hyper-threaded cores)
Cuckoo hash table
8 byte keys and 8 byte values
2 GB hash table, ~134.2 million entries
8-way set-associative
Workloads
Fill an empty table to 95% capacity
Random mixed reads and writes<br>
slide29. Multi-core scaling comparison (50% Insert) cuckoo: single-writer/multi-reader [Fan, NSDI’13]
cuckoo+: cuckoo with our algorithmic optimizations Number of threads Throughput (MOPS)<br>
slide30. Multi-core scaling comparison (50% Insert) cuckoo: single-writer/multi-reader [Fan, NSDI’13]
cuckoo+: cuckoo with our algorithmic optimizations Number of threads Throughput (MOPS)<br>
slide31. Multi-core scaling comparison (50% Insert) cuckoo: single-writer/multi-reader [Fan, NSDI’13]
cuckoo+: cuckoo with our algorithmic optimizations Number of threads Throughput (MOPS)<br>
slide32. Multi-core scaling comparison (50% Insert) cuckoo: single-writer/multi-reader [Fan, NSDI’13]
cuckoo+: cuckoo with our algorithmic optimizations Number of threads Throughput (MOPS)<br>
slide33. Multi-core scaling comparison (50% Insert) cuckoo: single-writer/multi-reader [Fan, NSDI’13]
cuckoo+: cuckoo with our algorithmic optimizations Number of threads Throughput (MOPS)<br>
slide34. Multi-core scaling comparison (10% Insert) cuckoo: single-writer/multi-reader [Fan, NSDI’13]
cuckoo+: cuckoo with our algorithmic optimizations Number of threads Throughput (MOPS)<br>
slide35. Multi-core scaling comparison (10% Insert) cuckoo: single-writer/multi-reader [Fan, NSDI’13]
cuckoo+: cuckoo with our algorithmic optimizations Number of threads Throughput (MOPS)<br>
slide36. Multi-core scaling comparison (10% Insert) cuckoo: single-writer/multi-reader [Fan, NSDI’13]
cuckoo+: cuckoo with our algorithmic optimizations Number of threads Throughput (MOPS)<br>
slide37. Multi-core scaling comparison (10% Insert) cuckoo: single-writer/multi-reader [Fan, NSDI’13]
cuckoo+: cuckoo with our algorithmic optimizations Number of threads Throughput (MOPS)<br>
slide38. Multi-core scaling comparison (10% Insert) cuckoo: single-writer/multi-reader [Fan, NSDI’13]
cuckoo+: cuckoo with our algorithmic optimizations Number of threads Throughput (MOPS)<br>
slide39. Factor analysis of Insert performance cuckoo: multi-reader single-writer cuckoo hashing [Fan, NSDI’13]
+TSX-glibc: use released Intel glibc TSX lock elision
+TSX*: replace TSX-glibc with our optimized implementation
+lock later: lock after discovering a cuckoo path
+BFS: breadth first search for an empty slot<br>
slide40. Lock elision enabled first and algorithmic optimizations applied later 100% Insert<br>
slide41. Algorithmic optimizations applied first and lock elision enabled later 100% Insert Both data structure and concurrency control optimizations are needed to achieve high performance<br>
slide42. Conclusion Concurrent cuckoo hash table
high memory efficiency
fast concurrent writes and reads
Lessons with hardware transactional memory
algorithmic optimizations are necessary<br>
slide43. Q & A Source code available: github.com/efficient/libcuckoo
fine-grained locking implementation Thanks!<br>