Introduction of Machine / Deep Learning Hung-yi

Published  . 0 views
↓ Download
Introduction of Machine / Deep Learning Hung-yi
1 / 1
Introduction of Machine / Deep Learning Hung-yi - slide 1 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 2 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 3 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 4 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 5 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 6 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 7 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 8 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 9 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 10 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 11 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 12 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 13 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 14 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 15 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 16 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 17 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 18 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 19 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 20 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 21 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 22 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 23 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 24 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 25 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 26 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 27 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 28 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 29 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 30 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 31 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 32 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 33 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 34 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 35 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 36 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 37 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 38 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 39 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 40 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 41 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 42 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 43 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 44 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 45 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 46 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 47 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 48 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 49 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 50 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 51 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 52 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 53 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 54 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 55 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 56 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 57 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 58 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 59 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 60 of 61 Introduction of Machine / Deep Learning Hung-yi - slide 61 of 61
Description: Introduction of Machine Deep Learning Hung-yi Lee 李宏毅 Machine Learning Looking for Function Speech Recognition Image Recognition Playing Go Cat How are you 5-5 (next move) Different types of Functions Regression: The function

Related Topics

Download Presentation

"Introduction of Machine / Deep Learning Hung-yi" is the property of its rightful owner. Permission is granted to download and print the materials on this website for personal, non-commercial use only, and to display it on your personal computer provided you do not modify the materials and that you retain all copyright notices contained in the materials. By downloading content from our website, you accept the terms of this agreement.

Presentation Transcript

slide1. Introduction of Machine / Deep Learning Hung-yi Lee 李宏毅<br>
slide2. Machine Learning ≈ Looking for Function Speech Recognition

Image Recognition

Playing Go “Cat” “How are you” “5-5” (next move)<br>
slide3. Different types of Functions Regression: The function outputs a scalar. Predict PM2.5 Spam
filtering Classification: Given options (classes), the function outputs the correct one.<br>
slide4. Different types of Functions Function (19 x 19 classes) Next move Each position is a class a position on the board Classification: Given options (classes), the function outputs the correct one. Playing GO<br>
slide5. Regression, Classification create something with structure (image, document) Structured Learning<br>
slide6. How to find a function? A Case Study<br>
slide7. YouTube Channel https://www.youtube.com/c/HungyiLeeNTU<br>
slide8. The function we want to find … no. of views on 2/26<br>
slide9. 1. Function with Unknown Parameters based on domain knowledge Model weight bias feature<br>
slide10. 2. Define Loss from Training Data Data from 2017/01/01 – 2020/12/31 Loss is a function of parameters Loss: how good a set of values is. 2017/01/01 01/02 12/31 2020/12/30 4.8k 4.9k 3.4k 9.8k 7.5k 01/03 …… 4.9k label 5.3k How good it is?<br>
slide11. 2. Define Loss from Training Data Data from 2017/01/01 – 2020/12/31 2017/01/01 01/02 12/31 2020/12/30 4.8k 4.9k 4.9k 7.5k 9.8k 3.4k 9.8k 7.5k 01/03 …… 5.4k<br>
slide12. 2. Define Loss from Training Data Loss: Cross-entropy<br>
slide13. 2. Define Loss from Training Data Error Surface<br>
slide14. 3. Optimization Compute Positive Negative Decrease w Increase w Source of image: http://chico386.pixnet.net/album/photo/171572850 Gradient Descent<br>
slide15. 3. Optimization Source of image: http://chico386.pixnet.net/album/photo/171572850 Gradient Descent Compute hyperparameters<br>
slide16. 3. Optimization Source of image: http://chico386.pixnet.net/album/photo/171572850 Gradient Descent Local minima global minima Compute Does local minima truly cause the problem?<br>
slide17. 3. Optimization Compute Can be done in one line in most deep learning frameworks<br>
slide18. 3. Optimization<br>
slide19. Machine Learning is so simple ……<br>
slide20. Machine Learning is so simple …… How about data of 2021 (unseen during training)? Training<br>
slide21. 2021/01/01 2021/02/14 Views
(k) Red: real no. of views
blue: estimated no. of views<br>
slide22. 2017 - 2020 2021 2017 - 2020 2017 - 2020 2021 2021 Linear models 2017 - 2020 2021<br>
slide23. Linear models have severe limitation. We need a more flexible model! Model Bias Linear models are too simple … we need more sophisticated modes.<br>
slide24. 0 1 2 3<br>
slide25. All Piecewise Linear Curves<br>
slide26. Beyond Piecewise Linear? To have good approximation, we need sufficient pieces. Approximate continuous curve by a piecewise linear curve.<br>
slide27. How to represent this function? Hard Sigmoid Sigmoid Function<br>
slide28. Change slopes Shift Change height<br>
slide29. sum of a set of + constant 0 2 3 red curve = 1 0<br>
slide30. New Model: More Features<br>
slide31. + + + 1 2 3 no. of features no. of sigmoid<br>
slide33. + + + 1 2 3<br>
slide34. + + + 1 2 3<br>
slide35. + + + 1 2 3 + +<br>
slide36. + + + 1 2 3 + +<br>
slide37. + + + 1 2 3 + + +<br>
slide38. …… + + feature Unknown parameters Function with unknown parameters<br>
slide39. Back to ML Framework<br>
slide40. Loss Loss: feature Loss is a function of parameters Loss means how good a set of values is. Given a set of values label<br>
slide41. Back to ML Framework<br>
slide42. Optimization of New Model gradient<br>
slide43. Optimization of New Model<br>
slide44. Optimization of New Model N B batch batch batch batch 1 epoch = see all the batches once update update update<br>
slide45. Optimization of New Model B batch batch batch batch Example 1 10,000 examples (N = 10,000)
Batch size is 10 (B = 10) How many update in 1 epoch? 1,000 updates Example 2 1,000 examples (N = 1,000)
Batch size is 100 (B = 100) How many update in 1 epoch? 10 updates N<br>
slide46. Back to ML Framework More variety of models …<br>
slide47. How to represent this function? Rectified Linear Unit (ReLU)<br>
slide48. Which one is better? Activation function<br>
slide49. Experimental Results<br>
slide50. Back to ML Framework Even more variety of models …<br>
slide51. + + + …… + + +<br>
slide52. Experimental Results Loss for multiple hidden layers
100 ReLU for each layer
input features are the no. of views in the past 56 days<br>
slide53. 2021/01/01 2021/02/14 3 layers Views
(k) Red: real no. of views
blue: estimated no. of views ?<br>
slide54. Back to ML Framework It is not fancy enough. Let’s give it a fancy name!<br>
slide55. + + + …… + + + Neuron Neural Network This mimics human brains … (???) Many layers means Deep hidden layer hidden layer Deep Learning<br>
slide56. 8 layers 19 layers 22 layers AlexNet (2012) VGG (2014) GoogleNet (2014) 16.4% 7.3% 6.7% http://cs231n.stanford.edu/slides/winter1516_lecture8.pdf Deep = Many hidden layers<br>
slide57. AlexNet (2012) VGG
(2014) GoogleNet
(2014) 152 layers 3.57% Residual Net
(2015) Taipei
101 101 layers 16.4% 7.3% 6.7% Deep = Many hidden layers Special
structure Why we want “Deep” network, not “Fat” network?<br>
slide58. Why don’t we go deeper? Loss for multiple hidden layers
100 ReLU for each layer
input features are the no. of views in the past 56 days<br>
slide59. Why don’t we go deeper? Loss for multiple hidden layers
100 ReLU for each layer
input features are the no. of views in the past 56 days Better on training data, worse on unseen data Overfitting<br>
slide60. Let’s predict no. of views today! If we want to select a model for predicting no. of views today, which one will you use? We will talk about model selection next time. <br>
slide61. To learn more …… https://youtu.be/Dr-WRlEFefw https://youtu.be/ibJpTrp5mcE Basic Introduction Backpropagation Computing gradients in an efficient way<br>