Deep Learning for Natural Language Processing

Published  . 0 views
↓ Download
Deep Learning for Natural Language Processing
1 / 1
Deep Learning for Natural Language Processing - slide 1 of 64 Deep Learning for Natural Language Processing - slide 2 of 64 Deep Learning for Natural Language Processing - slide 3 of 64 Deep Learning for Natural Language Processing - slide 4 of 64 Deep Learning for Natural Language Processing - slide 5 of 64 Deep Learning for Natural Language Processing - slide 6 of 64 Deep Learning for Natural Language Processing - slide 7 of 64 Deep Learning for Natural Language Processing - slide 8 of 64 Deep Learning for Natural Language Processing - slide 9 of 64 Deep Learning for Natural Language Processing - slide 10 of 64 Deep Learning for Natural Language Processing - slide 11 of 64 Deep Learning for Natural Language Processing - slide 12 of 64 Deep Learning for Natural Language Processing - slide 13 of 64 Deep Learning for Natural Language Processing - slide 14 of 64 Deep Learning for Natural Language Processing - slide 15 of 64 Deep Learning for Natural Language Processing - slide 16 of 64 Deep Learning for Natural Language Processing - slide 17 of 64 Deep Learning for Natural Language Processing - slide 18 of 64 Deep Learning for Natural Language Processing - slide 19 of 64 Deep Learning for Natural Language Processing - slide 20 of 64 Deep Learning for Natural Language Processing - slide 21 of 64 Deep Learning for Natural Language Processing - slide 22 of 64 Deep Learning for Natural Language Processing - slide 23 of 64 Deep Learning for Natural Language Processing - slide 24 of 64 Deep Learning for Natural Language Processing - slide 25 of 64 Deep Learning for Natural Language Processing - slide 26 of 64 Deep Learning for Natural Language Processing - slide 27 of 64 Deep Learning for Natural Language Processing - slide 28 of 64 Deep Learning for Natural Language Processing - slide 29 of 64 Deep Learning for Natural Language Processing - slide 30 of 64 Deep Learning for Natural Language Processing - slide 31 of 64 Deep Learning for Natural Language Processing - slide 32 of 64 Deep Learning for Natural Language Processing - slide 33 of 64 Deep Learning for Natural Language Processing - slide 34 of 64 Deep Learning for Natural Language Processing - slide 35 of 64 Deep Learning for Natural Language Processing - slide 36 of 64 Deep Learning for Natural Language Processing - slide 37 of 64 Deep Learning for Natural Language Processing - slide 38 of 64 Deep Learning for Natural Language Processing - slide 39 of 64 Deep Learning for Natural Language Processing - slide 40 of 64 Deep Learning for Natural Language Processing - slide 41 of 64 Deep Learning for Natural Language Processing - slide 42 of 64 Deep Learning for Natural Language Processing - slide 43 of 64 Deep Learning for Natural Language Processing - slide 44 of 64 Deep Learning for Natural Language Processing - slide 45 of 64 Deep Learning for Natural Language Processing - slide 46 of 64 Deep Learning for Natural Language Processing - slide 47 of 64 Deep Learning for Natural Language Processing - slide 48 of 64 Deep Learning for Natural Language Processing - slide 49 of 64 Deep Learning for Natural Language Processing - slide 50 of 64 Deep Learning for Natural Language Processing - slide 51 of 64 Deep Learning for Natural Language Processing - slide 52 of 64 Deep Learning for Natural Language Processing - slide 53 of 64 Deep Learning for Natural Language Processing - slide 54 of 64 Deep Learning for Natural Language Processing - slide 55 of 64 Deep Learning for Natural Language Processing - slide 56 of 64 Deep Learning for Natural Language Processing - slide 57 of 64 Deep Learning for Natural Language Processing - slide 58 of 64 Deep Learning for Natural Language Processing - slide 59 of 64 Deep Learning for Natural Language Processing - slide 60 of 64 Deep Learning for Natural Language Processing - slide 61 of 64 Deep Learning for Natural Language Processing - slide 62 of 64 Deep Learning for Natural Language Processing - slide 63 of 64 Deep Learning for Natural Language Processing - slide 64 of 64
Description: Deep Learning for Natural Language Processing DLNLP 5: Feed Forward Neural Networks Mihai Surdeanu and Marco A. Valenzuela-Escarcega 1 Overview Architecture of feed-forward neural networks (FFNNs) Learning algorithm for FFNNs Equations of

Related Topics

Download Presentation

"Deep Learning for Natural Language Processing" is the property of its rightful owner. Permission is granted to download and print the materials on this website for personal, non-commercial use only, and to display it on your personal computer provided you do not modify the materials and that you retain all copyright notices contained in the materials. By downloading content from our website, you accept the terms of this agreement.

Presentation Transcript

slide1. Deep Learning for Natural Language Processing DLNLP 5: Feed Forward Neural Networks Mihai Surdeanu and Marco A. Valenzuela-Escárcega 1<br>
slide2. Overview Architecture of feed-forward neural networks (FFNNs)
Learning algorithm for FFNNs
Equations of back-propagation 2<br>
slide3. Architecture 3<br>
slide4. Drawbacks of logistic regression Linear classifier
Feed-forward neural networks (FFNNs) can approximate any function! 4<br>
slide5. Example of non-linear classifier 5<br>
slide6. Drawbacks of logistic regression Linear classifier
Feed-forward neural networks (FFNNs) can approximate any function!
Hand-crafted features
We will address this limitation when we discuss word embeddings (chapter 8) 6<br>
slide7. Reminder: perceptron/logistic regression 7<br>
slide8. FFNN architecture 8<br>
slide9. Input layer 9<br>
slide10. Input layer Each input vector x describes one data point
Collection of explicit features (what we have seen so far), or
A numerical representation of the input (based on word embeddings, chapter 8) 10<br>
slide11. Intermediate layers 11<br>
slide12. Intermediate layers 12<br>
slide13. Intermediate layers 13 Output (or activation) of this neuron<br>
slide14. Intermediate layers 14 Inputs to this neuron, i.e., activations of neurons in the previous layer<br>
slide15. Intermediate layers 15 Weights of the edges connecting layer l - 1 with layer l<br>
slide16. Intermediate layers 16 Nonlinear function such as sigmoid (or tanh, ReLU)<br>
slide17. Output layer 17<br>
slide18. Output layer One neuron per class to be learned
For example, we need 3 neurons to learn a review classifier that outputs: positive, neutral, negative.
Output scores may be converted to a probability distribution using the softmax function
Softmax: takes a bunch of values and converts them to values in the [0, 1] interval such that they sum up to 1. 18<br>
slide19. FFNNs can be linear functions (when f is a pass-through function: f(x) = x) 19<br>
slide20. FFNNs can be linear functions (when f is a pass-through function: f(x) = x) 20<br>
slide21. Tensor notation Each layer can be described by this simpler equation that relies on matrices and vectors:

Remember:
Lower case, bold font = vector
Upper case, bold font = matrix 21<br>
slide22. Tensor notation Each layer can be described by this simpler equation that relies on matrices and vectors: 22 All activations in layer l All activations in layer l - 1<br>
slide23. Tensor notation Each layer can be described by this simpler equation that relies on matrices and vectors: 23 All weights connecting layer l – 1 with layer l<br>
slide24. FFNNs are generalizations of the perceptron and LR Perceptron
No intermediate layers
f(x) = x
Single output neuron; no softmax
Binary logistic regression
No intermediate layers
f(x) = sigmoid(x)
Single output neuron; no softmax
Multiclass logistic regression
No intermediate layers
f(x) = x
Multiple output neurons with softmax 24<br>
slide25. Learning algorithm 25<br>
slide26. Learning algorithm intuition Gradient descent, exactly the same as the one used for logistic regression!
“Knob turning”
”knob” = weight of an edge
If a neuron increases the probability of an incorrect prediction, its knobs will be turned down.
If a neuron increases the probability of a correct prediction, its knobs will be turned up. 26<br>
slide27. Intuition See this video until minute 9:30: https://www.youtube.com/watch?v=Ilg3gGewQ5U 27<br>
slide28. 28<br>
slide29. 29 Collection of all weights and biases in the network<br>
slide30. 30 One training example<br>
slide31. 31 Partial derivative of the cost function C for each parameter (weight or bias) in the network<br>
slide32. 32 Learning rate, which is a hyper parameter<br>
slide33. 33 But how do we efficiently compute these partial derivatives, especially for edges that are not directly connected to the output neurons? Back-propagation to the rescue!<br>
slide34. Intuition 34<br>
slide35. Equations of back-propagation 35<br>
slide36. Back-propagation intuition Recursively propagate neuron contributions to the output errors from the last layer back to the first 36<br>
slide37. Notations Notation simplification: denote Ci(Θ) as C

Error of neuron i in layer l: 37 The error of a neuron measures what impact a small change in its output z has on the cost C

Why use z instead of a?<br>
slide38. Notations Notation simplification: denote Ci(Θ) as C

Error of neuron i in layer l:
This is the key “knob”
The higher the error, the larger the parameter adjustments for this neuron 38<br>
slide39. Notations Notation simplification: denote Ci(Θ) as C

Error of neuron i in layer l:
This is the key “knob”
The higher the error, the larger the parameter adjustments for this neuron
L: index of the final network layer. Thus, is the error of neuron i in the last layer. 39<br>
slide40. The equations of back-propagation Compute the error of the neurons in the last layer (L)
From right-to-left, compute the errors of the neurons in all layers l, using the errors of the neurons in layer l + 1
Adjust the edge weights and biases in the entire network using these errors 40<br>
slide41. Warning: heaviest math in this book coming up! 41<br>
slide42. Backprop equation 1: Error of a neuron in the final layer In general:

This yields the following formula for the last layer: 42<br>
slide43. Proof of backprop equation 1 43<br>
slide44. Proof of backprop equation 1 44 Must loop over all k neurons in the last layer because C depends on all activations in the last layer. However, neuron i impacts only its own activation, so we can ignore the others.<br>
slide45. Proof of backprop equation 1 45 Chain rule!<br>
slide46. Example of equation 1: Binary logistic regression with MSE cost Mean squared error (MSE) cost:
Derivatives we need:

Then: 46<br>
slide47. Example of equation 1: Binary logistic regression with MSE cost 47 This is a good “knob.” For example, when the gold label y is 1 and the neuron output (s) is large, the error will be small. When the neuron is “confused,” e.g., its output (s) is 0.5, the error will be large.<br>
slide48. Example of equation 1: Binary logistic regression with MSE cost 48 Peeking ahead:
What’s the problem here?<br>
slide49. Backprop equation 2: Error of a neuron in an intermediate layer 49 Recursive back-propagation and quick to compute!

What type of programming is this?<br>
slide50. Visual helper for backprop equation 2 50<br>
slide51. Proof of backprop equation 2 51<br>
slide52. Proof of backprop equation 2 52 Chain rule again. We need to use all neurons in the next layer because they are all impacted by the current neuron<br>
slide53. Proof of backprop equation 2 53 Just the definition of<br>
slide54. Proof of backprop equation 2 54 Connections from neuron j to k can be ignored for i != j<br>
slide55. Proof of backprop equation 2 55 Chain rule again<br>
slide56. Backprop equation 3: Update for any bias term 56<br>
slide57. Proof for backprop equation 3 57 What rules have we applied here?<br>
slide58. Backprop equation 4 58<br>
slide59. Proof of backprop equation 4 59 What rules have we applied here?<br>
slide60. The 4 equations of back-propagation 60<br>
slide61. Math over! 61<br>
slide62. Wait a minute… There are auto-differentiation libraries nowadays that can compute the partial derivative wrt to any parameter directly... Why do we need to go through all this trouble?
What type of algorithm is back-propagation? 62<br>
slide63. Drawbacks of feed-forward neural networks Back-propagation is slow, especially for networks with >> parameters.
The brain probably does not learn using back-propagation (e.g., if it did, our vision would regularly black out during updates)
FFNNs are very powerful classifiers (i.e., can learn any function), but this can lead to overfitting (i.e., “hallucinating” a classifier)
So far, we discussed FFNNs that rely on hand-crafted features 63<br>
slide64. Take away Architecture of feed-forward neural networks (FFNNs)
Learning algorithm for FFNNs
Equations of back-propagation 64<br>