Artificial Neural Networks What are Artificial

Published  . 0 views
↓ Download
Artificial Neural Networks What are Artificial
1 / 1
Artificial Neural Networks What are Artificial - slide 1 of 29 Artificial Neural Networks What are Artificial - slide 2 of 29 Artificial Neural Networks What are Artificial - slide 3 of 29 Artificial Neural Networks What are Artificial - slide 4 of 29 Artificial Neural Networks What are Artificial - slide 5 of 29 Artificial Neural Networks What are Artificial - slide 6 of 29 Artificial Neural Networks What are Artificial - slide 7 of 29 Artificial Neural Networks What are Artificial - slide 8 of 29 Artificial Neural Networks What are Artificial - slide 9 of 29 Artificial Neural Networks What are Artificial - slide 10 of 29 Artificial Neural Networks What are Artificial - slide 11 of 29 Artificial Neural Networks What are Artificial - slide 12 of 29 Artificial Neural Networks What are Artificial - slide 13 of 29 Artificial Neural Networks What are Artificial - slide 14 of 29 Artificial Neural Networks What are Artificial - slide 15 of 29 Artificial Neural Networks What are Artificial - slide 16 of 29 Artificial Neural Networks What are Artificial - slide 17 of 29 Artificial Neural Networks What are Artificial - slide 18 of 29 Artificial Neural Networks What are Artificial - slide 19 of 29 Artificial Neural Networks What are Artificial - slide 20 of 29 Artificial Neural Networks What are Artificial - slide 21 of 29 Artificial Neural Networks What are Artificial - slide 22 of 29 Artificial Neural Networks What are Artificial - slide 23 of 29 Artificial Neural Networks What are Artificial - slide 24 of 29 Artificial Neural Networks What are Artificial - slide 25 of 29 Artificial Neural Networks What are Artificial - slide 26 of 29 Artificial Neural Networks What are Artificial - slide 27 of 29 Artificial Neural Networks What are Artificial - slide 28 of 29 Artificial Neural Networks What are Artificial - slide 29 of 29
Description: Artificial Neural Networks What are Artificial Neural Networks (ANN)? Colored neural network by Glosser.ca - Own work, Derivative of File:Artificial neural network.svg. Licensed under CC BY-SA 3.0 via Commons -

Related Topics

Download Presentation

"Artificial Neural Networks What are Artificial" is the property of its rightful owner. Permission is granted to download and print the materials on this website for personal, non-commercial use only, and to display it on your personal computer provided you do not modify the materials and that you retain all copyright notices contained in the materials. By downloading content from our website, you accept the terms of this agreement.

Presentation Transcript

slide1. Artificial Neural Networks<br>
slide2. What are Artificial Neural Networks (ANN)? "Colored neural network" by Glosser.ca - Own work, Derivative of File:Artificial neural network.svg. Licensed under CC BY-SA 3.0 via Commons - https://commons.wikimedia.org/wiki/File:Colored_neural_network.svg#/media/File:Colored_neural_network.svg<br>
slide3. Why ANN?<br>
slide4. Why ANN?<br>
slide5. Why ANN?<br>
slide6. Why ANN? Nature of target function is unknown.

Interpretability of function is not important.

Slow training time is ok.<br>
slide7. Perceptron<br>
slide8. Perceptron Model of a real neuron?<br>
slide9. LMS / Delta Rule for Learning a Perceptron model<br>
slide10. LMS / Delta Rule for Learning a Perceptron model Learning rate<br>
slide11. Demo of Simple Synthetic Dataset<br>
slide12. Demo of Simple Synthetic Dataset Convex problem! Guaranteed convergence!<br>
slide13. Problems with Perceptron ANN Only works for linearly separable data.
Solution? – multi layer

Very large terabyte dataset. A single gradient computation will take days 
Solution? – Stochastic Gradient Descent<br>
slide14. Stochastic Gradient Descent (SGD)<br>
slide15. Non-Linear Decision Boundary? Derivative of sigmoid?<br>
slide16. Questions about the Sigmoid Unit? How do we connect the neurons? For this lesson, linear chain – multilayer feedforward.
Outside this lesson:- Pretty much anything you like

How do we train? Backpropagation algorithm<br>
slide17. Backpropagation Algorithm Each layer does two things
Compute the derivative of E w.r.t. its parameters.
Why?

Compute the derivative of E w.r.t. its input.
The reason for this will be obvious when we do it.<br>
slide18. Dealing with Vector Data Partial derivatives change to gradients.

Scalar multiplication changes to vector matrix products or sometimes even tensor vector products.<br>
slide19. Problems Sigmoid units – many of them – vanishing gradients
ReLU units, pretraining using unsupervised learning.

Local optimum – non convex problem
Momentum, SGD, small initialization

Overfitting
Use validation data for early stopping, weight decay.

Lots of parameter tuning
Use several thousand computers to try several parameters and pick the best.

Lack of Interpretability
Do a D.Phil like me trying to interpret neurons in hidden layers.<br>
slide20. Demo on Face Pose Estimation<br>
slide21. Demo on Face Pose Estimation Input representation
Downsample image and divide by 255.

Output representation
1 of 4 encoding

Other learning parameters
Learning rate – 0.3, momentum – 0
Single sample SGD. Let’s see the code<br>
slide22. Demo on Face Pose Estimation Layer 2 weights Layer 1 weights Left Right Up Straight<br>
slide23. Expressive Power Two layers of sigmoid units – any Boolean function.

Two layer network with sigmoid units in the hidden layer and (unthresholded) linear units in the output layer - Any bounded continuous function. (Cybenko 1989, Hornik et. al. 1989)

A network of three layers, where the output layer again has linear units - Any function. (Cybenko 1988).

So multi layer Sigmoid Units are the ultimate supervised learning thing - right? Nope <br>
slide24. Deep Learning Sigmoid ANNs need to be very fat.

Instead we can go deep and thin. But then we have vanishing gradients!

Use ReLUs.<br>
slide25. Still too Many Parameters 1 Megapixel image over 1000 categories. A single layer network will itself need 1 billion parameters.

Convolutional Neural Networks help us scale to large images with very few parameters.<br>
slide26. Convolutional Neural Network<br>
slide27. Benefits of CNNs The number of weights is now much less than 1 million for a 1 mega pixel image.
The small number of weights can use different parts of the image as training data. Thus we have several orders of magnitude more data to train the fewer number of weights.
We get translation invariance for free.
Fewer parameters take less memory and thus all the computations can be carried out in memory in a GPU or across multiple processors.<br>
slide28. Thank you Feel free to email me your questions at aravindh.mahendran@new.ox.ac.uk Strongly recommend this book for basics<br>
slide29. References Cybenko 1989 - https://www.dartmouth.edu/~gvc/Cybenko_MCSS.pdf
Cybenko 1988 – Continuous Valued Neural Networks with two Hidden Layers are Sufficient (Technical Report), Department of Computer Science, Tufts University, Medford, MA
Fukushima 1980 - http://www.cs.princeton.edu/courses/archive/spr08/cos598B/Readings/Fukushima1980.pdf
Hinton 2006 - http://www.cs.toronto.edu/~fritz/absps/ncfast.pdf
Hornick et. al. 1989 - http://www.sciencedirect.com/science/article/pii/0893608089900208
Krizhevsky et. al. 2012 - http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf
Lecun 1998 - http://yann.lecun.com/exdb/publis/pdf/lecun-98.pdf
Tom Mitchell, Machine Learning, 1997<br>