Exploring Simple Siamese Representation Learning
Description: Exploring Simple Siamese Representation Learning Computer Aided Medical Procedures Masters Seminar: Deep Learning for Medical Applications by Xinlei Chen and Kaiming He Presented by Joaquín Gómez Sánchez Supervised by Azade Farshad January
Related Topics
Download Presentation
"Exploring Simple Siamese Representation Learning" is the property of its rightful owner. Permission is granted to download and print the materials on this website for personal, non-commercial use only, and to display it on your personal computer provided you do not modify the materials and that you retain all copyright notices contained in the materials. By downloading content from our website, you accept the terms of this agreement.
Presentation Transcript
slide1. Exploring Simple Siamese Representation Learning Computer Aided Medical Procedures
Master’s Seminar: Deep Learning for Medical Applications by Xinlei Chen and Kaiming He Presented by Joaquín Gómez Sánchez
Supervised by Azade Farshad
January 13th, 2022<br>
slide2. Exploring Simple Siamese Representation Learning Introduction<br>
slide3. What are Siamese Networks? January 14, 2022 Computer Aided Medical Procedures Slide 3 Neural Network I Neural Network II Image Shared Weights Distance Representations/Embeddings<br>
slide4. What is the problem with Siamese Networks? Problem: Collapsing solutions
Cause: Attributed to a lack of repulsive component in the optimization objective January 14, 2022 Computer Aided Medical Procedures Slide 4 Training progress<br>
slide5. Proposed solutions January 14, 2022 Computer Aided Medical Procedures Slide 5 Input image(s) Network architecture Pre/post representation Distance Different augmentations Different architectures
Different ways to share the weights W/ & w/o prediction or projection
W/ & w/o stop gradient Different distances<br>
slide6. Exploring Simple Siamese Representation Learning Related Work<br>
slide7. Contrastive Learning Based on: Attraction between positive sample pairs and the repulse of negative pairs.
Loss function example: Contrastive Loss Function (2005) [9]
Other popular contrastive learning loss functions: Noise Contrastive Estimation(2010), Triple Loss (2015), InfoNCE Loss (2018) January 14, 2022 Computer Aided Medical Procedures Slide 7<br>
slide8. Contrastive Learning: SimCLR (Simple framework for Contrastive Learning of visual Reprs.) [2] January 14, 2022 Computer Aided Medical Procedures Slide 8<br>
slide9. Clustering Based on: alternation of
Clustering the representation
Learning to predict the cluster assignments
Important fact: Not defined negative samples → Cluster centres play as negative prototypes January 14, 2022 Computer Aided Medical Procedures Slide 9<br>
slide10. Clustering: SwAV (SWapping Assignments between multiple Views of the same image) [4] January 14, 2022 Computer Aided Medical Procedures Slide 10 ResNet-50 based encoder Image augmentations Representations (projected into unit sphere) Online with Sinkhorn-Knopp alg.<br>
slide11. BYOL (Bootstrap Your Own Latent) [5] January 14, 2022 Computer Aided Medical Procedures Slide 11 ResNet-50 based encoders with Exponential Moving Average) Image augmentations Representations MSE between normalized predictions and target projections<br>
slide12. Exploring Simple Siamese Representation Learning Method<br>
slide13. SimSiam (Simple Siamese) [1] January 14, 2022 Computer Aided Medical Procedures Slide 13 Image augmentations ResNet-50 + Three-layers Projection MLP head encoder Two-layers MLP (bottleneck structure 2048-512-2048) Predictor output compared to encoder output with stop-grad<br>
slide14. SimSiam (Simple Siamese) [1] Loss: January 14, 2022 Computer Aided Medical Procedures Slide 14<br>
slide15. Baseline configuration Optimization:
SGD with cosine decay and schedule learning rate
Base learning rate: 0.05
Momentum: 0.9
Weight decay: 0.0001
Batch size: 256
Architecture aspects:
Encoder:
ResNet-50 + Projection MLP head (3 layers)
BN in each FC layer
Hidden FC of dim = 2.048.
Predictor:
Two layers MLP
BN in its hidden FC layer
Bottleneck structure: 2048-512-2048 January 14, 2022 Computer Aided Medical Procedures Slide 15<br>
slide16. Optimization problem:
Solution: Alternating algorithm An implementation of the Expectation-Maximization (EM)? January 14, 2022 Computer Aided Medical Procedures Slide 16 Network Augmentation<br>
slide17. SimSiam’s architecture vs. other options January 14, 2022 Computer Aided Medical Procedures Slide 17<br>
slide18. Exploring Simple Siamese Representation Learning Experiments and Results<br>
slide19. Why SimSiam’s solutions do not collapse? January 14, 2022 Computer Aided Medical Procedures Slide 19<br>
slide20. Why SimSiam’s solutions do not collapse? January 14, 2022 Computer Aided Medical Procedures Slide 20<br>
slide21. Why SimSiam’s solutions do not collapse? January 14, 2022 Computer Aided Medical Procedures Slide 21<br>
slide22. Comparison – ImageNet Classification January 14, 2022 Computer Aided Medical Procedures Slide 22<br>
slide23. Comparison – Transfer Learning January 14, 2022 Computer Aided Medical Procedures Slide 23<br>
slide24. Exploring Simple Siamese Representation Learning Conclusions & Student’s Review<br>
slide25. Conclusions New simple Siamese method
Model on par or improving previous complex models
Good model for learning representations for transfer learning January 14, 2022 Computer Aided Medical Procedures Slide 25<br>
slide26. Strengths Achieved simple state-of-the-art model
Achieved common architecture on which to make improvements January 14, 2022 Computer Aided Medical Procedures Slide 26 Weaknesses Collapsing solution prevention only shown through empirical observations
Comparison carried out only with ImageNet<br>
slide27. Suggestion for Future Work Application of the acquired improvement in NLP (embeddings)
Theoretical proof of collapsing solution prevention January 14, 2022 Computer Aided Medical Procedures Slide 27 Application to medical data Application of the model in medical applications in which data representations are required, e.g., obtaining image representations to compare while searching for similar ones in an image bank.<br>
slide28. Exploring Simple Siamese Representation Learning References<br>
slide29. [1] Chen, X., & He, K. (2021). Exploring simple siamese representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 15750-15758).
[2] Chen, T., Kornblith, S., Norouzi, M., & Hinton, G. (2020). A simple framework for contrastive learning of visual representations. In International conference on machine learning (pp. 1597-1607). PMLR.
[3] He, K., Fan, H., Wu, Y., Xie, S., & Girshick, R. (2020). Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 9729-9738).
[4] Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., & Joulin, A. (2020). Unsupervised learning of visual features by contrasting cluster assignments. arXiv:2006.09882.
[5] Grill, J. B., Strub, F., Altché, F., Tallec, C., Richemond, P. H., Buchatskaya, E., ... & Valko, M. (2020). Bootstrap your own latent: A new approach to self-supervised learning. arXiv:2006.07733.
[6] He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778).
[7] You, Y., Gitman, I., & Ginsburg, B. (2017). Large batch training of convolutional networks. arXiv:1708.03888.
[8] Deng, J., Dong, W., Socher, R., Li, L. J., Li, K., & Fei-Fei, L. (2009). Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition (pp. 248-255).
[9] Chopra, S., Hadsell, R., & LeCun, Y. (2005, June). Learning a similarity metric discriminatively, with application to face verification. In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05) (Vol. 1, pp. 539-546). IEEE. January 14, 2022 Computer Aided Medical Procedures Slide 29<br>
Master’s Seminar: Deep Learning for Medical Applications by Xinlei Chen and Kaiming He Presented by Joaquín Gómez Sánchez
Supervised by Azade Farshad
January 13th, 2022<br>
slide2. Exploring Simple Siamese Representation Learning Introduction<br>
slide3. What are Siamese Networks? January 14, 2022 Computer Aided Medical Procedures Slide 3 Neural Network I Neural Network II Image Shared Weights Distance Representations/Embeddings<br>
slide4. What is the problem with Siamese Networks? Problem: Collapsing solutions
Cause: Attributed to a lack of repulsive component in the optimization objective January 14, 2022 Computer Aided Medical Procedures Slide 4 Training progress<br>
slide5. Proposed solutions January 14, 2022 Computer Aided Medical Procedures Slide 5 Input image(s) Network architecture Pre/post representation Distance Different augmentations Different architectures
Different ways to share the weights W/ & w/o prediction or projection
W/ & w/o stop gradient Different distances<br>
slide6. Exploring Simple Siamese Representation Learning Related Work<br>
slide7. Contrastive Learning Based on: Attraction between positive sample pairs and the repulse of negative pairs.
Loss function example: Contrastive Loss Function (2005) [9]
Other popular contrastive learning loss functions: Noise Contrastive Estimation(2010), Triple Loss (2015), InfoNCE Loss (2018) January 14, 2022 Computer Aided Medical Procedures Slide 7<br>
slide8. Contrastive Learning: SimCLR (Simple framework for Contrastive Learning of visual Reprs.) [2] January 14, 2022 Computer Aided Medical Procedures Slide 8<br>
slide9. Clustering Based on: alternation of
Clustering the representation
Learning to predict the cluster assignments
Important fact: Not defined negative samples → Cluster centres play as negative prototypes January 14, 2022 Computer Aided Medical Procedures Slide 9<br>
slide10. Clustering: SwAV (SWapping Assignments between multiple Views of the same image) [4] January 14, 2022 Computer Aided Medical Procedures Slide 10 ResNet-50 based encoder Image augmentations Representations (projected into unit sphere) Online with Sinkhorn-Knopp alg.<br>
slide11. BYOL (Bootstrap Your Own Latent) [5] January 14, 2022 Computer Aided Medical Procedures Slide 11 ResNet-50 based encoders with Exponential Moving Average) Image augmentations Representations MSE between normalized predictions and target projections<br>
slide12. Exploring Simple Siamese Representation Learning Method<br>
slide13. SimSiam (Simple Siamese) [1] January 14, 2022 Computer Aided Medical Procedures Slide 13 Image augmentations ResNet-50 + Three-layers Projection MLP head encoder Two-layers MLP (bottleneck structure 2048-512-2048) Predictor output compared to encoder output with stop-grad<br>
slide14. SimSiam (Simple Siamese) [1] Loss: January 14, 2022 Computer Aided Medical Procedures Slide 14<br>
slide15. Baseline configuration Optimization:
SGD with cosine decay and schedule learning rate
Base learning rate: 0.05
Momentum: 0.9
Weight decay: 0.0001
Batch size: 256
Architecture aspects:
Encoder:
ResNet-50 + Projection MLP head (3 layers)
BN in each FC layer
Hidden FC of dim = 2.048.
Predictor:
Two layers MLP
BN in its hidden FC layer
Bottleneck structure: 2048-512-2048 January 14, 2022 Computer Aided Medical Procedures Slide 15<br>
slide16. Optimization problem:
Solution: Alternating algorithm An implementation of the Expectation-Maximization (EM)? January 14, 2022 Computer Aided Medical Procedures Slide 16 Network Augmentation<br>
slide17. SimSiam’s architecture vs. other options January 14, 2022 Computer Aided Medical Procedures Slide 17<br>
slide18. Exploring Simple Siamese Representation Learning Experiments and Results<br>
slide19. Why SimSiam’s solutions do not collapse? January 14, 2022 Computer Aided Medical Procedures Slide 19<br>
slide20. Why SimSiam’s solutions do not collapse? January 14, 2022 Computer Aided Medical Procedures Slide 20<br>
slide21. Why SimSiam’s solutions do not collapse? January 14, 2022 Computer Aided Medical Procedures Slide 21<br>
slide22. Comparison – ImageNet Classification January 14, 2022 Computer Aided Medical Procedures Slide 22<br>
slide23. Comparison – Transfer Learning January 14, 2022 Computer Aided Medical Procedures Slide 23<br>
slide24. Exploring Simple Siamese Representation Learning Conclusions & Student’s Review<br>
slide25. Conclusions New simple Siamese method
Model on par or improving previous complex models
Good model for learning representations for transfer learning January 14, 2022 Computer Aided Medical Procedures Slide 25<br>
slide26. Strengths Achieved simple state-of-the-art model
Achieved common architecture on which to make improvements January 14, 2022 Computer Aided Medical Procedures Slide 26 Weaknesses Collapsing solution prevention only shown through empirical observations
Comparison carried out only with ImageNet<br>
slide27. Suggestion for Future Work Application of the acquired improvement in NLP (embeddings)
Theoretical proof of collapsing solution prevention January 14, 2022 Computer Aided Medical Procedures Slide 27 Application to medical data Application of the model in medical applications in which data representations are required, e.g., obtaining image representations to compare while searching for similar ones in an image bank.<br>
slide28. Exploring Simple Siamese Representation Learning References<br>
slide29. [1] Chen, X., & He, K. (2021). Exploring simple siamese representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 15750-15758).
[2] Chen, T., Kornblith, S., Norouzi, M., & Hinton, G. (2020). A simple framework for contrastive learning of visual representations. In International conference on machine learning (pp. 1597-1607). PMLR.
[3] He, K., Fan, H., Wu, Y., Xie, S., & Girshick, R. (2020). Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 9729-9738).
[4] Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., & Joulin, A. (2020). Unsupervised learning of visual features by contrasting cluster assignments. arXiv:2006.09882.
[5] Grill, J. B., Strub, F., Altché, F., Tallec, C., Richemond, P. H., Buchatskaya, E., ... & Valko, M. (2020). Bootstrap your own latent: A new approach to self-supervised learning. arXiv:2006.07733.
[6] He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778).
[7] You, Y., Gitman, I., & Ginsburg, B. (2017). Large batch training of convolutional networks. arXiv:1708.03888.
[8] Deng, J., Dong, W., Socher, R., Li, L. J., Li, K., & Fei-Fei, L. (2009). Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition (pp. 248-255).
[9] Chopra, S., Hadsell, R., & LeCun, Y. (2005, June). Learning a similarity metric discriminatively, with application to face verification. In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05) (Vol. 1, pp. 539-546). IEEE. January 14, 2022 Computer Aided Medical Procedures Slide 29<br>