Bootstrap – The Statistician’s Magic Wand Saharon
Description: Bootstrap The Statisticians Magic Wand Saharon Rosset Bootstrap is a paradigm, not an algorithm An abstract view of statistics There is a world (unknown distribution) F We observe some data from the world, say 100 heights (z) and
Related Topics
Download Presentation
"Bootstrap – The Statistician’s Magic Wand Saharon" is the property of its rightful owner. Permission is granted to download and print the materials on this website for personal, non-commercial use only, and to display it on your personal computer provided you do not modify the materials and that you retain all copyright notices contained in the materials. By downloading content from our website, you accept the terms of this agreement.
Presentation Transcript
slide1. Bootstrap – The Statistician’s Magic Wand Saharon Rosset<br>
slide2. Bootstrap is a paradigm, not an algorithm<br>
slide3. An abstract view of statistics There is a “world” (=unknown distribution) F
We observe some data from the world, say 100 heights (z) and weights (y) of random people
We want to learn about some property of the world F, e.g.:
Mean of height
Correlation between height and weight
Variance of the empirical correlation between height and weight<br>
slide4. Standard statistical methodology Find a way to estimate the property of F of interest directly from the data
Mean height estimated by average
Correlation between height and weight estimated by empirical correlation
How do we estimate the variance of the correlation?
There are some formulae under some assumptions, but it gets complicated
Instead, we want to invent a general approach that will allow estimating every property of F relatively easily (hopefully, also well)<br>
slide5. Bootstrap idea: The Plug-in principle<br>
slide6. Graphical representation Real world Data X Statistics(X) Data X* Statistics(X*) Bootstrap world<br>
slide7. Is Bootstrap important in practice?<br>
slide8. Example: variance of empirical correlation<br>
slide9. The “double arrow” is the key to designing a bootstrap algorithm
The most standard approach: use the empirical distribution of the data
Drawing X* is drawing 100 pairs (z*,y*) with return from the original dataset
This is commonly referred to as “bootstrap sampling” or “nonparametric bootstrap”
But this is not the only approach, and often not the best one!<br>
slide10. Parametric Bootstrap example<br>
slide11. Concrete example<br>
slide12. Approach 1: standard non-parametric Bootstrap<br>
slide13. Approach 2: parametric Bootstrap using normal distribution<br>
slide14. Which one will be better here?<br>
slide15. Does Bootstrap always work?<br>
slide16. Hypothesis testing with Bootstrap<br>
slide17. Inference on phylogenetic treesFelsenstein (1985) Dataset of malaria genetic sequences from different organisms (11 species, sequences of length 221): Result of applying standard phylogenetic tree learning approach: Our inference goal: asses confidence in the 9-10 clade (subtree) – is it strongly supported by the data?<br>
slide18. Felsenstein’s Bootstrap of Phylogenetic trees<br>
slide19. Is this Bootstrap legit?<br>
slide20. Efron’s solution(s) In a beautiful paper, Efron et al. (1996, PNAS) reanalyze this problem and show:
That under some (quite complicated) assumptions Felsenstein’s approach can be considered a legitimate Bootstrap
That without these assumptions (but with some complicated math and geometry), an appropriate Bootstrap can be devised for the hypothesis testing view of the problem<br>
slide21. Efron’s hypothesis testing view<br>
slide22. A peek into Efron’s approach<br>
slide23. Comparing Bootstrap results of Felsenstein and Efron<br>
slide24. Summary Bootstrap is an extremely general and flexible paradigm for statistical inference
Allows us to handle complex situations with minimal assumptions and without complicated math
Doing theory (and also devising solutions for some problems) can get very complicated, though
Has been widely influential in science and industry
However, despite the conceptual simplicity it is often misunderstood and misapplied (well beyond Felsenstein)<br>
slide25. Thanks! saharon@post.tau.ac.il<br>
slide2. Bootstrap is a paradigm, not an algorithm<br>
slide3. An abstract view of statistics There is a “world” (=unknown distribution) F
We observe some data from the world, say 100 heights (z) and weights (y) of random people
We want to learn about some property of the world F, e.g.:
Mean of height
Correlation between height and weight
Variance of the empirical correlation between height and weight<br>
slide4. Standard statistical methodology Find a way to estimate the property of F of interest directly from the data
Mean height estimated by average
Correlation between height and weight estimated by empirical correlation
How do we estimate the variance of the correlation?
There are some formulae under some assumptions, but it gets complicated
Instead, we want to invent a general approach that will allow estimating every property of F relatively easily (hopefully, also well)<br>
slide5. Bootstrap idea: The Plug-in principle<br>
slide6. Graphical representation Real world Data X Statistics(X) Data X* Statistics(X*) Bootstrap world<br>
slide7. Is Bootstrap important in practice?<br>
slide8. Example: variance of empirical correlation<br>
slide9. The “double arrow” is the key to designing a bootstrap algorithm
The most standard approach: use the empirical distribution of the data
Drawing X* is drawing 100 pairs (z*,y*) with return from the original dataset
This is commonly referred to as “bootstrap sampling” or “nonparametric bootstrap”
But this is not the only approach, and often not the best one!<br>
slide10. Parametric Bootstrap example<br>
slide11. Concrete example<br>
slide12. Approach 1: standard non-parametric Bootstrap<br>
slide13. Approach 2: parametric Bootstrap using normal distribution<br>
slide14. Which one will be better here?<br>
slide15. Does Bootstrap always work?<br>
slide16. Hypothesis testing with Bootstrap<br>
slide17. Inference on phylogenetic treesFelsenstein (1985) Dataset of malaria genetic sequences from different organisms (11 species, sequences of length 221): Result of applying standard phylogenetic tree learning approach: Our inference goal: asses confidence in the 9-10 clade (subtree) – is it strongly supported by the data?<br>
slide18. Felsenstein’s Bootstrap of Phylogenetic trees<br>
slide19. Is this Bootstrap legit?<br>
slide20. Efron’s solution(s) In a beautiful paper, Efron et al. (1996, PNAS) reanalyze this problem and show:
That under some (quite complicated) assumptions Felsenstein’s approach can be considered a legitimate Bootstrap
That without these assumptions (but with some complicated math and geometry), an appropriate Bootstrap can be devised for the hypothesis testing view of the problem<br>
slide21. Efron’s hypothesis testing view<br>
slide22. A peek into Efron’s approach<br>
slide23. Comparing Bootstrap results of Felsenstein and Efron<br>
slide24. Summary Bootstrap is an extremely general and flexible paradigm for statistical inference
Allows us to handle complex situations with minimal assumptions and without complicated math
Doing theory (and also devising solutions for some problems) can get very complicated, though
Has been widely influential in science and industry
However, despite the conceptual simplicity it is often misunderstood and misapplied (well beyond Felsenstein)<br>
slide25. Thanks! saharon@post.tau.ac.il<br>