Policy Gradient as a Proxy for Dynamic Oracles in Constituency Parsing Daniel Fried and Dan Klein Policy Gradient as a Proxy for Dynamic Oracles in Constituency Parsing Daniel Fried and Dan Klein Policy Gradient as a Proxy for Dynamic
"Policy Gradient as a Proxy for Dynamic Oracles in" is the property of its rightful owner. Permission is granted to
download and print the materials on this website for personal, non-commercial use only, and to display it
on your personal computer provided you do not modify the materials and that you retain all copyright
notices contained in the materials. By downloading content from our website, you accept the terms of this
agreement.
Presentation Transcript
01
Policy Gradient as a Proxy for Dynamic Oracles in Constituency Parsing Daniel Fried and Dan Klein<br>
02
Policy Gradient as a Proxy for Dynamic Oracles in Constituency Parsing Daniel Fried and Dan Klein<br>
03
Policy Gradient as a Proxy for Dynamic Oracles in Constituency Parsing Daniel Fried and Dan Klein<br>
04
Parsing by Local Decisions The cat took a nap . NP NP VP S (S (NP The cat ) (VP …<br>
05
Non-local Consequences Exposure Bias Prediction True
Parse (S (NP The (S (VP (NP cat ?? [Ranzato et al. 2016; Wiseman and Rush 2016] … Loss-Evaluation Mismatch<br>
06
Dynamic Oracle Training Prediction
(sample, or greedy) True Parse (S (NP The (S (VP (NP cat … The The (NP Oracle The cat … Explore at training time. Supervise each state with an expert policy. [Goldberg & Nivre 2012; Ballesteros et al. 2016; inter alia]<br>
07
Dynamic Oracles Help! Expert Policies / Dynamic Oracles Daume III et al., 2009; Ross et al., 2011; Choi and Palmer, 2011; Goldberg and Nivre, 2012; Chang et al., 2015; Ballesteros et al., 2016; Stern et al. 2017 PTB Constituency Parsing F1<br>
08
What if we don’t have a dynamic oracle? Use reinforcement learning<br>
09
Reinforcement Learning Helps! (in other tasks) Auli and Gao, 2014; Ranzato et al., 2016; Shen et al., 2016 machine translation Xu et al., 2016; Wiseman and Rush, 2016; Edunov et al. 2017 machine translation several, including dependency parsing CCG parsing<br>
10
Policy Gradient Training [Williams, 1992] Minimize expected sequence-level cost: addresses exposure bias (compute by sampling) addresses loss mismatch(compute F1) compute in the same way as for the true tree Prediction True Parse<br>
11
Policy Gradient Training The cat took a nap. gradientfor candidate<br>
12
Experiments<br>
13
Setup x<br>
14
English PTB F1<br>
15
Training Efficiency PTB learning curves for the Top-Down parser<br>
16
French Treebank F1<br>
17
Chinese Penn Treebank v5.1 F1<br>
18
Conclusions Local decisions can have non-local consequences
Loss mismatch
Exposure bias
How to deal with the issues caused by local decisions?
Dynamic oracles: efficient, model specific
Policy gradient: slower to train, but general purpose<br>
19
Thank you!<br>
20
For Comparison: A Novel Oracle for RNNG (S (NP The man 1. Close current constituent if it’s a true constituent… … or it could never be a true constituent. 2. Otherwise, open the outermost unopened true constituent at this position. 3. Otherwise, shift the next word. (S (NP The man ) (VP had ) (S (NP The man ) (VP ) (S (NP The man ) (VP (S (NP The man ) (VP had …<br>
21
What if we don’t have a dynamic oracle? Define one<br>
22
For Comparison: A Novel Oracle for RNNG (S (NP The man 1. Close current constituent if it’s a true constituent… … or it could never be a true constituent. 2. Otherwise, open the outermost unopened true constituent at this position. 3. Otherwise, shift the next word. (S (NP The man ) (VP had ) (S (NP The man ) (VP ) (S (NP The man ) (VP (S (NP The man ) (VP had …<br>