Great Food, Lousy Service Topic Modeling for
Description: Great Food, Lousy Service Topic Modeling for Sentiment Analysis in Sparse Reviews Robin Melnick rmelnickstanford.edu Dan Preston dprestonstanford.edu OpenTable.com Short Characters Words Sparse An unexpected combination of Left-Bank
Related Topics
Download Presentation
"Great Food, Lousy Service Topic Modeling for" is the property of its rightful owner. Permission is granted to download and print the materials on this website for personal, non-commercial use only, and to display it on your personal computer provided you do not modify the materials and that you retain all copyright notices contained in the materials. By downloading content from our website, you accept the terms of this agreement.
Presentation Transcript
slide1. Great Food, Lousy Service Topic Modeling for Sentiment Analysis in Sparse Reviews Robin Melnick
rmelnick@stanford.edu Dan Preston
dpreston@stanford.edu<br>
slide2. OpenTable.com<br>
slide3. Short Characters Words<br>
slide4. Sparse “An unexpected combination of Left-Bank Paris and Lower Manhattan in Omaha. Divine. Inspirational and a great value.”
Food?
Ambiance?
Service?
Noise?<br>
slide5. Skewed<br>
slide6. Correlations<br>
slide7. SVM + Features, Features, Features! 30+ preprocessing and SVM classification features,
~50 configurations<br>
slide8. Key Features Stemming
Porter 1980 via NLTK
<fast>, <faster>, <fastest> <fast>
Negation processing
(enhanced approach from Pang et al. 2002)
“Not a great experience.” NOT_great
“They never disappoint!” NOT_disappoint
Net sentiment count
pos/neg lexicon (Harvard General Inquirer)
running +/- count
“Incredible(+) food, but our server was rude(-).” (0)<br>
slide9. Results (so far) Trained on 10,000 reviews
Tested on ~80,000 reviews
Accuracy
Baseline: 50.0%
Intermediate model: 56.6% (1.13x)
abs( average scoring delta ): 0.56<br>
slide10. Topic Modeling Hand-seeded topic-word list expanded via WordNet SynSets
sub-topic classifiers
topic-filtered n-grams
<soupFOOD was fantasticADJ>
<fantasticADJ soupFOOD was>
topic-word proximity filtering
both above <fantasticADJ/FOOD>.
Results:<br>
slide11. Word-Rating Distributions “worst” “mediocre” “decent” “solid” “exceeded”<br>
slide12. Frequency-Weighted Entropy Model Accuracy
Baseline: 50.0%
Intermediate model: 56.6%
Best (entropy) model: 58.6% (1.17x)
abs( average scoring delta ): 0.56 0.52<br>
rmelnick@stanford.edu Dan Preston
dpreston@stanford.edu<br>
slide2. OpenTable.com<br>
slide3. Short Characters Words<br>
slide4. Sparse “An unexpected combination of Left-Bank Paris and Lower Manhattan in Omaha. Divine. Inspirational and a great value.”
Food?
Ambiance?
Service?
Noise?<br>
slide5. Skewed<br>
slide6. Correlations<br>
slide7. SVM + Features, Features, Features! 30+ preprocessing and SVM classification features,
~50 configurations<br>
slide8. Key Features Stemming
Porter 1980 via NLTK
<fast>, <faster>, <fastest> <fast>
Negation processing
(enhanced approach from Pang et al. 2002)
“Not a great experience.” NOT_great
“They never disappoint!” NOT_disappoint
Net sentiment count
pos/neg lexicon (Harvard General Inquirer)
running +/- count
“Incredible(+) food, but our server was rude(-).” (0)<br>
slide9. Results (so far) Trained on 10,000 reviews
Tested on ~80,000 reviews
Accuracy
Baseline: 50.0%
Intermediate model: 56.6% (1.13x)
abs( average scoring delta ): 0.56<br>
slide10. Topic Modeling Hand-seeded topic-word list expanded via WordNet SynSets
sub-topic classifiers
topic-filtered n-grams
<soupFOOD was fantasticADJ>
<fantasticADJ soupFOOD was>
topic-word proximity filtering
both above <fantasticADJ/FOOD>.
Results:<br>
slide11. Word-Rating Distributions “worst” “mediocre” “decent” “solid” “exceeded”<br>
slide12. Frequency-Weighted Entropy Model Accuracy
Baseline: 50.0%
Intermediate model: 56.6%
Best (entropy) model: 58.6% (1.17x)
abs( average scoring delta ): 0.56 0.52<br>