Conversational AI Acknowledgement: Slides from

Published  . 0 views
↓ Download
Conversational AI Acknowledgement: Slides from
1 / 1
Conversational AI Acknowledgement: Slides from - slide 1 of 32 Conversational AI Acknowledgement: Slides from - slide 2 of 32 Conversational AI Acknowledgement: Slides from - slide 3 of 32 Conversational AI Acknowledgement: Slides from - slide 4 of 32 Conversational AI Acknowledgement: Slides from - slide 5 of 32 Conversational AI Acknowledgement: Slides from - slide 6 of 32 Conversational AI Acknowledgement: Slides from - slide 7 of 32 Conversational AI Acknowledgement: Slides from - slide 8 of 32 Conversational AI Acknowledgement: Slides from - slide 9 of 32 Conversational AI Acknowledgement: Slides from - slide 10 of 32 Conversational AI Acknowledgement: Slides from - slide 11 of 32 Conversational AI Acknowledgement: Slides from - slide 12 of 32 Conversational AI Acknowledgement: Slides from - slide 13 of 32 Conversational AI Acknowledgement: Slides from - slide 14 of 32 Conversational AI Acknowledgement: Slides from - slide 15 of 32 Conversational AI Acknowledgement: Slides from - slide 16 of 32 Conversational AI Acknowledgement: Slides from - slide 17 of 32 Conversational AI Acknowledgement: Slides from - slide 18 of 32 Conversational AI Acknowledgement: Slides from - slide 19 of 32 Conversational AI Acknowledgement: Slides from - slide 20 of 32 Conversational AI Acknowledgement: Slides from - slide 21 of 32 Conversational AI Acknowledgement: Slides from - slide 22 of 32 Conversational AI Acknowledgement: Slides from - slide 23 of 32 Conversational AI Acknowledgement: Slides from - slide 24 of 32 Conversational AI Acknowledgement: Slides from - slide 25 of 32 Conversational AI Acknowledgement: Slides from - slide 26 of 32 Conversational AI Acknowledgement: Slides from - slide 27 of 32 Conversational AI Acknowledgement: Slides from - slide 28 of 32 Conversational AI Acknowledgement: Slides from - slide 29 of 32 Conversational AI Acknowledgement: Slides from - slide 30 of 32 Conversational AI Acknowledgement: Slides from - slide 31 of 32 Conversational AI Acknowledgement: Slides from - slide 32 of 32
Description: Conversational AI Acknowledgement: Slides from Prof. Dilek Hakkani-Tür Conversational Language Understanding Dialogue State Tracking Response Generation Dialogue Policy BackEnd ActionKnowledge Providers Request(time) Traditional

Related Topics

Download Presentation

"Conversational AI Acknowledgement: Slides from" is the property of its rightful owner. Permission is granted to download and print the materials on this website for personal, non-commercial use only, and to display it on your personal computer provided you do not modify the materials and that you retain all copyright notices contained in the materials. By downloading content from our website, you accept the terms of this agreement.

Presentation Transcript

slide1. Conversational AI Acknowledgement: Slides from Prof. Dilek Hakkani-Tür<br>
slide2. Conversational
Language
Understanding Dialogue
State
Tracking Response Generation Dialogue 
Policy BackEnd Action/Knowledge Providers Request(time) Traditional Tasks/Pipelines - Task-Oriented Dialogues (TOD) Back-end
query

Response Book me a table at Cascal for 2 people restaurants
reserve_restaurant
Inform(
Rest._name: Cascal,
Num_people: 2) Sure, at what time do you want the reservation?<br>
slide3. Dialogue Policy Determines the next system action, given the dialogue state and possibly the output from the backend. Example actions that an agent can take:
Dialogue state may be missing required values to make an API call
i.e., to book a table, date and time may be required, but not specified yet
Dialogue state may have values with low confidence

The backend may return too many results:
Summarize: There are several restaurants in Mountain View, many of them serve seafood, Mexican, Chinese and Italian food. Most of them are around Castro Street.
Offer a few: I found many restaurants in Mountain View. Cascal is a highly recommended Spanish restaurant in the downtown/Castro Street area.
Request more constraints: There are several restaurant in Champaign, do you have a preferred cuisine or preferred region?
The backend may return no results:
Inform No Results: I couldn't find a restaurant that serves Himalayan food in Mountain View.<br>
slide4. Dialogue Policy - Actions Action is usually defined in the form of a dialogue act, slots and their values.
Some examples include:
Greeting(): Hi, how can I help you?
Request(time): At what time would you like to see the movie?
Inform(PhoneNumber=123 456 7890): The phone number for Cascal is 123 456 7890.
Confirm(Date=Monday, Time=5pm): Did I get it right? Did you say you want the booking on Monday at 5pm?<br>
slide5. Why Learn Dialogue Policies? Expensive and error-prone to build dialogue policies that scale to variations in:
User requirements (i.e., cuisine/location)
Dialogue contexts (i.e., Alexa vs Siri)
Back-end resources (i.e., British/United airways)
Save the day versus learn to plan for longer term<br>
slide6. Task-Oriented Dialogue as a Game USER
Has a goal (fixed/flexible) SYSTEM/AGENT
Has access to APIs
Can perform the task Book my flu shot with Dr. Shaw on Monday Dr. Shaw is available on Monday at October 6th at 5:15pm and 6pm. What time would you prefer? Games take many forms: Adversarial (Chess, Go, …), Cooperative (20 questions, Pictionary), Collaborative (Dialogue)
Large space of actions and states
Multi-action turns and flexible turn-taking<br>
slide7. Learning Dialogue Policies USER AGENT Action, a State, s Reward , r Aim to learn the optimum system action at each turn
For most relevant response via supervised (SL) and imitation learning (IL)
For optimal dialogue strategy via reinforcement learning (RL) Reward = f<br>
slide8. Learning Aialogue Policies<br>
slide9. Response Generation Realizes system actions in natural language<br>
slide10. Template-based Methods Define a set of rules to map system actions/semantic frames to natural language.<br>
slide11. Response Generation Modeling Recurrent neural nets for generation (Wen et al, 2015) https://arxiv.org/pdf/1508.01755.pdf

Issue: Repetitions
Din Tai Fung is a great Taiwanese restaurant that serves Taiwanese food.<br>
slide12. Response Generation Modeling (cont.) Added semantic conditioning (Wen et al, 2015) https://arxiv.org/pdf/1508.01745.pdf
Idea: using gate mechanism to control the generated semantics (dialogue act/slots)<br>
slide13. Response Generation Modeling (cont.) RNNs aware of context and attention (Dušek and Jurčíček, 2016) https://aclanthology.org/W16-3622.pdf
Added a context encoder to a sequence-to-sequence model: Adapting to users' way of speaking and providing context aware responses<br>
slide14. Response Generation Modeling (cont.) RNNs aware of context and attention (Dušek and Jurčíček, 2016) https://aclanthology.org/W16-3622.pdf
Added a context encoder to a sequence-to-sequence model: Adapting to users' way of speaking and providing context aware responses<br>
slide15. Response Generation Modeling (cont.) Slot values are delexicalized in the input and output, however they are important for surface realization.
The food quality is great, but the service is mediocre.
The food quality and service are both great.
Plan and generate: Slot values and sentence plans help generate natural outputs (Nayak et al, 2017) https://users.soe.ucsc.edu/~maw/papers/spd-intsp-v19.pdf<br>
slide16. The E2E dataset https://arxiv.org/pdf/1706.09254<br>
slide17. Response Generation Modeling (cont.) Hierarchical (tree-structured) meaning representations and constrained decoding Output is also a linearized, tree structured representation.<br>
slide18. Are Large Language Models All You Need for Task-Oriented Dialogue? Vojtěch Hudeček, Ondrej Dusek {hudecek, odusek}@ufal.mff.cuni.cz
Charles University, Faculty of Mathematics and Physics SIGdial 2023<br>
slide19. Introduction LLMs show outstanding performances in open-domain conversations, but maybe not as much in TOD setting
Evaluate LLMs on TOD without finetuning, focusing on zero-shot and few-shot

Similar to last week's paper, "Towards LLM-driven DST" by Feng et al, EMNLP 2023, but:
No fine-tuning
They also use LLMs for dialogue policy and response generation<br>
slide20. Approach<br>
slide21. Prompt Construction Domain detection prompt Task definition, domain description, dialogue history, user utterance, belief state with DB results<br>
slide22. Prompt Construction State tracking prompt Task definition, domain description, dialogue history, user utterance, belief state with DB results + retrieved examples in the few-shot cases<br>
slide23. Prompt Construction Response prompt Task definition, domain description, dialogue history, user utterance, belief state with DB results + retrieved examples in the few-shot cases<br>
slide24. Approach (Cont.) Domain Detection and State Tracking
Multi-domain
Response Generation
Delexicalized output
Context Storage
Encoded dialogue context
Similarity based retrieval using FAISS
Positive and negative examples<br>
slide25. Experiment Setup Datasets
MultiWOZ 2.2
SGD
Models
Tk-Instruct-11B
ChatGPT
Alpaca-LoRA-7B
GPT-NeoXT-Chat-Base-20B
OPT-IML-30B Variants
Zero-shot or few-shot (-zs- vs. -fs-)
Generated or oracle belief states (-gbs vs. -obs)
Evaluation Metrics
Accuracy for domain detection
Joint Goal Accuracy (JGA) and micro-F1 for state tracking
BLEU for response generation
Success rate
Human Evaluation<br>
slide26. Results – Domain Detection Accuracy<br>
slide27. Results – DST and Response Generation Llama3.1
GPT4<br>
slide28. Results – # of examples for success rates Examples stored to retrieve from, during inference.<br>
slide29. Human Evaluation 6 annotators, NLP/LING background, two strongest models
Randomly selected goals with minimal instructions, but allowing for clarification and corrections<br>
slide30. Error Analysis – Most erroneous behaviours Prompt-recoverable errors
Prompt engineering
Inherent errors
Not easily correctable by prompt engineering
Hallucination, the model offering results not in the database or response is not grounded in context. ~10-20% of examined interactions.<br>
slide31. Conclusion & Future Ideas LLMs are not performing well in terms of belief state tracking, even when provided with in-context few-shot examples
Some issues can be improved by prompt tuning and output parsing robust to irregularities
Single-turn evaluation is too rigid and does not show the whole picture!<br>
slide32. Limitations ChatGPT API
Model-specific prompting (not one prompt that works well with all models)
Possible data contamination in some models

No fine-tuning<br>