Multilingual Information Retrieval Doug Oard

Published  . 0 views
↓ Download
Multilingual Information Retrieval Doug Oard
1 / 1
Multilingual Information Retrieval Doug Oard - slide 1 of 53 Multilingual Information Retrieval Doug Oard - slide 2 of 53 Multilingual Information Retrieval Doug Oard - slide 3 of 53 Multilingual Information Retrieval Doug Oard - slide 4 of 53 Multilingual Information Retrieval Doug Oard - slide 5 of 53 Multilingual Information Retrieval Doug Oard - slide 6 of 53 Multilingual Information Retrieval Doug Oard - slide 7 of 53 Multilingual Information Retrieval Doug Oard - slide 8 of 53 Multilingual Information Retrieval Doug Oard - slide 9 of 53 Multilingual Information Retrieval Doug Oard - slide 10 of 53 Multilingual Information Retrieval Doug Oard - slide 11 of 53 Multilingual Information Retrieval Doug Oard - slide 12 of 53 Multilingual Information Retrieval Doug Oard - slide 13 of 53 Multilingual Information Retrieval Doug Oard - slide 14 of 53 Multilingual Information Retrieval Doug Oard - slide 15 of 53 Multilingual Information Retrieval Doug Oard - slide 16 of 53 Multilingual Information Retrieval Doug Oard - slide 17 of 53 Multilingual Information Retrieval Doug Oard - slide 18 of 53 Multilingual Information Retrieval Doug Oard - slide 19 of 53 Multilingual Information Retrieval Doug Oard - slide 20 of 53 Multilingual Information Retrieval Doug Oard - slide 21 of 53 Multilingual Information Retrieval Doug Oard - slide 22 of 53 Multilingual Information Retrieval Doug Oard - slide 23 of 53 Multilingual Information Retrieval Doug Oard - slide 24 of 53 Multilingual Information Retrieval Doug Oard - slide 25 of 53 Multilingual Information Retrieval Doug Oard - slide 26 of 53 Multilingual Information Retrieval Doug Oard - slide 27 of 53 Multilingual Information Retrieval Doug Oard - slide 28 of 53 Multilingual Information Retrieval Doug Oard - slide 29 of 53 Multilingual Information Retrieval Doug Oard - slide 30 of 53 Multilingual Information Retrieval Doug Oard - slide 31 of 53 Multilingual Information Retrieval Doug Oard - slide 32 of 53 Multilingual Information Retrieval Doug Oard - slide 33 of 53 Multilingual Information Retrieval Doug Oard - slide 34 of 53 Multilingual Information Retrieval Doug Oard - slide 35 of 53 Multilingual Information Retrieval Doug Oard - slide 36 of 53 Multilingual Information Retrieval Doug Oard - slide 37 of 53 Multilingual Information Retrieval Doug Oard - slide 38 of 53 Multilingual Information Retrieval Doug Oard - slide 39 of 53 Multilingual Information Retrieval Doug Oard - slide 40 of 53 Multilingual Information Retrieval Doug Oard - slide 41 of 53 Multilingual Information Retrieval Doug Oard - slide 42 of 53 Multilingual Information Retrieval Doug Oard - slide 43 of 53 Multilingual Information Retrieval Doug Oard - slide 44 of 53 Multilingual Information Retrieval Doug Oard - slide 45 of 53 Multilingual Information Retrieval Doug Oard - slide 46 of 53 Multilingual Information Retrieval Doug Oard - slide 47 of 53 Multilingual Information Retrieval Doug Oard - slide 48 of 53 Multilingual Information Retrieval Doug Oard - slide 49 of 53 Multilingual Information Retrieval Doug Oard - slide 50 of 53 Multilingual Information Retrieval Doug Oard - slide 51 of 53 Multilingual Information Retrieval Doug Oard - slide 52 of 53 Multilingual Information Retrieval Doug Oard - slide 53 of 53
Description: Multilingual Information Retrieval Doug Oard College of Information Studies and UMIACS University of Maryland, College Park USA January 14, 2019 AFIRM Global Trade Source: Wikipedia (mostly 2017 estimates) USA Japan South Korea Hong Kong

Related Topics

Download Presentation

"Multilingual Information Retrieval Doug Oard" is the property of its rightful owner. Permission is granted to download and print the materials on this website for personal, non-commercial use only, and to display it on your personal computer provided you do not modify the materials and that you retain all copyright notices contained in the materials. By downloading content from our website, you accept the terms of this agreement.

Presentation Transcript

slide1. Multilingual Information Retrieval Doug Oard
College of Information Studies and UMIACS
University of Maryland, College Park
USA January 14, 2019 AFIRM<br>
slide2. Global Trade Source: Wikipedia (mostly 2017 estimates) USA Japan South Korea Hong Kong China EU<br>
slide3. Most Widely-Spoken Languages Source: Ethnologue (SIL), 2018<br>
slide4. Web Pages Global Internet Users<br>
slide5. What Does “Multilingual” Mean? Mixed-language document
Document containing more than one language
Mixed-language collection
Collection of documents in different languages
Multi-monolingual systems
Can retrieve from a mixed-language collection
Cross-language system
Query in one language finds document in another
(Truly) multingual system
Queries can find documents in any language<br>
slide6. A Story in Two Parts IR from the ground up in any language
Focusing on document representation

Cross-Language IR
To the extent time allows<br>
slide7. Documents Query Hits Representation
Function Representation
Function Query Representation Document Representation Comparison
Function Index<br>
slide8. ASCII American Standard Code for Information Interchange

ANSI X3.4-1968 | 0 NUL | 32 SPACE | 64 @ | 96 ` |
| 1 SOH | 33 ! | 65 A | 97 a |
| 2 STX | 34 " | 66 B | 98 b |
| 3 ETX | 35 # | 67 C | 99 c |
| 4 EOT | 36 $ | 68 D | 100 d |
| 5 ENQ | 37 % | 69 E | 101 e |
| 6 ACK | 38 & | 70 F | 102 f |
| 7 BEL | 39 ' | 71 G | 103 g |
| 8 BS | 40 ( | 72 H | 104 h |
| 9 HT | 41 ) | 73 I | 105 i |
| 10 LF | 42 * | 74 J | 106 j |
| 11 VT | 43 + | 75 K | 107 k |
| 12 FF | 44 , | 76 L | 108 l |
| 13 CR | 45 - | 77 M | 109 m |
| 14 SO | 46 . | 78 N | 110 n |
| 15 SI | 47 / | 79 O | 111 o | | 16 DLE | 48 0 | 80 P | 112 p |
| 17 DC1 | 49 1 | 81 Q | 113 q |
| 18 DC2 | 50 2 | 82 R | 114 r |
| 19 DC3 | 51 3 | 83 S | 115 s |
| 20 DC4 | 52 4 | 84 T | 116 t |
| 21 NAK | 53 5 | 85 U | 117 u |
| 22 SYN | 54 6 | 86 V | 118 v |
| 23 ETB | 55 7 | 87 W | 119 w |
| 24 CAN | 56 8 | 88 X | 120 x |
| 25 EM | 57 9 | 89 Y | 121 y |
| 26 SUB | 58 : | 90 Z | 122 z |
| 27 ESC | 59 ; | 91 [ | 123 { |
| 28 FS | 60 < | 92 \ | 124 | |
| 29 GS | 61 = | 93 ] | 125 } |
| 30 RS | 62 > | 94 ^ | 126 ~ |
| 31 US | 64 ? | 95 _ | 127 DEL |<br>
slide9. The Latin-1 Character Set ISO 8859-1 8-bit characters for Western Europe
French, Spanish, Catalan, Galician, Basque, Portuguese, Italian, Albanian, Afrikaans, Dutch, German, Danish, Swedish, Norwegian, Finnish, Faroese, Icelandic, Irish, Scottish, and English Printable Characters, 7-bit ASCII Additional Defined Characters, ISO 8859-1<br>
slide10. Other ISO-8859 Character Sets -2 -3 -4 -5 -7 -6 -9 -8<br>
slide11. East Asian Character Sets More than 256 characters are needed
Two-byte encoding schemes (e.g., EUC) are used
Several countries have unique character sets
GB in Peoples Republic of China, BIG5 in Taiwan, JIS in Japan, KS in Korea, TCVN in Vietnam
Many characters appear in several languages
Research Libraries Group developed EACC
Unified “CJK” character set for USMARC records<br>
slide12. Unicode Single code for all the world’s characters
ISO Standard 10646
Separates “code space” from “encoding”
Code space extends Latin-1
The first 256 positions are identical
UTF-7 encoding will pass through email
Uses only the 64 printable ASCII characters
UTF-8 encoding is designed for disk file systems<br>
slide13. Limitations of Unicode Produces larger files than Latin-1
Fonts may be hard to obtain for some characters
Some characters have multiple representations
e.g., accents can be part of a character or separate
Some characters look identical when printed
But they come from unrelated languages
Encoding does not define the “sort order”<br>
slide14. Strings and Segments Retrieval is (often) a search for concepts
But what we actually search are character strings

What strings best represent concepts?
In English, words are often a good choice
Well-chosen phrases might also be helpful
In German, compounds may need to be split
Otherwise queries using constituent words would fail
In Chinese, word boundaries are not marked
Thissegmentationproblemissimilartothatofspeech<br>
slide15. Tokenization Words (from linguistics):
Morphemes are the units of meaning
Combined to make words
Anti (disestablishmentarian) ism

Tokens (from computer science)
Doug ’s running late !<br>
slide16. Morphological Segmentation Swahili Example Credit: Ramy Eskander<br>
slide17. Morphological Segmentation Somali Example Credit: Ramy Eskander<br>
slide18. Stemming Conflates words, usually preserving meaning
Rule-based suffix-stripping helps for English
{destroy, destroyed, destruction}: destr
Prefix-stripping is needed in some languages
Arabic: {alselam}: selam [Root: SLM (peace)]
Imperfect: goal is to usually be helpful
Overstemming
{centennial,century,center}: cent
Understamming:
{acquire,acquiring,acquired}: acquir
{acquisition}: acquis
Snowball: rule-based system for making stemmers<br>
slide19. Longest Substring Segmentation Greedy algorithm based on a lexicon

Start with a list of every possible term

For each unsegmented string
Remove the longest single substring in the list
Repeat until no substrings are found in the list<br>
slide20. Longest Substring Example Possible German compound term (!):
washington

List of German words:
ach, hin, hing, sei, ton, was, wasch

Longest substring segmentation
was-hing-ton
Roughly translates as “What tone is attached?”<br>
slide22. Probabilistic Segmentation For an input string c1 c2 c3 … cn
Try all possible partitions into w1 w2 w3 …
c1 c2 c3 … cn
c1 c2 c3 c3 … cn
c1 c2 c3 … cn
etc.
Choose the highest probability partition
Compute Pr(w1 w2 w3 ) using a language model
Challenges: search, probability estimation<br>
slide23. Non-Segmentation: N-gram Indexing Consider a Chinese document c1 c2 c3 … cn

Don’t segment (you could be wrong!)

Instead, treat every character bigram as a term
c1 c2 , c2 c3 , c3 c4 , … , cn-1 cn

Break up queries the same way<br>
slide24. A “Term” is Whatever You Index Word sense
Token
Word
Stem
Character n-gram
Phrase<br>
slide25. Summary A term is whatever you index
So the key is to index the right kind of terms!

Start by finding fundamental features
We have focused on character coded text
Same ideas apply to handwriting, OCR, and speech

Combine characters into easily recognized units
Words where possible, character n-grams otherwise

Apply further processing to optimize results
Stemming, phrases, …<br>
slide26. A Story in Two Parts IR from the ground up in any language
Focusing on document representation

Cross-Language IR
To the extent time allows<br>
slide27. Query-Language CLIR English
queries Results select examine<br>
slide28. Document-Language CLIR Somali
queries Somali
documents Results select examine<br>
slide29. Query vs. Document Translation Query translation
Efficient for short queries (not relevance feedback)
Limited context for ambiguous query terms

Document translation
Rapid support for interactive selection
Need only be done once (if query language is same)<br>
slide30. Indexing Time: Statistical Document Translation<br>
slide31. Language-Neutral Retrieval “Interlingual”
Retrieval 1: 0.91
2: 0.57
3: 0.36 Query
“Translation” Somali
Query
Terms English
Document
Terms Document
“Translation”<br>
slide32. Translation Evidence Lexical Resources
Phrase books, bilingual dictionaries, …
Large text collections
Translations (“parallel”)
Similar topics (“comparable”)
Similarity
Similar writing (if the character set is the same)
Similar pronunciation
People
May be able to guess topic from lousy translations<br>
slide33. Types of Lexical Resources Ontology
Organization of knowledge
Thesaurus
Ontology specialized to support search
Dictionary
Rich word list, designed for use by people
Lexicon
Rich word list, designed for use by a machine
Bilingual term list
Pairs of translation-equivalent terms<br>
slide35. Backoff Translation Lexicon might contain stems, surface forms, or some combination of the two. mangez mangez mangez mange mangez mange mangez mangent - eat - eats - eat - eat Document Translation Lexicon<br>
slide36. Hieroglyphic Egyptian Demotic Greek<br>
slide37. Types of Bilingual Corpora Parallel corpora: translation-equivalent pairs
Document pairs
Sentence pairs
Term pairs

Comparable corpora: topically related
Collection pairs
Document pairs<br>
slide38. Some Modern Rosetta Stones News:
DE-News (German-English)
Hong-Kong News, Xinhua News (Chinese-English)
Government:
Canadian Hansards (French-English)
Europarl (Danish, Dutch, English, Finnish, French, German, Greek, Italian, Portugese, Spanish, Swedish)
UN Treaties (Russian, English, Arabic, …)
Religion
Bible, Koran, Book of Mormon<br>
slide39. Word-Level Alignment Diverging opinions about planned tax reform Unterschiedliche Meinungen zur geplanten Steuerreform English German Madam President , I had asked the administration … English Señora Presidenta, había pedido a la administración del Parlamento … Spanish<br>
slide40. A Translation Model From word-aligned bilingual text, we induce a translation model

Example: where, p(探测|survey) = 0.4
p(试探|survey) = 0.3
p(测量|survey) = 0.25
p(样品|survey) = 0.05<br>
slide41. Using Multiple Translations Weighted Structured Query Translation
Takes advantage of multiple translations and translation probabilities
TF and DF of query term e are computed using TF and DF of its translations:<br>
slide42. BM-25<br>
slide43. Retrieval Effectiveness CLEF French<br>
slide44. Bilingual Query Expansion source language query Query
Translation results Source
Language
IR Target
Language
IR source language
collection target language
collection expanded
source language
query expanded
target language
terms Pre-translation expansion Post-translation expansion<br>
slide45. Query Expansion Effect Paul McNamee and James Mayfield, SIGIR-2002<br>
slide46. Cognate Matching Dictionary coverage is inherently limited
Translation of proper names
Translation of newly coined terms
Translation of unfamiliar technical terms

Strategy: model derivational translation
Orthography-based
Pronunciation-based<br>
slide47. Matching Orthographic Cognates Retain untranslatable words unchanged
Often works well between European languages

Rule-based systems
Even off-the-shelf spelling correction can help!

Subword (e.g., character-level) MT
Trained using a set of representative cognates<br>
slide48. Matching Phonetic Cognates Forward transliteration
Generate all potential transliterations

Reverse transliteration
Guess source string(s) that produced a transliteration

Match in phonetic space<br>
slide49. Cross-Language “Retrieval” Ranked List Query
Translation<br>
slide50. Uses of “MT” in CLIR Search Translated
Query Selection Ranked List Examination Document Use Document Query
Formulation Query
Translation<br>
slide51. Interactive Cross-Language Question Answering iCLEF 2004<br>
slide52. Questions, Grouped by Difficulty 8 Who is the managing director of the International Monetary Fund?
11 Who is the president of Burundi?
13 Of what team is Bobby Robson coach?
4 Who committed the terrorist attack in the Tokyo underground?
16 Who won the Nobel Prize for Literature in 1994?
6 When did Latvia gain independence?

14 When did the attack at the Saint-Michel underground station in Paris occur?
7 How many people were declared missing in the Philippines after the
typhoon “Angela”?
2 How many human genes are there?
10 How many people died of asphyxia in the Baku underground?
15 How many people live in Bombay?
12 What is Charles Millon's political party?

1 What year was Thomas Mann awarded the Nobel Prize?
3 Who is the German Minister for Economic Affairs?
9 When did Lenin die?
5 How much did the Channel Tunnel cost?<br>
slide53. For Further Reading Multilingual IR
Paul McNamee et al, Addressing Morphological Variation in Alphabetic Languages, SIGIR, 2009
African-Language IR
Open CLIR Challenge (Swahili), IARPA, 2018
Nkosana Malumba et al, AfriWeb: A Search Engine for a Marginalized Language, ICADL, 2015
Cross-Language IR
Jian-Yun Nie, Cross-Language Information Retrieval, Synthesis Lectures in HLT, Morgan&Claypool, 2010
Jianqiang Wang and Douglas W. Oard, Matching Meaning for Cross-Language Information Retrieval, Information Processing and Management, 2012<br>