Towards Semi Automatic Analysis of Spontaneous

Published  . 0 views
↓ Download
Towards Semi Automatic Analysis of Spontaneous
1 / 1
Towards Semi Automatic Analysis of Spontaneous - slide 1 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 2 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 3 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 4 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 5 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 6 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 7 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 8 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 9 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 10 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 11 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 12 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 13 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 14 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 15 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 16 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 17 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 18 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 19 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 20 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 21 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 22 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 23 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 24 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 25 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 26 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 27 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 28 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 29 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 30 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 31 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 32 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 33 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 34 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 35 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 36 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 37 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 38 of 39 Towards Semi Automatic Analysis of Spontaneous - slide 39 of 39
Description: Towards Semi Automatic Analysis of Spontaneous Language for Dutch Jan Odijk CLARIN Annual Conference 2020-10-05 1 Overview Problem and research question SASTA application Results Deviant Language Concluding Remarks Outlook for future work 2

Related Topics

Download Presentation

"Towards Semi Automatic Analysis of Spontaneous" is the property of its rightful owner. Permission is granted to download and print the materials on this website for personal, non-commercial use only, and to display it on your personal computer provided you do not modify the materials and that you retain all copyright notices contained in the materials. By downloading content from our website, you accept the terms of this agreement.

Presentation Transcript

slide1. Towards Semi Automatic Analysis of Spontaneous Language for Dutch Jan Odijk
CLARIN Annual Conference

2020-10-05 1<br>
slide2. Overview Problem and research question
SASTA application
Results
Deviant Language
Concluding Remarks
Outlook for future work 2<br>
slide3. Problem Language development (disorders), aphasia
Assessment includes spontaneous language sessions
Manual analysis
A lot effort /time, not paid for
 often not done or only partially
Can we automate this using language technology?
Goal: develop an application to semi-automatically grammatically analyse spontaneous language transcripts
Methods: TARSP, STAP and ASTA 3<br>
slide4. Problem TARSP (children 1-4), STAP (older children) and ASTA (aphasia patients)
Define language measures
E.g. occurrence of certain word classes,
# nouns, # compound nouns, # diminutives, …
# verbs, # finite verbs, # copulas, …
Constructions
Auxiliary + infinitive, adjective used attributively, …
Inversion, subordinate clauses, coordination, …
Derive scores and compare these to normative figures 4<br>
slide5. SASTA Application SASTA Application (SASTA project)
Input:
Transcript of a spontaneous language session
Formats: Sasta-specific (.docx, .txt) or CHAT
Or: annotated transcript
Parameter: select a method: TARSP, ASTA, STAP
Output:
Generic score form
Method-specific score forms (currently: TARSP, ASTA)
Annotated transcript, format based on the Schlichting 2005 Appendix 5<br>
slide6. SASTA Application SASTA Application
Web application
Public version soon here: https://sasta.hum.uu.nl
Server version for use in clinics
desktop application
Internal development:
Sastadev software 6<br>
slide7. SASTA Application: how does it work? Based on GrETEL, a treebank query application
Crucial Components
Upload facility (upload one’s own text or corpus)
Parse facility (Alpino parser)
Query facilities (Xpath query language)
Different arrangement of these components:
GrETEL: optimized for one query applied to a large treebank
SASTA: optimized for multiple queries applied to a small treebank 7<br>
slide8. SASTA Application: How does it work? Input text  CHAT format
CHAT processed:
Metadata
CHAT-annotations (result of the AnnCor project)
Each utterance automatically parsed by the Alpino-parser treebank
Method defined (in Excel format)
Method is read in
Its queries applied to the treebank
Queries generate scores and annotations
Illustration 8<br>
slide9. Results (4-sep) Data provided
Bronze reference
Data provided adapted:
By the developers, after applying SASTA
Currently checked independently by the original data providers
 Silver reference 9<br>
slide10. Results (4-sep) Comparison
Results v. Bronze
Indicator for quality (independent reference)
Results v. Silver
Indicator for quality (partially dependent reference)
Bronze v. Silver
Indicator for the quality of human annotation
Recall, Precision, F1-score
Results reflect defined core queries 10<br>
slide11. ASTA Scores (4-sep) 11<br>
slide12. TARSP (1-sep) 12<br>
slide13. STAP scores (25-sep) 13 Complexity analysis (STAP-form sheet 5)<br>
slide14. Deviant Language SASTA
mainly grammatical analysis
no analysis / correction of deviant language (which occurs a lot)
Partially manual annotation (Examples)
But it is incomplete, inconsistent, often absent
Deviant language negatively affects the grammatical analysis (Examples)
Can we (partially) automate this too? 14<br>
slide15. Deviant Language Examples Deviant transcriptions (Examples)
Usually to indicate a specific pronunciation*
(Regional) informal spoken language
ie-diminutives* (Examples)
Incorrect possessive pronouns (Examples)
Wrong pronunciation indicated in the transcript
Examples

(* = first version has been tested independently
+ = first version is available) 15<br>
slide16. Deviant Language False starts, repetitions, incomplete utterances+
Examples
Grammatical errors
(wrong) overgeneralisations* of verbal inflection
Examples
Wrong determiner+: Examples
Wrong auxiliary: Examples 16<br>
slide17. Other matters Limitations of Alpino
insufficient information on compounds*
Examples
Insufficient information on verbless utterances
Examples
Insufficient information on V1-sentence+
Examples
Analysis with words the child does not know at its age: Examples
Analysis as a construction the child does not know at its age: Examples 17<br>
slide18. Small experiment Schlichting 2005 appendix
automatic corrections (marked with *) improve recall (2 percent points) and precision (0.4 percent points)
(not integrated in Sasta yet, we are currently working on this) 18<br>
slide19. Concluding remarks Automating analysis of spontaneous language:
Promising results, though not perfect
Potential for further improvements
Its actual usefulness must be tested in clinical setting
If successful,
May have great societal impact
Enables / reduces costs / improves quality of such assessments
May have scientific impact
Many of the linguistic problems require sophisticated solutions
May benefit the CLARIN research infrastructure
Derivative program CHAMP-NL, CHAT iMProver for Dutch
Here applied to Dutch
But approach can be followed for other languages for which there is a parser, and a query system to query the structures generated by the parser 19<br>
slide20. Outlook SASTA project phase 2
is ASTA good enough to be useful / efficient / workable in the clinical setting?
Proposal prepared, meeting VKL in October
Cooperation with other partners
Auris, Hogeschool Utrecht, Itslanguage
Extend speech recognition, combine with speech recognition, alternative annotation interfaces 20<br>
slide21. Outlook SASTA+ (funded by Utrecht University)
Deal with deviant language
Generate multiple alternatives
Parse these alternatives
Compare the resulting parses and select the best
E.g. degree of grammatical cohesion
Automatic replacement of deviant pronunciation based on
Context
Other documents of the patient / the clinic / CHILDES corpora 21<br>
slide22. SASTA+ And the derivative product:
CHAMP-NL (CHAT iMProver for Dutch)
suggests improvements of CHAT files
Especially important since we also offer MOR en GRA analysis for CHAT files 22<br>
slide23. Thanks for your attention! 23<br>
slide24. Ref & Meer info Schlichting, Liesbeth (2005). TARSP. Taalontwikkelingsschaal van Nederlandse kinderen van 1-4 jaar. Amsterdam: Pearson.
http://portal.clarin.nl , http://www.clariah.nl
Augustinus, L. et al. 2017. GrETEL: A Tool for Example-Based Treebank Mining. In: Odijk. J and van Hessen. A. (eds.) 2017. CLARIN in the Low Countries. Pp. 269–280. London: Ubiquity Press. DOI: https://doi.org/10.5334/bbi.22 License: CC-BY 4.0
Odijk. J., van der Klis, M., and Spoel, S. (2018). Extensions to the GrETEL treebank query application. Proceedings of the 16th International Workshop on Treebanks and Linguistic Theories (TLT16) pp 46-55. Prague. http://aclweb.org/anthology/W/W17/W17-7608.pdf
Odijk & Van Hessen (eds.) 2017. CLARIN in the Low Countries. London: Ubiquity Press. (Open Access). DOI: http://dx.doi.org/10.5334/bbi
Odijk, J. et al. 2018. The AnnCor CHILDES Treebank. In Nicoletta Calzolari et al. (Eds.). Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018). Miyazaki. Japan (pp. 2275-2283). Paris. France: ELRA. 24<br>
slide25. TARSP Experiment 25 Method specifies language measures
Example (TARSP). illustrated in GrETEL:
Hww i (auxiliary with infinitive)
Query
Applied to the Schlichting 2005 Appendix
5 results: utterances 10, 11, 13, 22, 28
Schlichting’s analysis: 10, 11, 13, 22, 28<br>
slide26. CHAT Annotation Examples 26<br>
slide27. Deviant Language  Wrong analysis 27<br>
slide28. Deviant Transcription 28<br>
slide29. Ie-Diminutives 29<br>
slide30. Possessive pronouns 30<br>
slide31. Wrong pronunciation 31<br>
slide32. False starts, Repetitions, incomplete utterances 32<br>
slide33. (Wrong) Overgeneralisations 33<br>
slide34. Wrong Det / Aux 34<br>
slide35. Compounds 35<br>
slide36. V1 Utterances 36<br>
slide37. Wrong word/ construction analysis 37<br>
slide38. Verbless utterances 38<br>
slide39. SASTA Project Phase I SASTA project
Semi-Automatic Analysis of Spontaneous Language
Funded by CLARIAH, Stichting Taaltechnologie, Vereniging voor Klinische Linguïstiek (VKL), Vogellanden.
Team:
Utrecht University: Jan Odijk, Jelte van Boheemen, Sjoerd Eilander 
VKL: Rob Zwitserlood, Margo Zwitserlood
Vogellanden & VKL: Nina Blom, Elsbeth Boxum 39<br>