Individual Submission Summary
Share...

Direct link:

Latvian Language Resources for Cross-Lingual Language Technology and Digital Humanities

Fri, May 27, 4:00 to 5:30pm, Denny Hall, 112

Abstract

Machine-readable language resources (speech and text corpora, lexical databases, etc.) containing rich and interlingual linguistic annotations are crucial for data-driven and cross-lingual studies and technology development.
AI Lab at IMCS, University of Latvia has been creating various Latvian language resources for more than 20 years. In the last 5 years, compatibility, interoperability and integration of Latvian resources with resources being developed for other languages have been in the focus. This strategy has already paid off: Latvian language resources that have been created using interlingual models are being used in cross-lingual research and technology development by many groups and individuals around the world.
We will present four state-of-the-art resources developed for Latvian:
1) Latvian Treebank, a syntactically annotated text corpus which is compliant with the widely acknowledged Universal Dependencies grammar model that has been used to create more than 200 treebanks for more than 100 languages;
2) Latvian FrameNet, a frame-semantically annotated text corpus which is compliant with the widely acknowledged Berkeley FrameNet semantic model that has been used to created more than 20 framebanks for more than 15 languages;
3) Corpus of the Saeima, a structurally and grammatically annotated text corpus of the Latvian parliament debates which is compliant with the Text Encoding Initiative and Universal Dependencies annotation schemas and has been included in the multilingual CLARIN ParlaMint dataset of uniformly annotated parliamentary debates of many EU countries;
4) Tēzaurs.lv, a large and open lexical database which is being transformed into Latvian WordNet, compliant with and linked to Open Multilingual Wordnet.

Short Bio

Normunds Grūzītis, Dr. sc. comp., is a lead researcher and head of Artificial Intelligence Laboratory at Institute of Mathematics and Computer Science, University of Latvia. He is also a docent at the Faculty of Computing, University of Latvia. He has coordinated several research projects on natural language understanding and generation for Latvian, and creation of advanced Latvian language resources. In cooperation with industry partners, he has been involved in several language technology driven innovation projects. His main research interests are in computational linguistics and language technology, by combining knowledge-based and machine learning approaches. He received his PhD in computer science from the University of Latvia in 2011, followed by a postdoc position at University of Gothenburg, Sweden. He is an author or co-author of more than 45 peer-reviewed scientific publications.

Author