Search
Program Calendar
Browse By Day
Browse By Person
Browse By Division
Browse By Session Type
Search Tips
Personal Schedule
Sign In
X (Twitter)
Machine-readable language resources (speech and text corpora, lexical databases, etc.) containing rich and interlingual linguistic annotations are crucial for data-driven and cross-lingual studies and technology development.
AI Lab at IMCS, University of Latvia has been creating various Latvian language resources for more than 20 years. In the last 5 years, compatibility, interoperability and integration of Latvian resources with resources being developed for other languages have been in the focus. This strategy has already paid off: Latvian language resources that have been created using interlingual models are being used in cross-lingual research and technology development by many groups and individuals around the world.
We will present four state-of-the-art resources developed for Latvian:
1) Latvian Treebank, a syntactically annotated text corpus which is compliant with the widely acknowledged Universal Dependencies grammar model that has been used to create more than 200 treebanks for more than 100 languages;
2) Latvian FrameNet, a frame-semantically annotated text corpus which is compliant with the widely acknowledged Berkeley FrameNet semantic model that has been used to created more than 20 framebanks for more than 15 languages;
3) Corpus of the Saeima, a structurally and grammatically annotated text corpus of the Latvian parliament debates which is compliant with the Text Encoding Initiative and Universal Dependencies annotation schemas and has been included in the multilingual CLARIN ParlaMint dataset of uniformly annotated parliamentary debates of many EU countries;
4) Tēzaurs.lv, a large and open lexical database which is being transformed into Latvian WordNet, compliant with and linked to Open Multilingual Wordnet.
Normunds Grūzītis, Dr. sc. comp., is a lead researcher and head of Artificial Intelligence Laboratory at Institute of Mathematics and Computer Science, University of Latvia. He is also a docent at the Faculty of Computing, University of Latvia. He has coordinated several research projects on natural language understanding and generation for Latvian, and creation of advanced Latvian language resources. In cooperation with industry partners, he has been involved in several language technology driven innovation projects. His main research interests are in computational linguistics and language technology, by combining knowledge-based and machine learning approaches. He received his PhD in computer science from the University of Latvia in 2011, followed by a postdoc position at University of Gothenburg, Sweden. He is an author or co-author of more than 45 peer-reviewed scientific publications.