Dr. Nikolett Mus
Corpus-based grammar of under-described languages
I am a senior research fellow at the Research Centre for Linguistics in Budapest, Hungary.
My research lies at the intersection of syntax, information structure, corpus linguistics, and computational linguistics, with a particular interest in under-described and low-resource languages. I investigate how grammatical structure can be studied through corpus-based and computational approaches, combining theoretical syntax with language documentation and the development of linguistic resources.
Much of my research focuses on Uralic and Turkic languages, especially the syntax and information structure of verb-final languages, as well as the development of computational resources for corpus-based research. My work includes the creation of the Universal Dependencies Tundra Nenets Treebank, annotation methodologies, and computational tools for linguistic analysis.
More recently, my research has expanded to Hungarian Sign Language, where I am interested in corpus construction, annotation methodologies, grammatical description, and the development of sustainable research infrastructures that support long-term linguistic research on sign languages.
More broadly, I am interested in methodological questions arising in the documentation of languages without standardized written forms, including endangered spoken languages, sign languages, and non-standard language varieties such as the Csángó dialects of Hungarian spoken in Moldavia (Romania).
Through this website, I share my publications, software, corpora, and ongoing research projects.