This work package involves the methodological exploration of applying vector space models in lectometry research. On the basis of an evaluation of previous studies on lectometric measurements, we established a general framework for corpus-based lexical lectometry research that considers most if not all options for different steps. We also have experimented with the technical aspects of vector space models in order to see whether various parameter settings would lead to different and better results.
WP2 SYNCHRONIC CASE STUDY OF LINGUISTIC DISTANCES BETWEEN DIFFERENT CHINESE VARIETIES
This work package involved cases studies designed to measure the linguistic (i.e. lexical and grammatical) distances in the three language varieties of Chinese (i.e. Mainland Chinese, Taiwan Chinese, Singapore Chinese) at a certain time period. We used a subset of the Tagged Chinese Gigaword Version 2.0 (i.e. texts from 2000-2003) as the data resources for exploring the linguistic distances from a synchronic point of view.
WP3 DIACHRONIC CASE STUDY OF CONVERGENCE OR DIVERGENCE IN CHINESE AND THE COMPARISON WITH EUROPEAN LANGUAGES
Due to the lack of sufficient historical resources for Singapore Chinese, we restricted the diachronic examination to Mainland Chinese and Taiwan Chinese. The data availability for Chinese allows us for an investigation of the following perspectives:
Convergence/divergence between the formal strata between regiolects (i.e. Mainland Chinese vs. Taiwan Chinese)
Convergence/divergence between the formal strata in each regiolect.