A GENERAL FRAMEWORK FOR CORPUS-BASED LEXICAL LECTOMETRY RESEARCH
On the basis of an evaluation of previous studies on lectometric measurements, we established a general framework for corpus-based lexical lectometry research that considers most if not all options for different steps. For the general framework, we propose that a proper lexical lectometry research normally should involve the following distinct tasks: (1) compilation of a lectally stratified corpus; (2) sampling concepts as measuring points for lectometry; (3) identification of lexical expressions per concept; (4) disambiguation of lexical expressions in corpus data; (5) calculation of aggregated lexico-lectometric distances; (6) evaluation of measurement reliability and validity.
Another important part of the workflow is a visual representation of the token-based vector space models. We could use various dimensionality-reduction techniques, e.g. multidimensional scaling (MDS), t-distributed stochastic neighbour embedding (t-SNE), or uniform manifold approximation and projection (UMAP), to transform the numerical data in a high dimensional token-by-token similarity matrix into visual patterns on a two-dimensional space (i.e. token clouds) so that the low-dimensional representation retains some meaningful properties of the original data.
LECTALLY STRUCTURED ONOMASIOLOGICAL VARIATION IN THE CHINESE LEXICON
For both the non-weighted aggregation and frequency-weighted aggregation, we found a lectally structured onomasiological variation in the Chinese lexicon. More precisely, the uniformity scores indicate that the pairwise cross-lectal variation on an aggregated level follows an order of “lexical variation between Taiwan and Mainland > lexical variation between Singapore and Mainland > lexical variation between Singapore and Taiwan”.
We modelled the onomasiological choice between three near-synonymous Chinese causative markers, i.e., shi, ling and rang from a cross-variety perspective (i.e. Mainland, Taiwan, and Singapore Chinese). Our findings, which are different from the conclusions based on the studies of English and Dutch causative constructions, show that the choice of Chinese causative markers is mainly influenced by features of the causee and the effected predicate, while those of the causer do not have an important role to play in the alternation of shi, ling and rang. On top of that, the resulting models point to a significant lectal difference in the choice of shi and rang: Mainland Chinese uses shi more frequently while Singapore and especially Taiwan Chinese favour rang. Our model also shows a significant interaction effect between the language variety and the semantic class of cause.
A DIACHRONIC PERSPECTIVE ON LECTOMETRY
The preliminary findings show that the diachronic variation in Chinese lexicons is influenced by features of the concepts, such as concept salience, lexical fields, vagueness and affect. For instance, lexical fields like crime & law and transport seem to be less affected by the diachronic change, while subject to regiolectal difference; the stratificational distances among regiolects are changing progressively in the fields of technology and finance; the fields of social life/relations and entertainment have undergone diachronic change, but are more stable cross-lectally.