CORPUS SELECTION
Case studies in WP1 and WP2 used the Tagged Chinese Gigaword Version 2.0 (Huang 2009), which contains newspaper articles of Mainland Chinese, Taiwan Chinese and Singapore Chinese, as the data resource.
For case studies in WP3, we added two corpora, i.e. the Chinese Web Corpus (Mainland Chinese) and the Taiwan Chinese Web Corpus (Taiwan Chinese) on top of data from the Tagged Chinese Gigaword Version 2.0 (newspaper articles). We then took subsets of informal texts like bloggers and online forum articles of the web-based corpora.
SAMPLING LEXICAL VARIABLES AND LEXICAL VARIANTS
We combined both resource-driven selection and vector-driven selection for sampling concepts and the lexical expressions per concept.
Resource-driven: we relied on synsets extracted from Chinese Open Wordnet project
Vector-driven: we used a type-based distributional semantic algorithm (i.e. the Clustering-By-Committee algorithm, see De Pascale 2019) for extracting possible near-synonym candidates
We tried to include concepts covering various lexical fields and applied the selection criteria on internal uniformity (I<0.6) and crosslect uniformity (U<0.8) scores (see Geeraerts, Grondelaers & Speelman 1999).
EXTENDING WORDNET SYNSETS
There are some drawbacks of Chinese Open Wordnet. For instance, some variants might be missing for a certain concept; and it might be biased towards a particular language variety. Thus, we employed type-based vector space models to retrieve potential synonymous lexical variants for a concept.
Type-based vector spaces from different lects (but with the same contextual features) could be used to detect regional variation (cf. Perisman & Speelman 2009).
E.g. Concept POLICEMAN
Calculate co-occurence frequencies between the target word and contextual features (window size: 10L, 10R)
TOKEN-BASED VECTOR SPACE MODELS AS SEMANTIC CONTROL
(1) VISUALIZATION WITH TOKEN CLOUDS
We used the t-SNE technique to transform the numerical data in a high dimensional token-by-token similarity matrix into visual patterns on a two-dimensional space (i.e. token clouds) so that the low-dimensional representation retains some meaningful properties of the original data. Below are the token clouds for the concept COUNTERATTACK.
(2) EXCLUDING THE “OUT-OF-CONCEPT” TOKENS FOR THE LECTOMETRIC STUDY
We relied on cluster analysis to keep those variant tokens which are "in-concept", i.e. semantically equivalent, and to exclude those variant tokens which are "out-of-concept" with the help of cluster analysis. We also excluded clusters with tokens from 1 lect only (not a shared meaning across lects) and tokens of 1 variant only (specialized use without lexical variation).
CALCULATING WORD CHOICE UNIFORMITY ACROSS LECTS
We calculated word choice uniformity across lects by either weighting within profile or aggregation over concepts.
Profile-based linguistic uniformity was used as a measurement for linguistic distance between lects (cf. Speelman, Grondelaers & Geeraerts 2003).
The external uniformity value quantifies the difference between onomasiological profiles in two lects.
References
Geeraerts, Dirk, Stefan Grondelaers and Dirk Speelman. 1999. Convergentie en divergentie in de Nederlandse woordenschat: een onderzoek naar kleding- en voetbaltermen. Amsterdam: P.J. Meertens-Instituut.
Huang, Chu-Ren. 2009. Tagged Chinese Gigaword Version 2.0 LDC2009T14. Philadelphia: Linguistic Data Consortium.
Peirsman, Y., & Speelman, D. 2009. Word space models of lexical variation. In Proceedings of the Workshop on Geometrical Models of Natural Language Semantics (pp. 9-16).
Speelman, Dirk, Stefan Grondelaers and Dirk Geeraerts. 2003. Profile-Based Linguistic Uniformity as a Generic Method for Comparing Language Varieties. Computers and the Humanities 37(3). 317–337.