in collaboration with Maite Taboada
This project marries corpus-based quantitative methodologies to discourse analysis by investigating the relationship between text complexity and subjectivity as descriptive features in opinionated writing. The specific interest is on how the complexity of subjective text types interacts with discourse features (e.g. argumentative markers, sentiment words, modals) customarily used to characterise these text types.
The database is the Simon Fraser Opinion and Comments Corpus (SOCC) which comprises opinion articles and their corresponding reader comments from the Canadian online newspaper The Globe and Mail. To explore the interplay between different levels of text complexity and various markers of subjectivity, we employ conditional inference trees and random forests. Text complexity is assessed in terms of Kolmogorov complexity which measures the complexity of a text by the length of the shortest possible description of this text (Ehret 2018; Ehret 2017). In this project, we analyse complexity at the overall, morphological and syntactic level. Subjectivity is defined here as the linguistic expression of evaluation and opinion in language (e.g. Hunston and Thompson 2000) and is operationalised as the frequency of lexico-grammatical items which are used to convey subjectivity. Based on the extensive literature on the topic (e.g. Wiebe et al. 2004, Martin and White 1995, Biber and Finegan 1989, Halliday 1985), this set of subjectivity and argumentation markers comprises evaluative words, stance adverbials, connectives and modals.
Funding:
Alexander von Humboldt Foundation "Feodor-Lynen postdoctoral research fellowship"