|主講嘉賓|
Keynote Speech#1
A theoretical linguistic perspective on the teaching and learning of Taiwanese
An important objective of language teaching is to facilitate learners’ attainment of linguistic competence—defined in theoretical linguistics as the system of generative rules internalized and residing in the brain, and the mental lexicon. Successful internalization occurs when learners engage with primary linguistic data to form and empirically verify hypotheses regarding underlying structures. Although rules are more general and systematic, lexical items are usually arbitrary, imposing a substantial cognitive load on learners. The provision of meaningful input and contextualized application is critical to robust language development.
This presentation surveys the contemporary landscape of Taiwanese language instruction and discusses its challenges. To address some of these challenges, key insight from theoretical linguistics and language teaching pedagogy will be highlighted and possible adaptations proposed, whose applications will be demonstrated through reviewing of publicly available lesson plans and learning materials.
Keynote Speech#2
Revitalising Siraya
The revitalisation of Siraya in Taiwan has been a successful enterprise. It started a quarter of a century ago, and the language is now actively used by a core of language activists and taught in some twenty schools in the Tainan region. Members of the Siraya nation are currently very close to obtaining official recognition as one of the Indigenous peoples of Taiwan. However, the road to success has not always been easy, and the struggle is ongoing. In my presentation I compare the various Siraya revitalisation initiatives that have been undertaken, highlighting their aims, strengths and weaknesses, as well as the difficulties encountered by their protagonists and their prospects for success. An important challenge revivalists initially faced was the scepticism of academics, some of whom did not believe in the concept of language revitalisation, while others had no confidence in revitalisation endeavours undertaken by activists with no formal linguistic training. Others again argued that historically, the language being revitalised did not exactly belong to the region and population for which it was claimed. Or they objected against the methods the activists used in their efforts to adapt the Siraya language to the practical needs of current usage, sometimes simplifying its original spelling and grammar and adopting words from other languages to enable Siraya learners to communicate comprehensively. Other issues complicating the revitalisation of Siraya have been the existence of competing revitalisation projects and the failure of their leaders to work together. My evaluation of the above factors draws on my experience with the grammar, spelling and lexicon of Siraya as well as with the language's actual revival over the years; I also draw upon recent publications on language endangerment and conservation (see publications below).
Adelaar, Alexander. 2011. Siraya. Retrieving the phonology, grammar and lexicon of a dormant Formosan language. Berlin: De Gruyter Mouton.
Adelaar, Alexander. 2013. Reviving Siraya: a case for language engineering, Language Documentation and Conservation (Honolulu) 7:12-34.
Austin, K. Peter & Julia Sallabank. 2018. Language Documentation and Language Revitalization: Some Methodological Considerations (in Hinton, Huss and Roche)
Hinton, Leanne, Leena Huss & Gerald Roche (eds.). 2018. Handbook of Language Revitalisation. London: Routledge.
Olko, Justyna & Julia Sallabank (eds.). 2020. Revitalizing endangered languages: a practical guide. Cambridge: Cambridge University Press.
Keynote Speech#3
低資源語言語音 AI 的工程實踐:臺語、客語與原住民族語的經驗
Engineering Speech AI for Low-Resource Languages: Experiences from Taiwanese, Hakka, and Indigenous Languages in Taiwan
低資源語言的語音人工智慧開發,本質上是一個受到資料與資源高度限制的工程問題。相較於華語、英語等高資源語言,已擁有大規模人工轉寫語音資料、成熟的預訓練模型及完整的評測基準,臺語、客語與原住民族語所能取得的數位語音與文字資源相對有限,因此無法直接複製高資源語言依賴大量資料與大規模模型的開發模式,而需要建立不同的工程方法。
本演講將從工程觀點,分享臺灣低資源語言語音 AI 的開發經驗,並以臺語、客語與原住民族語為主要對象,探討幾個實際問題:在人工轉寫語音有限的條件下,如何建立與擴充訓練資料;如何有效利用既有的預訓練語音模型;如何結合真實語音與合成語音;如何建立一致且可重現的評測流程;以及如何將研究模型進一步轉化為可供後續研究與應用開發使用的語音 AI 系統。
本演講將從資料準備、語音資料擴充、預訓練模型調適、模型訓練、效能評測到實際部署等環節,說明低資源語音 AI 的完整工程思考。核心問題往往不只是如何提出新的神經網路架構,而是如何將有限的真實資料、預訓練模型、合成資料、運算資源及可靠的評測方法有效整合,形成一套穩定且可重現的開發流程。
這些工程經驗提供了一條發展臺語、客語與原住民族語語音 AI 的實務路徑。目標不是等待低資源語言擁有與高資源語言同等規模的資料,而是建立可重現、可擴充的工程方法,充分利用現有資源,逐步將有限的語言資料轉化為真正可使用的語音人工智慧技術。
Developing practical speech AI for low-resource languages is fundamentally an engineering problem under severe data and resource constraints. Unlike high-resource languages, for which large-scale transcribed speech corpora, pretrained models, and mature evaluation benchmarks are readily available, Taiwanese, Hakka, and Indigenous languages in Taiwan often have much more limited digital speech and text resources. Building usable speech technologies for these languages therefore requires a different engineering strategy.
In this talk, I will discuss our experience in developing speech AI technologies for low-resource languages in Taiwan, with particular attention to Taiwanese, Hakka, and Indigenous languages. The discussion focuses on several practical questions: how to construct and expand training data when manually transcribed speech is limited; how to make effective use of existing pretrained speech models; how to combine real and synthetic data; how to establish consistent and reproducible evaluation procedures; and how to turn research models into usable systems for subsequent research and application development.
The talk will present an engineering perspective on the complete development process, including data preparation, speech data expansion, pretrained-model adaptation, model training, evaluation, and deployment considerations. The central argument is that the key challenge of low-resource speech AI is often not the invention of a completely new neural architecture, but the systematic integration of limited data, pretrained models, synthetic data, computational resources, and reliable evaluation into an effective development pipeline.
These experiences suggest a practical pathway toward scalable speech AI for Taiwanese, Hakka, and Indigenous languages. Rather than relying on the availability of massive datasets, the goal is to develop reproducible engineering methodologies that can make effective use of the resources that are actually available and progressively transform them into practical speech technologies.