Chine-Yi Hsiang, National Chengchi University
Abstract
Using a dataset of 2,705 brand-partnered TikTok videos across 63 industries, this study develops an AI-powered framework to predict viewer engagement. Short-form content is represented through high-dimensional semantic and sentiment embeddings extracted from post descriptions and video transcriptions, combined with creator metadata. A MultiOutputClassifier jointly predicts four engagement outcomes using shared representation learning. Ablation analyses reveal that no single modality is sufficient alone; integrating text-based embeddings with sentiment signals and creator attributes yields the highest model performance. The framework offers a replicable, data-driven approach to modeling engagement in short videos. Future work will incorporate visual and auditory features to evaluate aesthetic and sensory effects on viewer behavior. The study aligns with the IS and Media track, presenting a multi-label, multimodal perspective that links platform-generated media content to user behavior in commercial settings.
Keywords
TikTok, Short-Video Commerce, Semantic Embeddings, Sentiment Embeddings