Here you'll find some documentation & data from previous student projects. Please note that in most instances the datasets are generated from non-probabilistic sampling methods (that is, students engaging in convenience sampling of their peers and connections). As such, analyzing these datasets may be useful for instructional purposes, but no findings should be taken as representative or informative of the broader population! Caveat emptor , and all that.
This dataset comes from a marketing-focused survey examining how young consumers recognize, follow, and respond to social media influencers. Respondents reported whether they recognized a set of nine influencers, whether they follow any of them, and whether they recall those influencers engaging in sponsorship activity. They also indicated which social media platforms they actively use (Instagram, YouTube, Twitch, TikTok, Facebook, Reddit) and whether they engaged in several influencer-related behaviors, such as purchasing recommended products or interacting with influencer content. The survey includes a block of parasocial relationship items—eight Likert-type statements tapping connection, affinity, and perceived closeness with a chosen influencer. Additional variables cover attention checks, demographics (including age, gender, employment, and college status), self-reported social media usage, and specific questions tied to influencer shopping behaviors (TikTok Shop, InstaShop).
Data collection occurred in November 2025. The final, validated dataset includes 1,342 respondents. Qualified respondents were under 46, identified at least one of the studied influencers & recalled that influencer sponsoring/selling a product. Quality control checks were also deployed (timing, ReCaptcha, attention checks). The sample is not representative of the broader social media influencer population; for example, about 44% of the respondents are SDSU students. 63% of the respondents are women.
Sentiment Analysis: The Influencer evaluation open-ended comments were analyzed using a dictionary-based sentiment method built around the VADER (Valence Aware Dictionary and sEntiment Reasoner) framework. In this approach, each written response was broken into individual words and compared to a predefined list of terms labeled as positive or negative. When a word appeared in the VADER dictionary, its emotional value contributed to overall sentiment. These contributions were added and then divided by the total number of words to calculate three key proportions: positive valence, negative valence, and neutral wording. A compound score was also created by combining all positive and negative contributions, resulting in an overall indicator showing whether the tone of the response leaned more positive or more negative. This method allows qualitative text to be converted into numerical values, making it easier to summarize patterns across many respondents. It is especially useful for exploring general tone differences across groups or identifying common emotional reactions. However, this approach has several limitations. It examines words individually and does not fully understand context, sarcasm, or sentence structure. Short responses may appear overly positive or negative because a single strong word can dominate the score. The dictionary may also miss slang, new expressions, or subtle emotional cues. As a result, these sentiment scores should be interpreted as broad indicators rather than precise reflections of each respondent's intended meaning.
PDF Printout of Survey
Survey Preview Link
Dataset - Excel File (n=1342) - Includes codebook, numerical form of raw data, labeled form of raw data.
*Excel file INCLUDES sentiment analysis of the OPEN text field; see final columns.
Dataset - SPSS File (n=1,342)
This study investigates clothing-shopping behaviors, priorities, and brand engagement among consumers who have shopped for clothes within the past year. Respondents report how often they shop online and in stores, how much they spent, and whether they view shopping as enjoyable or unpleasant. A key component is brand selection: participants indicate which retailers they shopped at, including a mix of fast-fashion options (Shein, Boohoo, Forever 21, H&M, Zara) and more sustainability-oriented or secondhand outlets (Patagonia, thrift/vintage stores, thredUP), along with mainstream and luxury labels. The survey then measures how frequently each selected brand is used. A large section assesses shopping priorities such as sustainability, affordability, ethical labor practices, comfort, personal style, trendiness, prestige, and social media influence. Demographic items follow, including age, gender identity, education, employment, income, and California/SDSU affiliation. Overall, the dataset provides a detailed view of how consumers balance convenience, values, and fashion motivations in their shopping decisions.
This survey explores several topics related to the coffeeshop habits of people who visit coffeeshops in San Diego. Most notably, it features brand perceptions and evaluations of three coffeeshops (Dark Horse, Better Buzz, and Starbucks). It also includes measures about coffee drinking and spending habits of respondents. Basic demographic information is captured as well.
Data collection occurred October 2023. The final, validated dataset included 970 survey responses. About 67% of the respondents are 18 to 24 years old, and nearly half of the respondents indicated they are currently attending SDSU. As such, any analysis conducted with this dataset should be approached with appropriate caution; clearly this is not a representative sample of the overall San Diego coffeeshop consumer market.
Word Codebook of the Qualtrics Survey
Dataset - Excel File (n=970) - Includes codebook, numerical form of raw data, labeled form of raw data.
The datafile comprises four primary sections. The first section dealt broadly with how participants consumed and purchased coffee. The second section was an experiment. This experiment manipulated whether the respondent was told that conceptual drawings and text for a coffee shop were generated by humans or with the aid of generative A.I. There were four levels of this independent variable; each level signaled increasing levels of utilization/reliance on generative A.I. to make the concepts for the coffee shop. All respondents saw the exact same coffee shop concept (same name, same slogan, same conceptual drawings of the location and coffee cup brand, and representative menu items). Respondents then provided open-ended and closed-ended evaluative responses about the coffee shop concept. The third section of the survey asked respondents about their preferences and motives toward plant-based non-dairy creamers. The final section included basic demographic questions and participant disclosures about their view towards technology and experience with using generative A.I.
Data collection occurred June 2023. The final, validated dataset included 493 survey responses.
Link to a PREVIEW of the Qualtrics Survey (A/B experiment)
Dataset - Excel File (n=493) - Includes codebook, numerical form of raw data, labeled form of raw data.
This study surveyed 18+ year old adults who had previously tried plant-based food alternatives to meat products. The study dealt with a range of issues related to this topic, including: reasons for eating plant-based alternatives, recognition of plant-based alternative brands, behaviors (spend & amount eaten) toward plant-based alternatives. There was also an A/B experiment about the potential consumer impact of a prominent vegan logo placed on product packaging. Some basic demographic information was included as well.
A PDF printout of the survey is here
Dataset - Excel File (n=397) - Includes codebook, numerical form of raw data, labeled form of raw data.
This study surveyed 21+ year old adults about alcohol consumption and perceptions of alcohol brands. Most of the specific brands studied were hard seltzer brands. The study also tended to focus on assessing the gendered assessment of each brand's identity; specifically, whether the brand was perceived as male, female, or gender fluid, using a mixture of previous brand personality measures and gender fluidity questions used in psychology and sociology studies. Participants were also asked about their perceptions of gender-neutral pronouns and other demographic characteristics.
A PDF printout of the survey is here
Dataset - Excel File (.xlsx) - Includes codebook, numerical form of raw data, labeled form of raw data.
This study surveyed 18+ year old adults about a variety of topics, most of them focusing on issues related to the COVID-19 pandemic and Black Lives Matter movement. Questions about Covid-19 dealt with emotional health, time spent on activities, and money spent on different product categories. Questions about BLM related to general awareness of the social movement, support related BLM, and attitudes about brands engaging in the BLM movement. Participants were also asked about their work status and other demographic characteristics. In total, 401 respondents were judged to have complete, valid responses to the survey.
A PDF printout of the survey is here
Dataset - Excel File (.xlsx) - Includes codebook, numerical form of raw data, labeled form of raw data.
This study surveyed Instagram users. The primary topic related to Instagram Influencers. Respondents were asked about their own level of Instagram usage, attitudes about Instagram Influencers, knowledge (both objective and subjective) about promotional policies related to Instagram Influencers, and some basic demographic information. In total, 372 respondents were judged to have complete, valid responses to the survey.
PREVIEW the survey here (takes you to a test link version of the study)
A PDF printout of the survey is here
Dataset - Excel File (.xlsx) - Includes codebook, numerical form of raw data, labeled form of raw data.