Unified multimodal models are AI systems that can both interpret and generate images and text, but it is unclear whether these abilities produce consistent factual behavior. This project asks whether a model gives the same factual answer when a concept is tested through different tasks, such as identifying an object and generating an image that reflects the same fact. We design an evaluation framework that changes one factual element at a time and compares model behavior across understanding and generation tasks. The evaluation records where models remain consistent, where they contradict themselves, and how performance changes as tasks become more difficult. We anticipate finding that strong performance on individual tasks can hide substantial inconsistencies across tasks, revealing weaknesses that standard evaluations overlook. These results could provide a clearer way to assess general-purpose multimodal AI systems and help researchers distinguish broad competence from task-specific success.
Student Major(s)/Minor: Data Science Major, Applied Mathematics Minor
Advisor: Dr. Jindong Wang
Cancer incidence and mortality vary substantially across Virginia, raising the question of how socioeconomic and environmental conditions contribute to these geographic disparities from 2011–2020. Cancer incidence, mortality, and stage-at-diagnosis data from the Virginia Cancer Registry were analyzed at county and health district levels and combined with U.S. Census socioeconomic measures and environmental exposure data, including fine particulate matter (PM2.5), nitrogen dioxide (NO₂), and water quality. Bayesian spatial models were used to estimate county-level cancer risk while accounting for geographic patterns and multiple risk factors. Results showed that lung cancer burden was highest in Southwest and Southside Virginia and was associated with smoking, poverty, and PM2.5 exposure. Colorectal cancer risk was elevated in rural Southside areas, while breast and prostate cancer showed distinct regional patterns. These findings can help identify high-risk communities and support targeted cancer prevention, screening, environmental health, and resource-allocation strategies.
Student Major(s): Human Health & Physiology Major; Biochemistry Minor
Advisor: Dr. Yi He
This project will explore integrating continual learning methods with reinforcement learning algorithms. A core issue in machine learning is that models need to be retrained repeatedly on data and struggle to retain useful skills when adapting to new tasks or environments, often requiring retraining from scratch, massive datasets, and expensive model training. Reinforcement learning most realistically reflects this problem. In reinforcement learning, an agent must determine how to maximize a reward signal through direct interaction and exploration of the environment. As the agent improves or the environment changes, the agent must adapt to new situations without forgetting previously acquired skills. Current applications of continual reinforcement learning focus on frameworks that restructure the inputs and outputs given to the model, rather than the capabilities of the model itself. These frameworks are often complex and make assumptions about how the introduction of tasks are structured. This project will instead focus on integrating models designed for continual learning into various reinforcement learning systems, identify which integrations are successful and why, and finally develop a continual learning model tailored for reinforcement learning.
Student Major(s)/Minor: Computer Science Major
Advisor: Dr. Haipeng Chen
In the digital age, information has become more accessible than ever, but it has also become more susceptible to abuse and misrepresentation thanks to the rapid-fire nature of modern social media. This project seeks to employ machine learning to investigate the reliability of online networks by modeling the “onset of failure”, that being when a post contains misleading or false information. The goal is to create a predictive model that will classify words and patterns in posts in order to produce a numerical reliability score. The model will be trained on public datasets containing instances from social media and news sources. By analyzing both the language used in posts, and the context in which they were rapidly spread, the project will be able to innovate a direct and measurable method of identifying unreliable statements, and contribute to a broader appreciation of credibility in digital communication.
Student Major(s)/Minor: Data Science Major, Music Minor
Advisor: Dr. Daniel Vasiliu
Ephemeral waterholes are an integral part of ecosystems in the Southern Africa savanna where access to water is limited. Studying the changes in surrounding vegetation in comparison with these waterholes can help to better understand drivers of wildlife movement and assess ecosystem health. For this project, Sentinel-2 satellite imagery was analyzed in Google Earth Engine. Fluctuations in NDVI (Normalized Difference Vegetation Index) around waterhole areas were compared between two time periods (normal rainfall in 2021-2022 and a drought year in 2024), as well as between different spatial zones (0-10m, 10-20m, 20-50m). Expected seasonal patterns in vegetation growth were observed, with increased greenness following wet seasons during periods of normal rainfall, as well as less observable seasonal distinctions and overall decreased NDVI during the drought period. Additionally, NDVI values tended to increase moving farther from the waterholes, and standard deviation of NDVI tended to be greater during wet seasons. These findings contribute to our understanding of vegetation patterns in these under-studied waterhole ecosystems.
Student Major(s)/Minor: Undeclared
Advisor: Dr. Jennifer Swenson
This project explores the disconnect between linking metgenomic data (the DNA that is commonly sequenced) and metatranscriptomic data (looking at the underlying RNA interactions of the environment we are sequencing) which tells us which species in the DNA are actively being expressed. We have focused on addressing the gap between the two types of data and the different species that are active within a sample, as well as investigating current methods of quantifying and dividing species distributions across different samples and timepoints. This led to analyses of current wastewater samples and leading genomic data tools such as Kraken, Metaphlan, and Humann in order to understand and optimize current research standings of MTX classification.
Student Major(s)/Minor: Computational & Applied Mathematics & Statistics (Mathematical Biology) Major
Advisor: Dr. Yanhai Xiong
As Large Language Models (LLMs) become more advanced, they must balance competing goals, such as being helpful versus harmless, or prioritizing accuracy against computational costs. Multi-Objective Reinforcement Learning (MORL) trains models on these multiple goals simultaneously. However, a major challenge is gradient conflict, where opposing objectives cause the model to prioritize the most dominant goal while silencing others. This project addresses how to effectively balance these conflicting objectives without unintentionally altering the user's intended priorities during training. By investigating advantage manipulation methods (such as GRPO and GDPO) and dynamic gradient weighting, we evaluate techniques to properly scale and combine different rewards. We anticipate that dynamic weighting will yield high average performance and greater robustness across diverse tasks. Ultimately, this approach will improve AI alignment, ensuring models behave exactly according to human-defined trade-offs across real-world applications.
Student Major(s)/Minor: Mathematics and Computer Science Major
Advisor: Dr. Haipeng Chen