Evaluating GenAI-Assisted Backlog Prioritization in VanVR
Comparing Product Owner Intuition and AI-Generated Reasoning
Sean Jeon ETEC 565 Assignment 2 July 2026
This site is still in progress. I’m building it to share my thoughts and studies, but for now, it’s mainly used for my MET coursework.
Sean Jeon ETEC 565 Assignment 2 July 2026
I work as a Product Owner and Product Manager on VanVR, an educational extended reality platform at the University of British Columbia. VanVR supports three-dimensional anatomy and pathology learning through web, augmented reality, and virtual reality tools. One of my responsibilities is managing the product backlog and helping the team decide which work should be completed first.
Backlog prioritization is not only a technical task. A feature may have strong learning value but require more time than the team can support. A less visible technical improvement may improve access and performance across the whole platform. Ethical access, accessibility, maintenance, stakeholder needs, and available staff must also be considered.
For this assignment, I tested whether ChatGPT could support the prioritization of eight VanVR backlog items. I asked it to arrange the items in an Impact–Effort Matrix, explain its reasoning, state assumptions, and identify missing information. I independently created my own matrix based on my Product Owner experience. The goal was not to let AI make the final decision. The goal was to test whether AI could provide a useful second perspective and improve professional discussion.
I developed the following heuristic from my Product Owner experience, the course readings, and the UBC teaching and learning guidelines. I created it before comparing the final matrices. After the test, I refined several questions so that the heuristic could examine both AI reasoning and my own professional judgment.
The heuristic also uses an adapted known–unknown structure. Lowes (2020) explains that information becomes useful through human interpretation and discussion. In this test, the four areas were:
Each criterion was scored from 0 to 2: 0 means the criterion was not met, 1 means it was partly met, and 2 means it was met well. The score is a discussion aid, not a pass-or-fail decision. A serious privacy, safety, accessibility, or ethical concern cannot be cancelled by a high total score.
This prototype, built with Custom GPTs, explores how AI can facilitate rather than make task prioritization decisions. It guides users through a structured heuristic process, encouraging them to evaluate AI-generated priorities, reflect on their own judgment, and make the final decision themselves.
Click here to test Collaborative Priority Facilitator
Suchman (2023) warns against treating AI as an independent and stable decision-maker. In this use case, ChatGPT is a tool within a larger human and organizational system. Its output depends on the prompt, its training data, and the context that was provided or left out.
Coleman (2021) explains that machine-learning systems can repeat existing categories and social conditions. In backlog prioritization, AI may favour visible features, common ideas about productivity, or majority users. It may give less attention to governance, accessibility, maintenance, or smaller user groups.
Crawford (2021) shows that AI also depends on energy, hardware, data, infrastructure, and human labour. For this reason, the heuristic asks whether GenAI use is proportionate to the task. The UBC guidelines also support privacy, accessibility, transparency, human review, and responsible use (The University of British Columbia [UBC], n.d.).
I used the following VanVR backlog items:
1. Improve the 3D annotation system
2. Add search, filter, and category functions
3. Improve the ethical access gate for anatomy and pathology content
4. Improve mobile augmented reality viewing
5. Improve 3D model compression and loading performance
6. Complete an accessibility review of the web interface
7. Develop multiplayer VR learning functions
8. Improve the screen capture and export function
No student names, medical records, confidential stakeholder information, unpublished research data, or private project records were entered into ChatGPT.
You are supporting a Product Owner for VanVR, an educational XR platform for anatomy and pathology learning. Review the eight backlog items and place each item in an Impact–Effort Matrix. Explain your reasoning, assumptions, uncertainties, stakeholder effects, accessibility concerns, ethical risks, and technical dependencies. Identify information that is missing and ask questions that should be answered before a final decision is made. Do not make the final prioritization decision. The Product Owner will make the final decision after consulting the team and stakeholders.
Figure 1 AI-Generated Impact–Effort Matrix for the VanVR Backlog
Note. Generated by ChatGPT from the VanVR backlog prompt.
Figure 2 Product Owner Impact–Effort Matrix for the VanVR Backlog
Note. Created by the author based mainly on professional experience, tacit project knowledge, and intuitive judgment.
The two matrices agreed on multiplayer VR and screen capture, but several other placements were different. The differences showed where AI used general product-management patterns and where I used VanVR-specific knowledge.
The most important finding was not only that the results were different. The AI and I also reached our placements in different ways. I often placed an item first through experience and intuition, and then found the reasons that explained my choice. The AI appeared to begin with general reasons and relationships, and then use them to support a placement.
Both approaches have risks. My intuition contains useful tacit knowledge, but I may decide first and then select reasons that support the decision. The AI gives a clear chain of reasons, but those reasons may be convincing without being grounded in the real VanVR context. The comparison made both kinds of uncertainty more visible.
This test showed that GenAI can support backlog prioritization, but its strongest value was not the final matrix. Its strongest value was helping me examine how I make decisions as a Product Owner. ChatGPT organized the eight items quickly, used a consistent format, and created useful questions about users, dependencies, accessibility, and technical work. This could help prepare a backlog-refinement meeting.
The comparison also showed that AI reasoning and professional intuition work differently. I often placed an item first through experience and pattern recognition, and then explained why it belonged there. For example, I rated model compression as very high impact and lower effort because I know the current VanVR asset workflow and available tools. This project knowledge was not included in the prompt. The AI appeared to work in the opposite direction. It first produced general reasons about user value, technical difficulty, and ethical importance, and then used those relationships to place an item.
Neither process was fully objective. My experience gave me useful tacit knowledge, but it also created a risk that I would decide first and find supporting reasons later. The AI made its reasoning more visible, but an organized explanation can sound more reliable than it is. Suchman (2023) argues that AI should not be treated as an independent decision-maker. In this test, the output depended on the prompt and on information the model did not receive.
The accessibility review was the clearest example. The AI rated it as high impact and lower effort, while I rated its immediate product impact lower and its effort higher. My estimate included possible testing, consultation, documentation, and development changes. However, the difference also revealed a weakness in the Impact–Effort Matrix itself. The AI was considering ethical and access impact, while I was thinking more about immediate product and development impact. Before placing tasks, the team should define whether “impact” means learning, user reach, ethics, compliance, technical performance, or strategy.
Coleman (2021) helped me consider how AI may repeat common ideas about productivity and value. It may favour visible features or majority users while giving less attention to governance, maintenance, or smaller user groups. My own professional habits can also repeat established priorities. The heuristic was useful because it asked both the AI and me to state assumptions, evidence, and missing information.
Sustainability was another important test. ChatGPT saved time by creating a second perspective, but the output still needed review and revision. Crawford (2021) reminds us that AI use depends on energy, hardware, data, infrastructure, and human labour. For a small backlog change, a spreadsheet or short team discussion may be enough. GenAI is more justified before a major refinement meeting, when its questions and alternative reasoning can add clear value.
My main takeaway is that GenAI should not select the final backlog priority. It can act as a structured second opinion that challenges intuition, reveals assumptions, and prepares questions. Final decisions should come from evidence and discussion with developers, educators, learners, and other stakeholders.
This test did not show that human intuition is always better than AI reasoning. It showed that they use different forms of knowledge and contain different blind spots. My judgment began with experience and intuition. The AI began with general reasons and relationships. The evaluation heuristic helped make both processes open to review.
For VanVR, GenAI is most useful as a second perspective before team discussion. It can help organize backlog items, reveal missing questions, and challenge an early decision. It should not replace Product Owner accountability or stakeholder consultation.
I independently developed the evaluation criteria based on my Product Owner experience, the course readings, and the UBC guidelines. I used ChatGPT to perform the backlog-prioritization test and later to support language editing and document organization. I independently created the Product Owner matrix and reviewed and revised all AI-assisted text. No personal student information, confidential medical information, unpublished research data, or private stakeholder records were entered into the tool.
Coleman, B. (2021). Technology of the surround. Catalyst: Feminism, Theory, Technoscience, 7(2), 1–21. https://doi.org/10.28968/cftt.v7i2.35973
Crawford, K. (2021). Atlas of AI: Power, politics, and the planetary costs of artificial intelligence. Yale University Press.
Lowes, R. (2020). Knowing you: Personal tutoring, learning analytics and the Johari Window. Frontiers in Education, 5, Article 101. https://doi.org/10.3389/feduc.2020.00101
Suchman, L. (2023). The uncontroversial “thingness” of AI. Big Data & Society, 10(2). https://doi.org/10.1177/20539517231206794
The University of British Columbia. (n.d.). Guidelines for all uses of GenAI in teaching & learning. Generative AI at UBC. Retrieved July 16, 2026, from https://genai.ubc.ca/guidance/teaching-learning-guidelines/guidelines-for-all-uses-of-genai-in-teaching-learning/
Untools. (n.d.). Impact–effort matrix. Retrieved July 16, 2026, from https://untools.co/impact-effort-matrix/
This prototype, built with Custom GPTs, explores how AI can facilitate rather than make task prioritization decisions. It guides users through a structured heuristic process, encouraging them to evaluate AI-generated priorities, reflect on their own judgment, and make the final decision themselves.
Click here to test Collaborative Priority Facilitator