This single-author preprint addresses how AI-assisted outputs should be evaluated in learning-intensive domains such as education, research, and professional work.
The central problem is proxy failure: a polished artifact may be useful, but it may no longer serve as credible evidence of the human understanding, judgment, or transfer ability that the task was meant to cultivate or certify. In other words, an output can look complete while the human capability that should stand behind it remains uncertain.
AI to Learn 2.0 does not propose a blanket ban on opaque AI. Instead, it allows opaque AI during exploration, drafting, hypothesis generation, and workflow design, while requiring that the released deliverable remain usable, auditable, transferable, and justifiable without the original large language model or cloud API.
A key feature of the framework is the distinction between artifact residual and capability residual. Artifact residual concerns how independently and transparently the final deliverable can stand on its own. Capability residual concerns whether the human author can still explain, justify, adapt, and transfer the work beyond the original AI-assisted process. The framework operationalizes this distinction through a five-part deliverable package, a seven-dimension maturity rubric, gate thresholds for critical dimensions, and a companion capability-evidence ladder.
Key Points
Generative AI can produce polished artifacts, but such artifacts may fail to provide reliable evidence of human understanding, judgment, or transfer ability.
AI to Learn 2.0 is a deliverable-oriented governance framework: it evaluates the final deliverable package rather than trying to police every individual interaction with AI.
The framework distinguishes artifact residual from capability residual, separating the independence and auditability of the output from the human capability that remains attributable to the author.
It operationalizes governance through a five-part deliverable package, a seven-dimension maturity rubric, gate thresholds on critical dimensions, and a capability-evidence ladder.
Opaque AI is allowed during exploration, drafting, hypothesis generation, and workflow design, but the released deliverable must be usable, auditable, transferable, and justifiable without the original LLM or cloud API.
In learning-intensive settings, the framework also requires context-appropriate human-attributable evidence of explanation or transfer.
Worked scoring across contrastive cases, including coursework substitution, symbolic-regression governance, teacher-audited national-exam practice forms, and a self-hosted lecture-to-quiz pipeline, shows how the framework separates polished substitution workflows from bounded, auditable, and handoff-ready AI-assisted workflows.
Taken together, AI to Learn 2.0 is proposed as a governance instrument for structured third-party review where capability preservation, accountability, and validity boundaries matter.
The main contribution of this preprint is to move the question from “Was AI used?” to “Can the final deliverable stand on its own, and can the human author still explain, justify, and transfer the work?”
This distinction is important because banning AI entirely may reduce practical relevance, while judging only the polished artifact may fail to verify learning or professional competence. AI to Learn 2.0 offers a middle path: it preserves the usefulness of AI assistance while keeping human capability, accountability, and validity boundaries visible to reviewers.
Seine A. Shintani. AI to Learn 2.0: A Deliverable-Oriented Governance Framework and Maturity Rubric for Opaque AI in Learning-Intensive Domains.
DOI: 10.48550/arXiv.2604.19751
arXiv: 2604.19751 [cs.AI]
Keywords: AI to Learn 2.0, AI governance, opaque AI, learning-intensive domains, deliverable-oriented assessment, maturity rubric, capability evidence, accountability, auditability