GenAI Methods Session 1: How can scientific problems inspire fundamentally new foundations of GenAI?
Anuj Karpatne
Virginia Tech
From AI for Science to the Science of AI: How Scientific Knowledge Can Guide the Next Generation of Generative AI
As Generative AI creates unprecedented opportunities for scientific discovery, scientific applications are also exposing fundamental limitations of current data-driven GenAI systems. This talk will argue that science should not only benefit from advances in AI but also serve as a driving force for inspiring the next generation of GenAI foundations. Building on recent advances in the rapidly growing field of Knowledge-Guided Machine Learning (KGML), this talk will present a roadmap for integrating scientific knowledge through several enabling foundations of GenAI, including model architectures, self-supervised learning, post-training, and reinforcement learning. The talk will discuss promising strategies for building knowledge-guided foundations of GenAI in the context of a variety of scientific applications. This includes modeling the quality of water in lakes across the US and discovering novel biological traits linked with evolution from biodiversity images. The talk will conclude with a discussion of how these ideas establish a two-way relationship between AI and science, where advances in AI accelerate scientific discovery while scientific challenges reshape the future foundations of generative AI by revealing the principles required for GenAI systems to be more reliable, trustworthy, and capable of reasoning beyond observed data.
Amarda Shehu
George Mason University
What Autonomous Scientific Discovery Demands of GenAI Foundations
Generative models have now made proposals cheap, but testing remains expensive. What we choose to test sets the rate of discovery. In an autonomous lab, that choice is a policy: which candidate, which tool, in what order, and under what budget. Situating discovery in an autonomous lab suggests new foundations for generative AI, such as, calibration comparable across modalities and conditioned on earlier results, explanation of decision paths, including skipped tools, and benchmarks that score decisions per unit budget. In this talk, I will offer an information market-based approach with closed-loop allocation policy as inferential arbitrage.
Charuleka Varadharajan
Lawrence Berkeley Laboratory
Using Earth System Challenges to Drive New Approaches for Generative AI
Earth science poses complex challenges that current generative AI models have difficulty in addressing. These include nonlinear physical-chemical-biological feedbacks, extreme spatial heterogeneity, sparse data (particularly in the subsurface), and nonstationary regimes where past data may not be sufficient to predict the future. Scientific discovery and decision-making demand integrating multi-scale physical laws and process-based reasoning with data-driven methods, alongside rigorous uncertainty quantification. This talk will describe how these needs can drive new generative AI approaches, calling for co-design of observation, data, and modeling workflows to yield more robust, physically consistent methods that go beyond current models.
Alina Zare
University of Florida
The Strength to Say "I Don't Know": Why AI Needs Competency Awareness for Scientific Discovery
As AI transforms scientific discovery and becomes a regular scientific collaborator, users must recognize the limits of its competence. Scientific applications often involve novel, ambiguous, or data-limited scenarios where confident predictions can be misleading. This talk presents methods for enabling competency-aware AI that identifies out-of-distribution and ambiguous inputs, supporting more trustworthy scientific discovery.
GenAI Methods Session 2: How do we know when GenAI is actually advancing science?
Aryan Deshwal
University of Minnesota
AI-driven Adaptive Experimental Design for Accelerating Scientific Discovery
Many problems in scientific discovery involve searching over a large number of candidates, where evaluating each candidate is expensive because it requires significant resources. For example, searching the space of materials for a desired property while minimizing the total resource-cost of physical wet-lab experiments or computationally expensive simulations for their evaluation. The key challenge is how to select the sequence of experiments to uncover high-quality solutions for a given resource budget. In this talk, I will discuss adaptive experiment design algorithms to tackle this broad challenge in a variety of complex real-world settings where there are usually multiple objectives, multi-fidelity experiments, feasibility discovery, batch selection and black-box constraints. I will also present results on applying these algorithms to solve problems in domains including nanoporous materials discovery, surfactant design, electronic design automation and additive manufacturing.
Nikunj Oza
NASA
GenAI-enabled Science for NASA
NASA performs and funds research and development for Earth Science, Space Science, Aeronautics, and Human Space Exploration, and many related areas. These areas and their branches cover a large range in terms of many attributes of science: data characteristics (availability, number and types of features, confidentiality, uncertainty, subjectivity), time spans of relevance, domain knowledge, computational needs, etc. This talk will describe this range and some of what is needed for GenAI to contribute to sciences depending on where they fall in this range.
Manish Parashar
University of Utah
Addressing Data Challenges for AI-Enabled Science
The current data renaissance—accelerated by rapid advances in artificial intelligence—is reshaping research, scholarship, and scientific discovery across disciplines. In the age of AI, data is no longer merely an asset or byproduct of computation; it is the foundation of discovery itself. Yet realizing the full promise of data‑driven science requires far more than scale alone. It demands a transdisciplinary approach that integrates diverse data, robust computational infrastructure, and expertise spanning institutions and domains.
Despite unprecedented growth in digital data and access to powerful computing platforms, persistent challenges remain. Barriers in discovery, access, interoperability, governance, and long‑term sustainability continue to limit the impact and reach of data‑intensive science. Even more critically, data‑driven discovery depends on trust. Without clear provenance, transparent data curation and transformation, and confidence that data is fit for purpose and responsibly reused, the foundations of AI‑enabled science remain fragile.
Ram Sriram
NIST
Measurement Science for AI-Enabled Scientific Discovery
Artificial intelligence is rapidly becoming an integral component of scientific discovery, supporting data analysis, hypothesis generation, experimental design, and autonomous scientific workflows. As AI assumes a larger role in the scientific process, rigorous measurement and evaluation become essential for establishing confidence in AI-generated predictions, hypotheses, and decisions. This talk will discuss the emerging role of measurement science (AI metrology) in enabling trustworthy AI for science, including methods for evaluating accuracy, robustness, uncertainty, groundedness, explainability, and human oversight. It will also examine the need for standardized benchmarks, reference datasets, and evaluation frameworks that support reproducibility and interoperability across scientific domains. Finally, the talk will discuss how advances in AI measurement science can provide the foundation for the safe and effective integration of increasingly capable AI systems into the scientific discovery process.
The Future of GenAI for Materials Research
Vincenzo Lordi
Lawrence Livermore National Laboratory
Frontiers of GenAI for Materials Science: Where are we and where are we headed?
I will provide my perspectives on how generative and agentic AI are pushing a new frontier in materials science productivity, where computation increasingly optimizes and accelerates the effort of human experts, assisting and removing humans from the most mundane tasks. The use of AI agents to automate laboratories and orchestrate simulations both create efficiencies and lower the barriers for research. In the future, we expect increased machine intelligence to shift toward nearly-autonomous research including hypothesis generation and testing; however, human ingenuity and creativity likely will continue to guide progress. Will a future exist where AI can autonomously create new knowledge, or will humans need to intervene to prevent AI rot? We sit at a crossroads of such developments for materials science and science more broadly.
Jin Qian
Berkeley National Laboratory
From Generative AI to Digital Twins: Building Trustworthy AI for Materials Discovery
Generative AI has the potential to fundamentally reshape how materials are discovered. Yet the central challenge is no longer simply generating candidates, it is determining which predictions are scientifically meaningful, unique, and trustworthy. This talk explores Digital Twins as a unifying paradigm in which generative AI, physics-based simulations, experiments, and autonomous laboratories continuously inform one another via bidirectional feedback loops. I will discuss open questions surrounding novelty, validation, and scientific trust. I argue that the next generation of AI for materials research will be defined by systems that can reason, learn from feedback, and evolve alongside real experiments.
Taylor Sparks
University of Utah
From Generating Materials to Designing Them: Reinforcement Learning for Scientific Foundation Models
Large generative models are rapidly becoming capable of proposing new materials, but generation alone is rarely sufficient for scientific discovery. In this talk, I will present our recent CrysText framework as an example of how scientific foundation models can represent crystal structures as a language modeling problem, then discuss how reinforcement learning provides a natural mechanism for steering these models toward application-specific objectives. I will conclude with future challenges, including extending generative AI to heterogeneous materials, defects, disordered systems, and synthesis-aware materials design.
Subramanian Sankaranarayanan
University of Illinois Chicago/Argonne National Laboratory
Materials Discovery Cloud – A multimodal discovery engine for microelectronics
Materials discovery for microelectronics requires connecting complex processing conditions, nanoscale structure, defects, interfaces, and device-level function across large multimodal datasets. We are developing the "Materials Discovery Cloud" to serve as an AI-native discovery engine that integrates high-throughput synthesis, in situ and operando characterization, physics-informed simulations, and autonomous learning to accelerate the design of next-generation microelectronic materials. Building on the “AlphaFold for Microelectronics” concept, the platform aims to decode how defects, disorder, interfaces, strain, and processing pathways determine electronic, thermal, optical, and switching behavior. By combining multimodal foundation models, uncertainty-aware active learning, automated workflows, and human-in-the-loop validation, the Materials Discovery Cloud will enable rapid prediction, optimization, and experimental realization of materials with targeted microelectronic functionality. This framework provides a scalable path toward closed-loop materials discovery, transforming fragmented data streams into actionable knowledge for devices ranging from memory and logic to sensors, interconnects, and quantum-relevant materials.
GenAI in Physics and Astrophysics
Stella Offner
University of Texas Austin
AI Reaches for the Stars
I will discuss the road to full AI integration into astronomical workflows: why the field of astronomy is a particularly fertile ground for AI development, the potential revolutionary gains possible for scientific analysis and discovery, how AI developments within astronomy can accelerate AI advancements more broadly, and the current roadblocks to fully AI-integrated astronomy.
Nhan Tran
Fermi Lab
AI and particle physics: from collisions to collaborations
Particle physics is complex in two very different ways. Physically, detectors and accelerators are extraordinarily complex systems — temporally, spatially, extreme in radiation and temperature. Institutionally, the collaborations that build and run these experiments are just as complex: thousands of contributors spread globally across institutions and generations, leaving behind a sprawling codebase and informal records in logs, wikis, and chat archives that no single person could ever hold in their head. This talk introduces the field and generative AI as a tool to address this complexity.
GenAI and Healthcare
Tarek Haddad
Medtronic
Foundation Models and Generative AI in Medical Devices: Opportunities and Validation Challenges
Foundation models are enabling the development of higher-quality patient-facing medical device algorithms by learning from large, diverse datasets and supporting more robust and generalizable models. Generative AI also has significant potential to support medical device development through activities such as documentation, regulatory submissions, evidence synthesis, and engineering workflows. However, the use of generative AI and AI agents in patient-facing applications remains at an earlier stage. New approaches are needed to validate their behavior, verify reliability and consistency, manage uncertainty, and ensure safe performance across diverse clinical scenarios before these systems can be deployed with confidence.
Hamid Tizhoosh
Mayo Clinic
Generative AI: Between Euphoria and Promise
Generative AI has sparked unprecedented enthusiasm by demonstrating remarkable capabilities in producing human-like text, images, code, and other forms of content. Its potential to accelerate innovation, enhance creativity, and transform domains ranging from business to medicine has fueled expectations of a new era of intelligent systems. Yet this excitement is accompanied by fundamental challenges. Hallucinations, the absence of reliable source attribution, embedded biases, and limited explainability raise critical concerns, particularly in high-stakes applications such as healthcare. More recently, multimodal AI has been promoted as a promising step forward by integrating text, images, and other data types to achieve richer contextual understanding. While multimodality offers important advances, it does not by itself resolve the core limitations of generative AI and may even introduce new forms of complexity and error. This talk critically examines whether the current enthusiasm surrounding generative AI reflects genuine long-term promise or temporary euphoria. It argues that trustworthy AI will require not only more capable models, but also evidence-grounded reasoning, transparent attribution, rigorous validation, and clinically meaningful integration into real-world workflows.
Rui Zhang
University of Minnesota
From Generative AI Models to Trustworthy Learning Health Systems
Generative AI offers new opportunities for clinical reasoning and biomedical discovery, but reliable real-world impact requires more than fluent model outputs. This talk will highlight two complementary areas of our work: fine-grained evaluation of medical reasoning and uncertainty-aware clinical decision support. These examples illustrate the importance of evidence grounding, explicit uncertainty, external validation, and human–AI collaboration. The talk will conclude with a perspective on how rigorous evaluation and governance can help translate generative AI models into trustworthy learning health systems.
Generative AI for Agriculture
Rahul Ramachandran
NASA
From Foundation Models to Agentic Workflows: A NASA Strategy for Accelerated Scientific Discovery
NASA's Office of the Chief Science Data Officer addresses the challenges of managing and integrating its vast, siloed, multidisciplinary datasets by leveraging foundation models and agentic workflows through a two-pillar AI strategy: Data-Driven AI Models and Enhancing Scientific Workflows. A key component is the "5+1" roadmap, which focuses on pretraining domain-specific foundation models using flagship datasets to reduce downstream costs, alleviate data bottlenecks, and enable cross-disciplinary synthesis. Additionally, the Accelerated Knowledge Discovery (AKD) concept enhances the research lifecycle by combining language models, agentic orchestration, and researcher expertise. These strategic initiatives are mirrored in the broader scientific community, as highlighted by the Second ESA–NASA International Workshop held in May 2026, where emphasis shifted from the feasibility of training geospatial foundation models to questions of model evaluation, trust, and their integration into end-to-end scientific workflows.
Brian Stucky
USDA Agricultural Research Service
Open-source AI workflows and the future of agricultural research
Open-source software has transformed the life sciences by catalyzing rapid analytical innovation, the adoption of advanced data science and machine learning methods, and a revolution in accessibility and reproducibility. Now, the future of open source in science seems unclear. Over the next few years, some of the biggest leaps in AI-driven scientific progress could come from agentic AI workflows, which automate long-running, complex research tasks that synthesize many sources of information. Agricultural research is often highly multidisciplinary and therefore especially well positioned to benefit. However, the AI models used in these workflows, the data and methods used to develop them, and the agentic “harnesses” supporting their use for scientific research are often proprietary and closed, which threatens innovation, reproducibility, and research security. We need innovative research exploring agentic agricultural research workflows built with open models and supporting software, along with new approaches to make open models and workflows easily accessible to agricultural researchers. Pursuing these goals will help protect research security and ensure that our research outputs are of maximum benefit to farmers, ranchers, and consumers.
Raju Vatsavai
North Carolina State
Embeddings as the Backbone of Modern GeoAI: From Foundation Models to Real-World Applications
The advent of foundation models has paved the way for building cost-effective, downstream geospatial machine learning applications. However, these models are often pretrained on diverse input modalities with varying spatial resolutions, shapes, and architectural designs. Such heterogeneity, combined with differences in parameter counts, significantly affects downstream task performance. In this study, we conduct a systematic evaluation of both general vision and geospatial-specific foundation models, using embeddings extracted from two benchmark datasets across three downstream machine learning tasks. Our analysis provides key insights and offers practical guidelines for effectively integrating embeddings into geospatial AI workflows.
Dimitris Zermas
Sentera
GenAI for Drone Imagery: Lessons from Production Agriculture
Deploying deep learning at scale across tens of millions of acres of drone imagery surfaces constraints that benchmark research rarely encounters. At John Deere and Sentera, production vision models for weed detection, stand count, and crop health must handle unique data distributions, sub-centimeter objects, and agronomic variability that differs by crop, geography, and growth stage. This talk frames where GenAI fits in that pipeline today and identifies three open research problems: generating images that obey biological rules, achieving cross-crop generalization without bespoke pipelines, and building evaluation methods to detect synthetic data bias before deployment. Closing these gaps requires academia and industry working together with agronomic fidelity as the standard.