Abstracts within each theme are listed alphabetically. Click the down-arrows on the right to expand the text, and outside the box to collapse.
Emma Braccini (TU Dresden (TUD)): The Pragmatic Dimension of Scientific Understanding: Understanding-for Climate Extremes
Scientific understanding is traditionally conceived as a robust form of conceptual grasp: the ability to apprehend the relevant relations and structures of given phenomena and to situate them in a coherent theoretical framework. Within this classical view, understanding is treated as a theoretical aim of science, sharply distinguished from applied aims, which evaluate scientific outputs primarily in terms of their practical usefulness (e.g., de Regt, 2017; Khalifa, 2017).
I argue that, unlike other epistemic aims, scientific understanding can serve both as a theoretical achievement and as an action-guiding goal, particularly in some domains like climate change research, pandemic science, and analyses of economic collapse, where understanding is inseparable from responding to urgent, high-stakes, and deeply uncertain conditions characterized by unpredictability.
Looking at established accounts (including Strevens’s (2013) distinctions between understanding-that and understanding-with, and Parker’s (2014) distinction between understanding-why and “a broader notion” suited to complex systems), I suggest that existing frameworks do not fully capture how understanding operates in applied scientific practice. While some forms of understanding are rightly treated as theoretical aims used to evaluate and develop theories, others play a distinct practical role that remains underexplored.
To address this gap, I introduce a novel concept of scientific understanding, called understanding-for. This notion consists in the ability of scientists to use their knowledge of a phenomenon to intervene, adjust conditions, and shape the environments in which they operate in order to achieve specific goals. This pragmatic dimension cannot be fully accounted for by traditional categories of scientific understanding, none of which capture the capacity to respond effectively to a phenomenon. Scientific understanding is not merely a matter of possessing explanatory knowledge, but also of being able to anticipate phenomena and acquire adaptive strategies in relation to them.
Understanding-for is developed through the case study of climate extremes, an area in which international organizations such as the United Nations and the IPCC have emphasized the urgent need for both theoretical insight and practical capacities to respond. Drawing on non-explanatory sources of understanding (e.g., Lipton, 2008), including the manipulation of climate models and the use of analogies and exemplars, this work shows how scientists can achieve a pragmatic form of understanding that supports decision-making and mitigation strategies.
In this context, understanding why such events occur or how they unfold is insufficient on its own. What is also required is a practical understanding, conceived as a competence that enables scientists to apply knowledge to modify, adapt to, or effectively respond to ongoing phenomena. Accordingly, understanding-for enables, for example, the design of evacuation protocols for vulnerable areas and the prioritization of critical infrastructure to be secured first, thereby enhancing the capacity to safeguard lives and reduce damage.
This paper contributes to expanding the notion of scientific understanding by conceiving it also as an ability to use scientific knowledge and information, in specific contexts and under specific circumstances, to act and respond pragmatically to a phenomenon. Only by integrating this dimension can scientific understanding adequately meet the demands posed by a rapidly changing climate.
Benjamin Santos Genta (New York University): No Reproducibility For Systematic Reviews? No Problem.
Systematic reviews are rigorous syntheses of the available scientific evidence with respect to a particular research question. Alongside meta-analyses, systematic reviews are crucial for informing clinical guidelines and regulatory decisions about drug approvals.
Conducting systematic reviews is notoriously labor- and time-intensive, often taking months or even years to complete. The most time-intensive parts involve the screening and selection of relevant the studies to synthesize: researchers often screen thousands of papers to determine which are suitable for inclusion in the review. To accelerate this process, researchers have begun to turn to artificial intelligence (AI) models to automate these time-intensive stages. Tools now being marketed for this purpose include ASReview, Elicit, and Consensus.
Practitioners are concerned about the reproducibility of AI-assisted reviews. For instance, researchers have found that running the same query at different times will yield different results (Bernard et al. 2025), and that full transparency is often unachievable in practice because of the memory demand that would be required. This lack of reproducibility seems to be an issue for the epistemic weight we should assign to reviews. Their credibility is taken to depend on the reproducibility of their methods and conclusions; for this reason, reviews are highly standardized: each step of the process is meant to be transparently documented so that others can, in principle, reproduce the process and its outputs. Even philosophers who have challenged the idea that aiming for reproducibility is necessary in all research contexts—such as Leonelli (2018) and Feest (2019)—typically accept that reproducibility is both desirable and achievable in highly standardized environments, such as randomized controlled trials (and, presumably, systematic reviews).
In this paper, I challenge that this lack of reproducibility is an issue. Using distinctions by Machery (2020), I first show that AI-assisted reviews do not introduce substantially greater theoretical or practical risks of irreproducibility than traditional reviews. Even in human-led reviews, the procedural reproducibility that is demanded of AI-assisted reviews is often unattainable in practice.
While this lack of reproducibility might seem to leave (both AI and human) reviews in an epistemically precarious position, I then argue that this position can be partially or fully recouped by running robustness and sensitivity analyses on the reviews—this can be done even without a strict kind of reproducibility. Traditional human-led systematic reviews are not put under the scrutiny of robustness analyses because of the time that it takes to conduct the screening of the relevant studies. Yet, with AI-assisted reviews, rapid robustness and sensitivity analyses are becoming possible and should be conducted. I sketch how these analyses could be done—for example, prompt robustness (where slightly different prompts are used to select the studies), model-robustness (where different AI tools are used), and quality-robustness (testing whether conclusions change when borderline ‘edge-case’ studies are included or not).
The upshot is a revision to review standards: if minimal reproducibility is often unattainable, then AI robustness testing should become a central criterion of credibility for systematic reviews.
Patrick Greenough (University of St Andrews): Being is Connecting: The Case Study of Concepts
An essentially networked entity (an ENE) is an entity that is defined (in whole or part) by its place within a network. Thus, the identity conditions of an ENE are (wholly or partly) to be specified in terms of its position within a structured system of relations—that is, by its place within a constitutive network. Prime exemplars include: species, social groups, institutions, languages, brands, genres, trends, nations, colours, numbers, logical constants, and even persons. Indeed, any genealogical entity whose identity is fixed (in part) by its place within a taxonomy, lineage, or evolutionary history will count as an ENE.
We are interested in those ENE which are subject to the forces of change and evolution. This kind of ENE faces a serious challenge: the problem of over-connection whereby: (1) ENE are not just defined by their local connections to other entities in the network but by their global connections to all nodes. (2) Consequently, a change in the nature of one node will change the nature of all nodes. Upshot: the identity of a (flexible) ENE is hyper-sensitive to any and all changes in the network—no matter how small and distant (cf. Fodor 1998).
Why think that such hyper-sensitivity is a bad thing? Perhaps it a good thing to be expected and embraced. In other words, why is the Problem of Over-Connection a problem and not simply an articulation of a key feature of ENEs? The main answer to be developed in this paper/talk is: hyper-sensitivity is simply not an inherent feature of a (typical) ENE. This answer is good news for those who think that kind of hypersensitivity is not just implausible or unexpected but deeply problematic. Here the thought is that hypersensitivity makes the identity of ENEs too unstable to feature in useful explanations, too easy to misidentify, and too hostage to distant factors. While we have sympathy for these arguments, they do not play a role in the main message of our story. That message is: ENE's are simply not sensitive to any and all changes in the network.
Our central case study involves concepts, since concepts are networked entities par excellence. Arguably, what we say about concepts carries over to ENEs more generally. There are four key features of the view on offer: Firstly, concepts are flexible entities (which are subject to the laws of elasticity). Secondly, distant nodes in a constitutive network are not constitutively connected to each other via a chain of constitutively connected concepts. Rather, there is a shifting disconnection point in such a series of concepts. Thirdly, given that concepts are to be seen as elastic, it is plausible that any force of revision/replacement to a node in the network is to be seen as a decaying force which does not proliferate unchecked across the whole network. Finally, this still allows for a form of conceptual holism whereby a concept can indeed be connected to every other concept in the network—just not always definitionally so.
Mason Majszak (NYU): The Amalgamation of Evidence in Deeply Uncertain Contexts
Many scientific domains are characterized by complex systems and contexts of deep uncertainty. This can be seen in climate science, macroeconomics, public health, and others, where experts can’t agree on the specific model of certain systems nor on the probability of specific outcomes and scenarios stemming from changes in these systems. In these contexts, expert elicitation protocols have been employed to produce claims under deep uncertainty. However, elicitation protocols, as evidence amalgamation procedures, have gone undiscussed within philosophy of science.
In this paper, I will examine how elicitation protocols are used within a frontier topic in climate science marred by deep uncertainty, climate tipping points. Climate tipping points can be characterized as “critical threshold[s] beyond which a system reorganizes, often abruptly and/or irreversibly” (IPCC 2021, 2251). Due to the limits of computer models (Terpstra et al. 2025), claims in this context often rely, in part, on expert judgment (Lam and Majszak 2022). In practice, an expert elicitation protocol (Kriegler et al. 2009) has been used to produce claims related to tipping points. This elicitation protocol can be seen as a form of evidence amalgamation where experts must individually weigh and amalgamate different lines of evidence to produce their judgments by answering questions. However, these protocols present a unique circumstance where the individual judgments are seen as new pieces of evidence, which are then aggregated into a single judgment. In this way, there are two distinct levels of evidence amalgamation taking place.
I will first argue that the process of producing individual judgments, done by a single expert, can be seen as an instance of classical Bayesian updating (Bovens and Hartmann 2003; Fletcher et al. 2018). However, I contend that the subsequent amalgamation, done through the pooling of judgments into a collective judgment, should not. Within contexts of deep uncertainty, experts do not agree on the model of the system, thus there is no single well-defined conception of the system that each judgment is referencing. Here, the amalgamation of different judgments, while not being able to support specific claims about the way a single system will change, can be used to ground claims about how various iterations of the system could change. Thus, I contend that evidence amalgamation at this second level should not be understood as an aggregation procedure of evidence within a shared model, but rather as a procedure which treats each individual judgment as a heterogeneous model of a single target.
This two-level conception of aggregation can then provide guidance on what questions can/should be asked of experts, and how the collective judgments should be viewed for decision-making. For instance, if a protocol focuses on well-defined scenarios and potential impacts, rather than producing precise probabilities, the collective judgment, as being produced by a set of heterogenous models, can be examined from a robustness like reasoning (Lloyd 2010; Majszak 2025). This could potentially improve the epistemic support for claims regarding the extent of the possibility space for climate tipping point impacts, supporting the justification for decision-making and policy intervention.
Chryssi Malouchou (University of Edinburgh): Identifying heuristic indispensability - deployment realism and different kinds of 'indispensabilities'
To overcome the antirealist challenge from the history of science, deployment realism (the ‘divide et impera’ strategy, in Psillos’ (1999) terms) argues that theoretical components that play an indispensable role in bringing about a theory’s novel predictions are typically retained across theory-change and worthy of realist commitment. Yet, despite the crucial role that the notion plays in supporting deployment realism, the meaning of ‘indispensability’ has remained ambiguous. In this paper, my aim is to delve into and clarify the notion of ‘indispensability’ underpinning deployment realism.
More precisely, I aim to show that ‘indispensability’ is a blanket term that covers three distinct notions: material indispensability, explanatory indispensability, and heuristic indispensability. Material indispensability is a mechanical notion that captures any kind of knowledge claims that are minimally sufficient to lead to the predicted phenomenon. Explanatory indispensability refers to the claims that provide a minimally sufficient explanation of the predicted phenomenon. Finally, heuristic indispensability refers to the claims strictly without which scientists could no longer predict the phenomenon from the theory at hand, in a practical sense. Current accounts conflate these different notions, and, as a result, struggle to provide an effective recipe for identifying indispensable components that lives up to the promise of realism.
More specifically, based on these distinctions, I try to show the following. First, I argue that both material and explanatory indispensabilities cannot be used to support deployment realism, as they fail to guarantee belief in the distinctively theoretical parts of theories. Second, I attempt to show that current deployment realist strategies used to identify indispensable assumptions either collapse into material or explanatory indispensabilities, or attempt to safeguard belief in the distinctively theoretical parts of theories in an arbitrary way. Thus, I argue that heuristic indispensability is the only kind of indispensability that can support deployment realism. I end the paper by arguing that a counterfactual-based strategy (in line with a proposal made by Baron (2016)) can be more effective for identifying (heuristically) indispensable assumptions. This counterfactual strategy allows us to safeguard belief in the distinctively theoretical parts of theories in a non-arbitrary way.
Lucy Mason (University of Pittsburgh): Methodological Intersubjectivity
Intersubjective agreement is a cornerstone of scientific practice, particularly agreement on measurement results and experimental data. Science could not function without shared results, verification, and replication. At the same time, measurement involves irreducibly individual, subjective, contingent, or perspectival elements tied to specific instruments, experimental contexts, or agents.
Structuralist accounts of measurement have identified calibration as a prominent mechanism by which we achieve intersubjectivity (e.g. Mari et al. 2012). However the exact role of calibration, and the nature of the intersubjectivity it creates, is in need of unpacking. In this paper, I formally define methodological intersubjectivity, describing how measurements are coordinated and how we construct shared measures through calibration (as well as how we define and model measurable properties, which is – as I will show – a process closely intertwined with calibration). I will look particularly at time measurements in a relativistic world, where time is fundamentally local yet global intersubjective standards – such as International Atomic Time (TAI) – are nonetheless achieved. Time measurements both exemplify and constitute a large part of the methodologies involved in calibration, as time is a fundamental property widely used across science and integral to our international system of units. Examining the roles of institutions such as the International Bureau of Weights and Measures (BIPM), Global Navigation Satellite Systems (GNSS), and the International Earth Rotation and Reference Systems Service (IERS), I argue that these services can be understood through the application of an expanded form of Massimi’s (2022) notion of the interlacing of different communities and perspectives, extending it from historical lineage to the coordination of individual measurements and the concrete links formed between measuring devices.
Analysing calibration in this way pushes back against the recent characterisation of intersubjectivity in structuralist accounts of measurements as subject-independence. Instead, measurements should be understood relative to complex networks of many agents and measurement devices working together. This point links philosophy of measurement to ideas in feminist philosophy of science about the social nature of knowledge. Exploring how this manifests in the case of calibration gives us a better handle on recognising how values and subjectivity feature in measurement at both the individual and institutional scales, and the interplay between these levels.
I will also push this point further to explore the physical nature of these networks in addition to their epistemic role. We define measurable properties to be by nature intersubjective and create tangible connections between measuring devices through calibration, establishing intersubjectivity not just on an epistemic level but as a physical reality. This also allows us to design measurement techniques that go far beyond individual instruments. In this sense, a single measurement device is akin to a mushroom showing above ground with a vast mycelium below. Work done to modularise calibration for ease of use often disguises this fact and prevents proper recognition of its significance. This is particularly relevant for discussions about the perspectival nature of measurements and observation, raised for example in Giere (2006) and in recent discussions of intersubjectivity in the philosophy of quantum mechanics.
Johannes Nyström (Stockholm University): Globalizing the no miracles argument
According to the core premise of the no-miracles argument (NMA), the best explanation of the predictive success demonstrated by many scientific theories is that those successful theories are at least approximately true (Putnam 1975, Boyd 1981). The NMA then proceeds via an abductive step to the conclusion that scientific realism is probably true.
On the basis of a reconstruction of the NMA in Bayesian epistemology, Howson (2000, 2013) claims that the argument falls prey to the base rate fallacy and is therefore logically invalid. Since it does not take into account base rate information about approximate truth, there is no way to assess whether the realist implication of predictive success is in fact a sufficient justification of scientific realism.
In response to Howson, Dawid and Hartmann (2018) claim that his reconstruction of the NMA is incomplete. They extend Howson's Bayesian model of the argument by explicitly taking into account the success rate of scientific theories in the targeted research field. They show that if this rate is sufficiently high, the argument does not fall prey to the base rate fallacy. They call this the frequency-based NMA.
In this paper, I show that, as formalized by Dawid and Hartmann, the frequency-based NMA is susceptible the so-called look-elsewhere effect (LEE). The LEE occurs when the significance of a strong data set is in fact strongly reduced by the large size of the parameter space being searched for a result. In the context of the frequency-based NMA, focusing on a scientific research field (parameter value) that demonstrates a high frequency of success and ignoring insignificant results in other fields amounts to falling prey to the LEE.
In fact, the existence of the LEE allows the anti-realist to reject the logical validity of the NMA, as formalized by Dawid and Hartmann, on the basis of the familiar charge that the argument is the result of a partisan selection process (e.g., van Fraasen 1980, Magnus and Callender 2004). Hence, understanding and controlling for the impact of the LEE is of crucial important for the status of the NMA as a logically valid argument in a Bayesian framework. The main aim of the paper is to demonstrate why the LEE arises in this context, and to discuss the implications of ignoring it. In particular, I show that the characteristics of the scientific enterprise and the realist position are exactly of the kind that gives rise to a potentially strong LEE. The effect may thereby be expected to play a significant role in this context.
Finally, I argue that, given the structure of the realist position, the only plausible way to resolve concerns about the LEE is to extend the data input of Dawid and Hartmann’s model of the frequency-based NMA beyond the targeted research field. In effect, this amounts to adopting a fully global NMA, in the sense of Carrier (2004) and Henderson (2017). I explain why this move allows the disagreement on the selection process that stems from the LEE to be resolved.
Malvina Ongaro (Politecnico di Milano): A notion of uncertainty as conflict of reasons
Uncertainty is a pervasive feature of life. The pervasiveness of uncertainty makes it central to many debates in both philosophy and the sciences, which means that the concept has been addressed from different perspectives and with a plurality of labels. And yet, while there are several frameworks to identify different types of uncertainty, to design instruments to represent them more or less quantitatively, and to explore their implications for scientific practice and for decision-making, there are almost no general accounts of what the concept of uncertainty in itself is. In this paper, I propose an account of uncertainty as based on conflicting reasons that goes beyond a purely epistemic notion of uncertainty. This can be used to ground a taxonomy of types of uncertainty which highlights how they may each require different approaches.
First, I present the account of uncertainty in standard Bayesian decision theory. The default view claims that all the uncertainty faced by the agent can be represented with a single probability function. This view is limited both in the severity and in the quality of the uncertainty that it covers. Criticisms along the first line have a long history and have led to refinements and expansions of the tools of the theory, but studies on the qualitative variety of uncertainty are still limited. Looking at this variety supports the idea that we need a more general account of uncertainty, one that goes beyond ignorance about the actual state of the world.
Thus, I argue that uncertainty can concern both cognitive and non-cognitive attitudes. I present a definition of uncertainty as having conflicting reasons for mutually exclusive alternatives attitudes. I take attitudes to be intentional mental states with propositional content, and I include choice among them.
Given this account, dealing with uncertainty means dealing with conflicting reasons. Thus, I move to conflict and present two broad types of conflict: what I call \textit{amenable} conflict, which tends to reduce with the improving of epistemic and cognitive conditions, and \textit{radical} conflict, which persists even under ideal circumstances. Given the nature of this distinction, we should not expect to resolve uncertainty arising from radical conflict through empirical inquiry or logical analysis.
I suggest that radical conflict is always possible between reasons supporting non-cognitive attitudes, and that it is possible with reasons supporting cognitive ones in three cases: when the content of the attitude does not have a truth value, when its truth value is epistemically unaccessible, or when it can be practically considered unaccessible within the horizons of the decision at hand.
This discussion is used to ground a typology of uncertainty on the nature and the content of the attitudes involved. Given that not all these types of uncertainty can be resolved with evidence or reflection, then the classification supports the claim that we should not treat all the uncertainty faced in decisions in the same way.
Arkadiusz Synowczyk (Università della Svizzera italiana): Inductive Relations Without Rules: A Response to Segarra
Norton’s [1, 2] material theory of induction (MTI) holds that inductive arguments are warranted by domain-specific facts rather than universal rules. Segarra [3] has recently argued that, while MTI correctly identifies the source of warrant, it lacks an account of the inductive support relation. In response, he proposes a hybrid theory of induction (HTI) on which support is explicated in terms of rules whose legitimacy is determined by Norton’s material facts. This paper defends two claims. First, if Norton’s critique of rule-based approaches is sound, then HTI inherits the very problem it is meant to solve. Second, a satisfactory response to the insufficiency of MTI requires, by Norton’s own strategy, treating inductive support as grounded in a direct experiential relation to the conclusion.
Norton’s central objection to rule-based approaches is that they cannot generate evidential force from within [1, 2]. Bayesian calculi, for example, define quantities such as P(H|E) and impose coherence constraints on updating, yet nothing in the formalism determines which priors, likelihoods, hypothesis spaces, or independence assumptions are warranted. To have evidential force rather than remain a purely formal structure, Bayesian accounts must invoke external assumptions that either presuppose that evidential relations obtain or themselves constitute claims about what supports what [1]. If this diagnosis is correct, then rule-invoking accounts, including HTI, can at best model inductive support once it is secured. They cannot establish that a rule constitutes a relation of support rather than merely codifies prior commitments.
The upshot is that responding to the insufficiency of MTI and HTI requires abandoning Segarra’s [3] assumption that only rules can supply an account of inductive support. This, in turn, depends on identifying the role of rules in the problem of induction and what Norton’s rejection of them opens up. The sceptic treats non-ampliative inference as comparatively unproblematic and measures induction against this standard. Since inductive premises do not entail their conclusions, she insists that if such reasoning is warranted it must be underwritten by a principle that renders it valid. If one follows Norton [1, 2] in resisting this demand, one thereby gives up entailment as the standard for support. The crucial question is then why entailment was required in the first place.
As Hume [4] himself acknowledges, the demand for entailment arises because the connection between premises and conclusion in induction is taken to be neither logical equivalence nor direct experience. Logical equivalence is an implausible option. The remaining candidate is a direct but defeasible experience with the conclusion as its content, an inferential seeming that p [5]. This strategy, however, is one that both Norton [2] and Segarra [3] appear to reject, since both insist that their project concerns the logic rather than the epistemology of induction and therefore avoid appeals to facts about consciousness. I conclude that once formal approaches are denied evidential self-sufficiency, this separation cannot be sustained. For better or worse for MTI-inspired theories, the logic and epistemology of induction coincide at a foundational level.
Martin Voggenauer (Radboud University): Naturalized Grounding and the Fate of Causal Exclusion
In this talk, I explore how an appropriate conception of naturalized grounding—concerning how higher-level causes are grounded in lower-level analogues—might offer a way out of the generalized exclusion problem.
The exclusion argument claims that lower-level causation reduces higher-level causation to mere epiphenomena. Given the causal closure of the lower level—that every lower-level effect has a sufficient cause within it—and the rejection of systematic overdetermination, higher-level causation seems epiphenomenal. Since each lower-level effect is already fully caused by a lower-level cause, there seems to be no causal work left for higher-level causes.
The argument’s force depends on supervenience, which modally relates higher-level phenomena to lower-level realizers. One strategy for rejecting the exclusion argument replaces supervenience with the non-modal asymmetric concept of grounding. Kroedel & Schulz (2016) and Stenwall (2021) argue that this shift renders the causal overdetermination invoked by the exclusion argument weak or harmless. Yet their approach merely reinterprets overdetermination rather than eliminating it.
Instead, I argue that grounding, properly understood, can dissolve causal overdetermination. Focusing on event causation, I propose that higher-level event tokens are grounded not only in their actual lower-level realizers but in a broader set of possible realizers. To this end, I introduce naturalized grounding between events based on multiple rather than actual constitution. Multiple constitution holds that an entity can be simultaneously constituted by different collections of parts (Kistler 2009; Jones 2015). Though often applied to objects, the concept extends to events as relata of causal relationships.
Accordingly, a higher-level event token is not identical with the lower-level tokens that realize it but corresponds to a range of possible realizations capable of yielding the same higher-level outcome. What makes an event higher-level is its abstraction from the specific details of its actual realization. A higher-level event thus encompasses both its actual lower-level realization and others that could yield the same outcome. This notion of multiple constitution concerns event tokens, not types. Because higher-level events are multiply constituted, the actual constitution relation is asymmetric: lower-level event tokens constitute the higher-level event token, but the latter is not reducible to any single realization.
This framework neutralizes the exclusion argument by rejecting the assumption that a higher-level cause of a higher-level event must also cause its underlying lower-level events. While higher-level effects are grounded in lower-level events, those lower-level events are caused by specific lower-level causes, not by the higher-level cause itself. Given multiple constitution, higher-level causes encompass distinct possible lower-level realizations that yield distinct lower-level effects yet the same higher-level outcome. This blocks the inference that higher-level causation entails causation of its actual lower-level realization, thereby avoiding systematic overdetermination of lower-level and higher-level effects by both levels of causes.
Ultimately, the exclusion argument’s success depends on how we understand the grounding relation between higher- and lower-level causal relata. If grounding is conceived in appropriately naturalized terms, higher-level causation can be genuinely efficacious rather than merely epiphenomenal.
Tvrtko Vrdoljak (University of Cambridge): When Abnormality Misfires: Causal Selection and Disciplinary Baselines
Explanatory practice routinely singles out some factors as causes while treating others as background conditions. Abnormality-based accounts of causal selection, most influentially associated with Hart and Honoré (1985), explain this practice by appealing to deviations from a normal course of events. Such accounts promise to vindicate causal selection as principled rather than arbitrary or merely pragmatic. However, they rely on an often-unexamined presupposition, namely that the baseline of normal conditions relative to which abnormality is assessed is itself epistemically well-founded.
This paper argues that abnormality-based causal selection can systematically misfire when standards of normality are fixed upstream by disciplinary convention rather than by responsiveness to the explanandum. In such cases, entire categories of causally relevant factors may be pre-classified as background conditions independently of their actual role in producing the phenomenon to be explained. The result is not simply that abnormality judgments are context-relative, but that the context itself becomes epistemically rigid and insufficiently responsive to evidence.
I develop this argument through a case study drawn from historical explanation. The paper offers a reconstruction of Kyle Harper’s The Fate of Rome (2017), an account of the Roman Empire’s collapse which integrates climatic variability and pandemic disease alongside more familiar political, economic, and military factors. Standard historiographic explanations typically treat climate and disease as part of the normal background of social life, foregrounding intra-social dynamics as the primary difference-makers. Harper’s account reverses this orientation. Drawing on an expanded evidential base, he treats climatic and epidemiological disruptions as abnormal shocks that rendered existing social stressors fatal. Interpreted through the lens of abnormality-based causal selection, this shift is best understood not as the addition of new causal variables, but as a revision of inherited assumptions about what counts as normal in historical explanation.
Methodologically, the paper proceeds by rational reconstruction, redescribing Harper’s explanatory practice in the conceptual vocabulary of philosophy of science in order to make explicit its epistemic logic. The analysis shows how sociocentrism functions as a salient example of baseline-fixing. By tacitly treating nonhuman dynamics as normal background conditions, sociocentric explanation systematically distorts judgments of abnormality and causal salience. The paper further argues that triangulation, understood as the coordination of methodologically independent lines of evidence, plays a corrective role in this process. By integrating textual sources, archaeology, palaeoclimate proxies, and ancient DNA, Harper’s evidential strategy makes previously obscured deviations empirically salient, thereby enabling the recalibration of baseline expectations.
The broader philosophical upshot is twofold. First, abnormality-based causal selection is only as epistemically reliable as the baselines on which it depends. Second, disciplinary domain restrictions, even when motivated by reasonable pragmatic or organisational considerations, introduce a structural epistemic vulnerability when they function as mechanisms for fixing standards of normality in advance of inquiry. The paper thus connects debates about causal selection to the role of discipline-based assumptions in shaping explanatory judgment, identifying a general failure mode of abnormality-based explanation and clarifying the conditions under which it can be corrected.
Karla Weingarten (Radboud University): On the Productiveness of Idealizations: A Case Study in Effective Theories
Within the plethora of models that scientists employ daily, one would be hard-pressed to find a model that is an exact copy of its target. Most models deliberately idealize: they omit unnecessary details or isolate a specific mechanism that is present in the target. Such simplifying or abstracting distortions are introduced to serve the goal of a modelling process. Recent accounts of representation accordingly do not hinge on the similarity of the two systems, but rather on the model's capacity to facilitate inferences about the target which can lead, for example, to scientific understanding (Suárez, 2024). In other words, these accounts emphasize the model-to-target pipeline. At the same time, idealizations are generally considered falsehoods and thus thought not to actively contribute to a model's success. This focus on the distortive target-to-model relation renders idealized models apparently epistemologically inferior to more complete alternatives.
In this talk, I push back against this view and argue that when we ask whether idealized models can provide understanding, we should pay more attention to the model-to-target inferences rather than focusing exclusively on the target-to-model distortions. Building on work in the literalism vs. factivism debate (Frigg & Nguyen, 2021; Lawler, 2021), I reason that idealizations carry information in two distinct ways: in addition to the extrinsic element that is a target falsehood, idealizations also have an intrinsic informational content that can be considered to be independent of the specific target system studied, and that remains truthful even when put in relation to that target. This makes idealizations not merely distortions of the target, but carriers of true information. As such, idealizations do not hinder but can actively contribute to scientific understanding, and thus be labelled productive (Weingarten, forthcoming).
I support this reasoning with a case study of effective field theories (EFTs). These non-fundamental theories contain a cutoff Λ below which they provide an accurate description of the phenomenon studied but above which they lose validity. EFTs are usually not considered to convey fundamental descriptions of our universe: the cutoff can be interpreted as a limit to our knowledge (Rivat & Grinbaum, 2020). Can true understanding of a phenomenon be obtained even when the theory consulted is known to have a limited validity? Or is understanding, after all, obtainable only through fundamental theories?
This practical situation mirrors the theoretical considerations made in this talk. We can consider the cutoff Λ inherent to an EFT to be an idealization. It contributes productively to understanding by carrying information about the relevant energy scales and interactions for the target phenomenon modelled. Not only can we conclude that EFTs can provide scientific understanding but also that the cutoff can be characterized as productive; without this distortion, the same understanding would not be attainable. Understanding is not correlated to the number of idealizations a model contains, or, in the EFT context, the fundamentality of a theory, but to the inferences it allows. More complete models are not immediately superior; idealized models can provide understanding just as well.
Yuyou Wu (University of Leeds): Reconciling Artifactualism and Representationalism: An Enactive Account of Scientific Models
Philosophers have long sought to explain what scientific models are and how they yield knowledge. Traditional representationalism treats models as representations of target systems, deriving epistemic value from this relationship. However, attempts to explicate this relationship run into persistent difficulties, leading some to declare it a “dead end” (Sanches de Oliveira [2018]). Dyadic accounts—such as isomorphism or similarity—miss representation's directionality and capacity for error, while triadic, agent-dependent accounts risk arbitrariness, undermining their capacity to explain how we learn about targets from models.
A recent response from artifactualism treats models as primarily material tools, emphasizing how concrete engagement with them scaffolds scientific inference (e.g., Knuuttila [2011]; Sanches de Oliveira [2018], [2022]). Yet, scientific practice does not involve a random exploration of a model’s countless sensorimotor affordances—the possibilities for action it offers us as embodied agents with specific sensory and motor capacities. Rather, it is systematically guided and constrained. The selection and coordination of specific sensorimotor engagements, as well as the inferences made about the world based on them, are guided by representational norms and meanings. Existing artifactualist accounts, by deliberately avoiding representational terminology, lack the resources to explain such systematic selection and coordination. Furthermore, supplementing them with traditional representationalism simply returns us to the aforementioned difficulties.
I propose a solution via an enactive account. Enactivism—a framework within embodied cognitive science—understands cognition not as the manipulation of internal neural representations, but as the exercise of skillful know-how in embodied, situated action; a process termed “enaction” or “sense-making” (e.g., Varela et al. [1991]; Thompson [2007]). Through this activity, an organism cultivates adaptive skills, becoming attuned to relevant environmental aspects and transforming that environment into a domain of salience and meaning relative to its embodied capacities. Thus, cognition is a history of structural coupling that “brings forth” a domain of significance, where meanings and norms are co-constituted through dynamic interaction rather than being pre-given.
Building upon this foundation, I argue that many scientific models should be understood as enacted representations—representations “brought forth” through the embodied, situated practice of scientific modeling. They are representational artifacts whose specific content, relations, and norms are not pre-specified but instead emerge from this practice as an advanced form of embodied sense-making. This process is itself situated within and conditioned by scientific practice as a history of structural coupling between humans and nature, through which we explore and become increasingly attuned to nature’s varied aspects by developing relevant, embodied skillful know-how.
This enactive account reconciles artifactualism and representationalism, incorporating the former's focus on material engagement and the latter's commitment to representational content. Crucially, it explains their integration: representational relations, meanings, and norms are emergent order parameters that arise from and, in turn, impose top-down constraints on lower-level constituent processes—whether sensorimotor, inferential, or interpretive. This explains the selective, coordinated patterns of material engagement with models as concrete artifacts. By grounding representation in embodied practice, this account provides a coherent, naturalistic framework that captures how models, as inextricably artifactual-representational entities, enable scientific learning and understanding.
Ziren Yang (University of Miami): Surrogative Reasoning under Referential Indeterminacy
Many accounts of scientific representation presuppose that models represent targets that exist and possess at least minimal determinacy, what I call the “determinacy presumption.” Similarity accounts require models and targets to be similar in relevant respects. Structuralist accounts appeal to structural mappings between models and targets. Some inferentialist accounts presuppose denotation relations between models and targets, as does the DEKI account. This presumption grounds representational relations and enables accuracy assessments: accurate representation requires that targets possess the modeled features, while misrepresentation occurs when they lack them.
Exploratory models with referential indeterminacy (ERI models) challenge this presumption. These models target hypothetical entities whose existence and specification are epistemically uncertain. Early cosmological simulations modeled hypothetical cold dark matter universes, despite deep uncertainty about the initial conditions and the nature of dark matter. Origin-of-life models target self-assembling lipid vesicles whose existence, composition, and evolutionary capacity remain unclear. What the target is—or whether there is a target at all—remains unresolved.
Despite lacking determinate targets, ERI models function as representations in scientific practice. Scientists draw inferences from model dynamics, assess assumptions as claims about physical or prebiotic possibility, and refine conceptions of potential targets. Yet standard accounts cannot explain how these models represent when targets are indeterminate. How can ERI models function as scientific representations when their targets lack the determinacy required for representational relations and accuracy judgments?
I argue that scientific representation admits a zetetic dimension: representation operative under determinacy asymmetry between relatively determinate model systems and indeterminate targets. Models are determinate because modelers construct them through explicit assumptions. Targets remain indeterminate because their existence and properties are uncertain. This asymmetry enables exploratory modeling: models can investigate possible target specifications precisely because targets remain indeterminate.
In the zetetic dimension, representation is evaluation-based rather than accuracy-based. Models are evaluated by assessing whether they legitimately support inquiry into possible target specifications. Evaluation proceeds through independently warranted empirical and theoretical resources—established principles, observed phenomena, mathematical derivations—whose relevance to indeterminate targets is justified by the inquiry’s specific aims and background assumptions. Through such evaluative practices, targets are progressively specified, rendering representation norm-governed despite referential indeterminacy.
Zetetic representation is graded: it weakens as targets become more determinate, while standard representational relations strengthen. The zetetic view extends rather than replaces existing theories. Standard accounts apply when targets are sufficiently determinate; zetetic representation applies when determinacy asymmetry is significant. Scientific representation exists along a continuum from wholly exploratory to wholly descriptive.
Jonathan Bain (NYU): The Explanatory Role of Entanglement Wedge Reconstruction in the Black Hole Information Loss Paradox
This talk addresses a puzzle associated with a proposal (Penington 2020, Almheiri et al. 2019) that claims to resolve the black hole information loss paradox. The proposal consists of two parts: The first part uses the RT formula from the AdS/CFT correspondence to derive the Page curve for the entanglement entropy of the Hawking radiation emitted by an evaporating black hole, and this suggests that information is not lost. The second part appeals to another result from AdS/CFT, namely, entanglement wedge reconstruction (EWR), to explain how information escapes by becoming encoded in the radiation during the late stage of evaporation. The proposal insists that both parts are necessary to resolve the paradox. This is puzzling since one can show that the RT formula is equivalent to EWR (Harlow 2017). That is, if the RT formula holds as a relation between the bulk and the boundary in AdS/CFT, then so does EWR, and vice-versa. Thus, one would think, if EWR explains information escape, so too should the RT formula. I indicate how this puzzle can be addressed by keeping track of different versions of the paradox, and by considering the types of explanation of information escape that EWR and the RT formula might be said to provide.
I first set the stage with reviews of what Wallace (2020) calls the Page time and firewall paradoxes, and the derivation of the Page curve using the RT formula. I then review EWR and how it is supposed to resolve the Page time and firewall paradoxes. Here I consider the suggestion that EWR provides a microphysical explanation of information escape that the RT formula by itself does not provide (Cinti & Sanchioni 2024). I then review Harlow's (2017) proof of the equivalence between EWR and the RT formula, which suggests that anything EWR can explain, the RT formula can explain, too. My conclusion is that EWR does indeed provide an explanation of information escape and how the information loss paradox can be resolved (in either its Page time or firewall versions). But so, too, and to the same extent, does the derivation of the Page curve using the RT formula.
Jonathan Fay (University of Bristol): Whence the desire to close the universe
Recent scholarship, especially de Swart (2020), has shown that cosmologists in the 1970s and 1980s often favoured flat or closed cosmological models for essentially non-experimental reasons, a prejudice that helped shape attitudes toward dark matter in this period. Some of these preferences were motivated by appeals to Mach’s principle. Since previous scholarship has tended to conflate arguments grounded in Mach’s principle, I will closely examine and disentangle two Machian arguments for particular cosmic geometries, one favouring a closed universe and the other favouring the flat Einstein–de Sitter model. In both cases, I will assesses their historical roots, philosophical motivations, and scientific legitimacy.
The first case concerns the long-standing argument for a closed universe—popularised in the renaissance period by John Wheeler—which was rooted in Einstein’s attempt to eliminate boundary conditions in general relativity on Machian grounds. Drawing on a close reading of Wheeler’s early relativity notebooks and responding to questions raised by Blum and Brill (2020), I show that Wheeler’s views were substantially influenced by Hermann Weyl’s (1924) dialogue Massenträgheit und Kosmos despite Weyl’s explicit objections to the argument for closure. I argue that the Einstein–Wheeler case for closure relied on assumptions drawn from pre-expansion cosmology that fail to carry over to the modern context, yet persisted due to their superficial philosophical appeal.
The second case examines a lesser-known argument for a flat universe associated with Dennis Sciama, as well as Hermann Bondi. This argument derives from a Machian interpretation of frame-dragging effects in general relativity, according to which inertial forces in rotating frames should be identified with gravitational effects generated by cosmic matter. In linearised models, this leads to a constraint requiring the quantity Gρτ2 to remain constant over cosmological evolution, a condition uniquely satisfied by the Einstein-de Sitter model among FLRW cosmologies. I show that while Sciama’s early work on a linearised model (Sciama, 1953) attracted some interest in this line of reasoning, later attempts to extend it to the full non-linear theory using Green's function-based methods failed to produce any useful results, leading to the dissolution of Sciama’s Machian programme. In this case the programme failed to carry over assumptions rooted in classical conceptions of spacetime to the general relativistic domain, however questions remain concerning whether the programme may be salvaged in the context of speculative theories of quantum gravity.
Although both arguments claimed to derive from Mach’s principle, we show that they differed wildly in their methodology and the context in which they were developed. More generally, this study illustrates how, despite their eventual failure, philosophical speculations concerning Mach’s principle not only motivated the initial development of general relativity by Einstein, but continued to be influential both by generating interest and influencing the parameters of research well into the renaissance period of the 1960s-70s, at a time when the reputation of cosmology began shifting to that of a precision, empirical science.
Antonio Ferreiro (Utrecht University): Experiments, models, and unification, or why gravitational waves are not ripples of spacetime
In 2016 the LIGO collaboration announced the first direct detection of gravitational waves using kilometer-scale laser interferometers of extraordinary sensitivity. These instruments go far beyond the classical Michelson-Morley experiment: their successful operation depends essentially on quantum theory, both in the modelling of radiation shot noise and in the use of squeezed quantum states to reduce this noise to the quantum limit. Gravitational wave detection is therefore not merely a triumph of general relativity, but an achievement that relies on the nontrivial interaction between general relativity and quantum field theory within a single modelling framework.
In this talk I argue that a careful analysis of how gravitational-wave detection is modelled in LIGO forces a revision of the standard interpretation of gravitational waves as literal “ripples of spacetime.” Instead, within the relevant regime, gravitational waves should be understood as matter-like fields interacting with the interferometer. This conclusion follows not from speculative considerations about quantum gravity, but from the concrete structure required to model the detector-wave interaction.
I begin by reconstructing the canonical model of gravitational wave detection, focusing on the interaction between a weak gravitational perturbation and the quantum optical degrees of freedom of the interferometer. I argue that general relativity and quantum field theory are unified here at the level of models rather than theories. Extending Maudlin’s account of unification [1], I distinguish three forms: trivial unification (e.g. Schrödinger-Newton model or low energy quantum gravity), theoretical unification (as in a full theory of quantum gravity), and dynamical unification. In the latter, modelling the target system requires an epistemic bridge structure between otherwise independent theories. I suggest that LIGO detection exemplifies this third form.
Indeed, the modelling of quantum optical fields in the interferometer requires additional spacetime structure, most notably a time-like Killing field [2], to define positive-frequency modes and a Hilbert space for the quantum field. This requirement effectively places us in the weak-field limit of general relativity, where the spacetime metric is treated as fixed and approximately flat, and the gravitational perturbation appears as a field propagating on that background. In this regime, the gravitational field cannot coherently be interpreted as part of spacetime structure itself, since that structure is already fixed by the background metric underwriting the quantum theory (for an analogue case in atomic system see [3]).
I argue that under these modelling assumptions, and given the surplus structure required for dynamical unification, the gravitational wave functions analogously to matter: it carries energy, couples to quantum fields, and is represented as a dynamical entity distinct from spacetime geometry. This challenges the traditional spacetime-matter dichotomy [4] and supports a reinterpretation of general relativity as a theory that unifies spacetime and matter, rather than merely geometrizing gravity. I conclude by suggesting that this perspective aligns closely with Einstein’s own understanding of general relativity as a unificatory theory of gravity and inertia [5], and shows how general relativity can remain autonomous while being dynamically unified with quantum field theory.
Alexander Franklin (King's College London): Against Rigour: In Favour of Informal Standards for Reduction in Philosophy of Physics
Many would regard rigour as self-evidently a virtue in philosophical theorising. By considering two vignettes from recent debates over reduction in philosophy of physics I'll argue that the pursuit of rigour has distracted philosophers from the questions that motivated their research: in both contexts I'll suggest that less rigorous analysis would better enable the explanation and understanding of the reductive nature of inter-theoretic relations. I'll claim that these case studies exemplify aspects of the broadside against rigour developed in Wilson (2021) as well as themes from Wallace (2022).
The case studies I’ll explore engage with questions of inter-theoretic relations. Reductionism is the thesis that facts about fundamental physics are in some sense sufficient to determine all other facts. In my view debates over reduction and reductionism are enlightening and philosophically fruitful notwithstanding this lack of precise definition. How such debates proceed, at least in many cases, is that some exemplar that's especially difficult to reduce is posited in the form of an explanatory gap. The business of seeking to account for that phenomenon, to explain it, or to demonstrate its consistency with the more fundamental description is the business of reduction. My contention in this talk is that reductive projects would fare better (i.e. would deliver more insights, shed more light, and be more swiftly resolved) were the more rigorous and formal requirements on reduction to be dropped in favour of the much looser goals of closing explanatory gaps and seeking to understand the features of the less fundamental description in more fundamental terms.
The Nagel model of reduction provides a clear exemplar for how reduction might be achieved. It's rigorous insofar as if it purports to offer a logically valid argument for the truth of reduction in each context to which it's applied. My primary claim is that, pace Dizadji-Bahmani et al. (2010) the Nagel model should not serve as a regulative ideal for reduction. I support this claim by examining the case of phase transitions and arguing that the debate has been captured by disagreements over whether Nagelian bridge laws may include limits or not (cf. Nguyen (2024). This has distracted from the more interesting questions, addressed in other parts of the literature concerning how the qualitative discontinuity found in first order phase transitions can be explained from the bottom up.
My second example involves Nickles reduction and the limit-based approach to understanding the relationship between classical and quantum physics. In response to Feintzeig’s (2022) claim that he’s provided such a reduction, I argue that the hbar limit is neither necessary nor sufficient for the instantiation of classical behaviour in quantum physics. I’ll show that decoherence theory is better suited to the quantum-classical reduction: different models of decoherence are applicable in different physical settings, whereas the appeal to limits, while more readily rigorisable, is incompatible with such context sensitivity.
Overall, I claim that the reduction debate is worthwhile and fruitful but ill-served by the rigorous models employed as exemplars in the literature.
Toby Friend (SUNY at Buffalo): Is there a Theory of Everything?
A Theory of Everything (TOE) can be defined (following Friend 2025b) as collection of laws Λ such that,
1. all members of Λ are true,
2. everything that happens in nature is covered by a member of Λ,
3. the predictions of Λ are maximally determinate (incl. determinate probability distributions), and
4. Λ is unified.
That there is a TOE awaiting discovery is a central working assumption within the foundations of physics. But it is a hypothesis which should be under significant scrutiny, not least because of the significant economic costs research into it often gives rise to. Challenges to the existence of a TOE come from multiple sources. Here I will discuss two, and offer responses to them which draw on work in causal discovery and cosmology (respectively).
One challenge concerns whether true and exact physical laws cover all causal relationships. Cartwright (1994) has famously queried this regarding relationships we already know of. But it is problematic also if the future discovery of such relationships remains a live possibility.
In response, I point out that causal relations not covered by true, exact laws are incompatible with the Principle of the Nomological Character of Causality (PNCC). Although it is often unremarked upon, PNCC is a consequence of many philosophical accounts of laws. Indeed, it arguable that Cartwright’s own metaphysics of causation implies the principle (Friend 2025b). Moreover, the justification for PNCC can arguably be given on a priori grounds, since causal relationships cannot be discovered without assuming that causal relationships give rise to predictable type-level probabilistic independencies and dependencies (Markov and Minimality conditions) that we are willing to treat as broadly lawlike.
A second, less often discussed challenge to the existence of a TOE comes from the possibility of idlers: events which have no causal role at all. In contrast with the preceding challenge, many philosophical accounts of laws are consistent with idlers, and hence with laws that are not universal (Friend 2025a). Moreover, there does not appear to be any self-evident principles that prohibits them.
I’ll consider and critique a number of potential extant motivations for rejecting idlers, including an epistemology-constrained ontology and the so-called ‘Eleatic Principle’. I’ll then suggest another motivation, based on the following principle.
* Causal Connectedness of Worldmates (CCW). Two entities x and y (properties, particulars, events) are worldmates if and only if they are causally connected (one is the cause of the other or they have a common cause).
I will argue that there are empirical and philosophical grounds for preferring CCW over the typical spatiotemporal-connectedness criterion for worldmates (Lewis 1986).
1. Inflationary cosmology, particularly the propensity for ‘bubble’ universes, suggests that spatiotemporal continuity might not be a necessary criterion for worldmates.
2. Conversely, Everettian many-worlds theory is not uncontroversially a view about different worlds in the metaphysical sense. So spatiotemporal continuity may not be a sufficient criterion either.
3. AdS/CFT correspondence and various views in quantum gravity suggest that the physical space (if any) in which matter exists is not spatiotemporal.
4. More crucially, there are persuasive arguments from scrutiny of relativity theories that spacetime (and by extension, any physical space) is a ‘non-entity’ (Brown and Pooley 2006).
Antonios Papaioannou (University of Rennes): Predictions in an Uncertain World
While predictions were traditionally regarded as simply an observational consequence derived from a clear theoretical background, modern computational models require revisiting the standard conception. That is because the model outputs that function as predictions emerge from a complex interplay of theoretical assumptions, empirical inputs, and modelling decisions that can contribute to our epistemic ignorance regarding the model, as they each introduce different uncertainties that can compound and couple. In these high-uncertainty computational domains, where models can become highly complex and non-linear, scientists tend to rely more heavily on predictive modelling strategies as a way to justify the success of resulting models and keep the discipline grounded on empirical data, despite the latter also being scarce and uncertain. This, however, can lead to potentially misleading claims about the predictive power of these models, since the modelling uncertainties can mask underlying issues behind the guise of an apparent agreement between model outputs and data. When uncertainties make it impossible to ascertain which model components contribute to a specific output, insisting that such models can deliver physical insights solely based on empirical agreement is unwarranted.
Our aim is to show that, in these contexts, determining when a model can be considered as having predictive power must be reframed into identifying what exactly needs to be known about the model and about how a particular output is produced such that something can be learned from comparing its outputs to empirical data, and a meaningful statement can be made about whether the latter agree or disagree. In order to achieve this, we demonstrate how different types of uncertainties can hinder the interpretation of comparisons between model outputs and empirical data. As an example of a high-uncertainty context, we take the discipline of astrochemistry, utilising the case of Infrared Dark Cloud Models, to showcase some of the problems with ascribing predictive power to highly uncertain models prematurely and how it can obstruct scientific progress by ignoring our lack of understanding. We argue that the problem cannot be solved by suggesting exact criteria for defining what a prediction is in these contexts, but by determining when we have reached a sufficiently good understanding of how an output arises, considering all the different inputs, assumptions, and parameters of the model.
The literature has extensively addressed model validation and confirmation, following the work by Oreskes et al. (1994), including efforts within the computational sciences themselves to develop the necessary tools to meet these challenges e.g. Oberkampf & Roy (2010). Similarly documented are the challenges that arise when simulations become epistemically opaque (Humphreys, 2004) and how that opacity can be linked to understanding (Beisbart, 2021). Despite this rich body of work, the literature has not systematically examined how understanding of a model’s internal workings is necessary for the successful observational agreement to carry epistemic significance. This work intends to fill this gap by articulating the epistemic conditions that must be satisfied for a model output to legitimately function as an insightful prediction.
Dominic Ryder (London School of Economics): No Alternative to String Theory? A Critique of Non-Empirical Confirmation in Quantum Gravity
Contemporary fundamental physics, and in particular quantum gravity research, faces an epistemological crisis. The characteristic energy scale of quantum gravity theories is far greater than those we can test empirically, and there exist practically insurmountable barriers to testing these energy scales in the future. It is therefore unclear how any such theory could receive scientific confirmation.
Richard Dawid has proposed a solution to this epistemological crisis by claiming that scientific theories can be confirmed non- or meta-empirically. A cornerstone of Dawid’s proposal is the no-alternatives argument: if scientists have worked long and hard and failed to develop an alternative to a given theory, then that theory receives some degree of confirmation. Dawid applies this to string theory, a theory of quantum gravity that unifies all known physical interactions. He claims that scientists have failed to develop an alternative to string theory and as such string theory can be meta-empirically confirmed. In this talk I argue that string theory has alternatives and so meta-empirical confirmation cannot be applied to it.
The first aim of this talk is to clarify the “rules” of the no-alternatives argument, which determine what can and cannot count as an alternative. A clear and unified articulation of these rules has not been given in the literature, because Dawid has developed the argument over a number of years and multiple publications. The rules distinguish scientificality conditions from methodological choices: the former are conditions a theory must satisfy to be scientific, and the latter are optional features of any particular theory. I follow Dawid in arguing that alternatives can only be excluded by violating scientificality conditions.
The second aim of this talk is to motivate a family of alternatives to string theory. This family includes any quantum gravity theory coupled to the matter content of the Standard Model of Particle Physics (QG + SM). I provide two examples: loop quantum gravity + SM and asymptotically safe gravity + SM. The asymptotically safe gravity research programme has gained much attention in the physics literature in the past 15 years, and a sub-aim of this talk is to introduce it more widely to the philosophers of science. These quantum gravity theories are briefly described along with their coupling to the SM.
To assess whether the members of this family are alternatives to string theory, I consider five arguments that may exclude these as alternatives according to the rules of the no-alternatives argument: (i) unification, (ii) mathematical consistency, (iii) well-definedness, (iv) effectiveness and (v) UV-completeness. In my discussion of argument (iv) effectiveness, I introduce a novel taxonomy for reasons to consider the SM effective, which is a surprising lacuna in the literature on the epistemology of the SM.
Analysing the arguments (i)-(v), I find that for each, either: A) It is a methodological choice and so cannot exclude alternatives, B) It fails to eliminate QG + SM, or C) It also eliminates string theory as an alternative. Therefore, I conclude that string theory has alternatives and it cannot be confirmed by the no-alternatives argument.
Simone Salzano (LMU Munich): Demystifying Effective Realism
Difficulties in reconciling the structure and aims of effective field theories (EFTs) with debates on scientific realism have recently motivated the development of new philosophical frameworks. One prominent proposal is effective realism (Williams, 2019; Fraser, 2020), which aims to preserve realist commitment while acknowledging the scale dependence of some of our best scientific theories. Although this approach is gaining traction in philosophical discussions on both non-relativistic quantum mechanics (Fraser and Vickers, 2025) and quantum field theory, Ladyman and Lorenzetti (forthcoming) suggest that these existing versions remain metaphysically underarticulated. Particularly, they warn that without clearly specified metaphysical commitments, such positions risk collapsing into antirealism, treating their notion of effective ontology as a merely instrumental tool. In this respect, they claim that effective realism is best vindicated when embedded within an ontic structural realist framework, based on real patterns and a scale-relative ontology (Ladyman and Ross, 2007).
This paper challenges that diagnosis, arguing that effective realism possesses far greater philosophical depth and systematic potential than has been previously recognized. To make my case, I introduce and develop a threefold distinction between weak, medium, and strong forms of effective realism, each associated with different metaphysical commitments.
(i) Weak ER aligns partially with phenomena-first approaches (Massimi 2022). Particularly, it locates realist commitment in stable events and their modal robustness as established through scientific inferential practices, rather than in an ontology of dispositional properties or unobservable entities.
(ii) Medium ER allows for a plurality of modal frameworks and can incorporate the standard version of ontic structural realism, as well as alternative variants that attribute causal powers to structures (Esfeld, 2009).
(iii) Strong ER advances a thicker, layered metaphysics inspired by EFT practice, while remaining cautious about interlevel metaphysical relations at play, in line with recent arguments that grounding does not adequately capture EFTs intertheoretic relations (McKenzie, 2024).
As I will contend, carefully distinguishing these variants advances debates on realism in contemporary physics by clarifying the conceptual space and showing that effective realism need not collapse into structural realism.
Rupert Smith (University of Bristol): Miasma Theory and the Preservation of Modal Structure: A Test Case for Ontic Structural Realism
The case of Miasma theory has been used to challenge ontic structural realism (OSR), but this paper argues that, given an appropriate articulation of the commitment to objective modal structure in OSR, this case of theory change supports OSR. The most sophisticated version of miasma theory was developed in the mid-1800s, and according to the theory, diseases, notably cholera, were caused and transmitted by miasmas, or 'bad airs', that were produced via rotting organic matter. Miasmas do not exist, and the hypothesis that the 19th century cholera outbreak in London was caused by miasmas is false: cholera was revealed to be transmitted not by 'bad air' but by 'bad water'. Nonetheless, Tulodziecki (2021) argues that Miasma theory (specifically this 19th century formulation linking cholera outbreaks to miasmas) provides a case study of a non-mathematical theory which had novel predictive success, but was not approximately true and was eventually abandoned. Hence, Miasma theory has the potential to undermine the connection realists make between predictive success and approximate truth. Tulodziecki claims there is no plausible structural continuity in this case of theory change, challenging structural realism. Tulodziecki thus presents a test case for whether there is continuity of modal structure on theory change for all predictively successful theories, including non-mathematical theories, as defenders of OSR claim. This case study illuminates the extent to which OSR can accommodate, previously overlooked case studies of predictively successful theories that posit entities that are not accepted by current science.
Tulodziecki is not clear about exactly which form of structural realism she is engaging with, nor about whether she regards structural realism as a kind of selective realism (defenders of OSR argue it is not), and she makes no reference to modal structure in her analysis. This paper shows that selective realism cannot account for the continuity on theory change in this case, but OSR understood in terms of modal structure can. Ultimately, while the putative law of miasma theory linking soil elevation to cholera mortality was abandoned, and there is no continuity of reference for the term 'miasma', there is a limited continuity in the representation of objective modal structure, in the form of the counterfactuals, explanations and predictions. Notably, there is a non-accidental correlation between the environments where a modern epidemiologist would say the air is filled with pathogens according to germ theory and where miasma theorists would say there is a high concentration of miasmas. Hence, the case study of miasma bolsters this formulation of OSR against competitors that eschew talk of modal structure, and shows how advocates of OSR can extend their account of representation of modal structure beyond physics.
Yongwoo Yi (LMU Munich): Symmetry Undermines Separability
There are good reasons to think that the global phase of the wave function (or equivalently, global U(1) transformations) is not physically meaningful (Wallace 2022; Gao 2024). This talk explores how this fact should be accommodated by examining two influential approaches to symmetry: reduction and sophistication. As a case study of reduction, I consider geometric quantum mechanics formulated on projective Hilbert space (Ashtekar and Schilling 1997). As a case study of sophistication, I consider an approach that modifies the wave function's value space to treat symmetry-related models as representing the same physical possibility. On this view, the global phase can be a genuine property, but its identity is determined entirely by the network of relative phases.
Although global U(1) symmetry is mathematically simple, I argue that taking it seriously has far-reaching metaphysical consequences. From these case studies, I draw three conclusions. First, some reductionist treatments of global U(1) symmetry suggest a substantially different picture of quantum reality. Second, High-Dimensionalism in quantum mechanics (Albert 1996, 2013; Ney 2021) cannot preserve separability under any plausible treatment of global phase. If correct, this undermines what has often been regarded as the view's strongest motivation. Third, the argument points to a more general moral: whenever a symmetry renders only comparative quantities (e.g. relative phase) physically meaningful, it poses a systematic challenge to separability.
Rebecca R. Cuciniello (University of Genova (FINO)): Plasticity is a Blind Spot in Contemporary Theories of Biological Agency
This paper argues that phenotypic plasticity, as it is understood in recent evolutionary theory, constitutes a theoretical blind spot for current accounts of organismal agency and goal-directedness. I examine the epistemic implications of this limit and suggest a way forward.
In the first part, I start by outlining the theoretical landscape. In the last decades, a view of organisms as goal-directed agents has emerged, with a growing number of researchers arguing that recognising organismal agency is necessary to move beyond the Modern Synthesis and recentre the organism within evolutionary causation (Lala et al., 2019). Within philosophy of biology, two accounts of organismal agency have emerged: the ecological approach (EA; Walsh, 2015), and the theory of biological autonomy (TBA; Moreno & Mossio, 2015). Although they differ in how they define goals and goal-directed activities, they both assume that organismal goals are determinate, describable, and explanatory of behaviour, physiology, and development.
I then show how these theories take plasticity - the organism’s ability to modify its phenotype depending on internal and external context - as a hallmark of organismal agency. Given certain goals, plastic responses express how “the end modifies its means” (Moczek, 2022). Plasticity thus embodies the agent’s robustness, its ability to maintain certain ecological or organisational invariants (goals) through compensatory changes (means).
Following, I contrast this with recent work in evolutionary theory - particularly in evo-devo and the Extended Evolutionary Synthesis - which takes plasticity as a driver of evolutionary change (Pfenning, 2021). From exaptation to phenotypic accommodation, plasticity is better described as repurposing i.e. using the same means for different goals or functions. If plasticity-led evolution shows how organisms change their own goals, then contemporary teleological accounts, which treat goals as invariants and plasticity as robustness, overlook the creative character of organismal agency (Tahar, 2022).
In the second part, I draw out two important epistemic implications of this “mismatch” for theories of goal-directedness such as EA and TBA.
1. The first concerns normative function (and malfunction) attribution. Normative and functional behaviours are generally described as counterfactually dependent on invariant goals across contexts. If no solid such criteria are identifiable, then functional attributions are precarious, context-dependent, and do not fully explain the plasticity of organisms.
2. The second concerns the possible bridge between philosophical and scientific theories. Despite the frequent mutual reference and programmatic support, the two rely on distinct understandings of plasticity (robustness vs creativity), and thus of organismal agency. Particularly, evolutionary works seeking to centre organismal agency within evolutionary causation - i.e. to explain how plasticity, but also niche construction, shape evolution - cannot rely on the concepts offered by EA or TBA.
Concluding, I suggest that teleological conceptions today might be hitting a theoretical limit. The study of plasticity highlights how, if organismal agents’ goals themselves are plastic, it is hard to see how they still count as “goals” within contemporary frameworks. A thorough revision of teleological concepts of agency, goal, and function might be required.
Danny Davis (The University of Leeds): Procedural Objectivity and the Aims of Conservation
Conservation science is a goal-oriented discipline (Soulé 1985), meaning it’s founded upon certain aims or outcomes that it seeks to bring about. These aims are captured by a variety of concepts, such as ‘biodiversity’, ‘ecosystem health’, ‘naturalness’, and others. This paper will begin by arguing that these concepts are thoroughly value-laden, such that their very definition presupposes value-judgements. Claims regarding conservation aims are therefore “mixed claims” (Alexandrova 2018, p.424), the truth of which depends on the evaluative standard used to define the relevant aim.
For example, consider the following claims:
(a) Restoring lost species will produce a more natural ecosystem.
(b) Sheep grazing has led to ecosystem degradation in the uplands of mid-Wales.
(c) Species invasions often increase biodiversity.
Claim (a) depends on whether we judge independence from human purpose and intention to be the value underlying naturalness (Katz 1997), or whether it is instead historical continuity that is important. In the case of (b), value conflicts between rewilding organisations, who claim grazing has caused the Welsh uplands to become “bleak and broken” (Monbiot 2014, p.66), and local communities, who see sheep as integral to these ecosystems, have led to differing verdicts regarding the health of Welsh upland ecosystems. Claim (c) depends on the scale at which biodiversity should be measured. If biomass productivity is prioritised, then a local or 'alpha' measure of biodiversity will be favoured, supporting claim (c), whereas if option and heritage value are prioritised, then ‘beta’ and ‘gamma’ measures will be preferred, leading to the rejection of claim (c) (Lean 2021).
Given this value-dependence, the problem arises of how to decide between competing interpretations of these concepts such that claims regarding the aims of conservation can avoid accusations of relativism and be made more objective. To tackle this problem, I suggest a dual strategy. Firstly, I argue that value-claims should be distinguished from mere subjective preferences, since values, as opposed to mere preferences, are potentially mistaken and subject to rational constraints (Williams 1972, p.17-18). Secondly, I will endorse arguments offered by feminist philosophers against the value-free conception of objectivity (Longino 2004, Harding 2015) and suggest Alexandrova’s “procedural objectivity” (Alexandrova 2019, p.436) offers an account more suited to conservation science, arguing that deliberative procedures can contribute to the objectivity of conservation science by ensuring underlying values are made explicit and subjected to scrutiny from diverse perspectives.
Finally, I will respond to potential objections to the application of procedural objectivity to conservation science. I will argue that the procedural account can successfully address concerns about unequal power dynamics within procedures by ensuring a focus on procedural justice, thereby giving due recognition to diverse knowledge systems rather than assimilating them into mainstream discourse (Reid et al. 2020, Ruano-Chamorro et al. 2021). I will also respond to the charge of anthropocentrism, arguing that so long as the non-anthropocentric values of stakeholders are given weight, procedures will not necessarily be anthropocentric, as well as suggesting the possibility of assigning representatives for non-human populations within procedures.
Zaza Doborjginidze (Hong Kong University of Science and Technology): The Organizational Account of Biological Functions is a Dispositional Account of Biological Functions
Two main theoretical traditions have dominated recent philosophical debate on biological functions. On the one hand, historical or etiological accounts (Wright 1973; Millikan 1989; Neander 1991; Godfrey-Smith 1994) define the existence of function diachronically, in terms of its selection history. This approach explains the normative dimension of function - distinguishing what a trait should do from what it merely does - but at the cost of disconnecting function from a trait’s current causal capacities, a problem known as epiphenomenalism (Christensen Bickhard 2002). On the other hand, dispositional accounts (Cummins 1975; Craver 2001; Davies 2001) define functions synchronically, as the causal contributions that a component makes to the capacities of a containing higher-level system. While this preserves the link between function and its current causal role, it struggles to provide a non-arbitrary, normative basis for distinguishing proper functions from mere accidental beneficial effects, known as the normativity problem. In recent years, the organizational account (OA) has emerged as a third contender, promising to merge the two traditions (Mossio et al. 2009). It proposes that genuine functions arise only within self-maintaining differentiated systems characterized by organizational closure.
OA defines organizational closure as a network of mutually dependent constraints - a set of components in which each part’s activity helps maintain the conditions for the existence of every other part, including itself and the whole system. By grounding function in the intrinsic, self-producing teleology of the living organism, OA aims to capture both normativity (a trait ought to contribute to the system’s closure) and present relevance (its function is its current role in this selfsustaining whole). Artiga and Mart´ınez (2016), however, argue that OA’s attempts to explain cross-generational traits, based on the notion of “encompassing systems”, ultimately reduce it to an etiological theory. In explaining crossgenerational traits OA must refer to past lineage contributions, exposing it to the same epiphenomenalism that plagues historical accounts. If correct, OA would merely be a variant of selected effects theory, not an alternative. Mossio and Saborido (2016) respond to the criticism by claiming that in OA, closure refers to an atemporal, relational structure of mutual dependence rather than a historical lineage. As a result, a crossgenerational trait’s function is defined by its role in the ongoing self-maintenance of an organization spanning generations. This relational interpretation of OA, I argue, covertly changes OA’s ontological assumptions from intrinsic, ontologically objective teleology to an arbitrarily defined contribution to higher-level systems, thereby transforming it into a version of a dispositional account. From here, I argue, Organizational Account has two options, either accept splitting of concept of function into evolutionary and physiological functions, and apply their theory to the latter, or to accept that OA is not seperate category, but rather a version of dispositional account. Organizational Account has two options, either accept splitting of concept of function into evolutionary and physiological functions, and apply their theory to the latter, or to accept that OA is not seperate category, but rather a version of dispositional account.
Rebecca C. Mann (The University of Sydney): Metabolic Wholes: The Organism as Having a Centred Metabolic Network
The concept ‘organism’ is central to the biological sciences and is often used to describe a very particular sort of biological entity, one that has a specific level of physiological organisation. The organism concept grounds many other often-discussed topics in biology, including population dynamics, social interactions, and ecological systems. Despite its operational importance, there remains a lack of consensus regarding how we should characterise the organism. There are four levels at which we can identify this disagreement. The first is whether the concept of the organism and the evolutionary individual are the same or distinct. In this paper, I adopt a view that treats evolutionary individuality and organismality as representing distinct concepts, picking out different sorts of entities in biology (see also Godfrey-Smith 2013).
The second level of disagreement concerns whether we should be considering criteria of organismality abstractly or materially—whether the material realisation of notions such as cooperation, integration, or organisation matters when defining the organism concept (see also Baedke 2025; DiFrisco 2019).
The third level concerns which specific criteria should underlie the concept “organism”. Is the organism a spatially bounded whole, an immunological whole, or a metabolic whole (Cf. Gould and Lloyd 1999; Pradeu 2010; O’Malley 2020)? Here, I argue for a metabolic account of the organism, basing this account in a metabolic view of life. There are multiple metabolic accounts of the organism describing the organism as a metabolically integrated whole that persists over time (see Dupré and O’Malley 2009; Godfrey-Smith 2013; Skillings 2016). However, there has been limited work done to spell out exactly what this kind of metabolic integration looks like in real entities, particularly in comparison to other metabolic interactions, such as those found in ecosystems.
This raises a fourth level of inquiry: in real organisms, what are the precise material features of metabolism that make it organismal metabolism?
In this paper, I argue for a materially restricted concept of the organism, grounded by particular biological functions, substances, and/or processes, instead of more abstract concepts that rely solely on generalised notions like cooperation or integration. Organisms are the quintessential living entity and to be alive is to metabolise. As such, we should turn to accounts of metabolism to define the organism. However, the presence of metabolism alone is not enough to describe the organism.
In this paper, I develop an account of the organism as an entity that has a centred metabolic network. I identify two main features indicative of a centred metabolic network: (1) metabolism centred around the production of core intermediary metabolites (Friedlander et al. 2015; Itoh et al. 2024); and (2) the compartmentalisation of metabolism, with the sharing of intermediary metabolites between compartments (Bar-Peled and Kory 2022).
An organism could thus be defined as a cooperative collection of biological parts that are unified by a centred metabolic network, self-maintaining and resisting entropy by turning energy from the environment into usable energy for the whole organism.
Patrick McGivern (University of Wollongong): Transitions in biological individuality and mesoscale structure
In this paper, I examine the extent to which questions concerning transitions in biological individuality can be clarified (and generalised) by comparing them with questions concerning other types of transition in natural systems, in particular, transitions associated with the emergence of mesoscale structures in multiscale systems in physics (systems exhibiting structures on a variety of intermediate scales between the atomic and the macroscopic). Evolutionary transitions in biological individuality involve “the evolution of a higher level biological unit out of formerly-free living units” (Okasha 2022). Ideas about transitions in this sense are central to discussions of major evolutionary transitions more generally (Maynard Smith and Szathmary 1995; Clarke 2014, 2025). While questions concerning transitions in individuality may appear to be uniquely biological, many of the underlying concerns regarding the enduring status of collectives and the emergence of higher-level individuals can also be found in other settings, including more general biological ones (such as questions concerning the emergence of coordinated behaviour in animal collectives) and in non-biological ones (such as questions concerning hierarchical structures in materials science). For example, Bourrat has recently proposed a ‘coarse-graining’ account of biological individuality intended to capture a ‘quasi-ontological’ sense in which higher-level individuals emerge from collections lower-level ones: this occurs when those higher-level individuals are in some sense coarse-grained summaries of lower-level evolutionary processes (Bourrat 2023). This suggests an analogy with questions concerning hierarchical structures in materials science, where the relationship between structures at intermediate mesoscales can be understood in terms of a variety of coarse-graining procedures (Batterman 2021). Importantly, these mesoscale structures invite questions that are very similar to those associated with biological individuality, for instance concerning the ontological status and explanatory significance of such structures, and concerning the extent to which transitions between structures fall into natural kinds that can be described in terms of broad generalities. I argue that considering these questions concerning transitions between mesoscale structures provides a useful way of framing corresponding questions concerning transitions in biological individuality.
Francisco Navarro (Durham University): Ecological units as moral patients: challenges to ecological health normativism
Attributing health to ecological units (e.g., ecosystems, communities) is often assumed to support environmental ethics. The commitment behind this assumption is that health is a normative or partly normative concept: not merely a description of states, but also a value-laden concept in which those states are judged as good (advantageous, beneficial) for an entity for its own sake. Protecting what is good for an entity is a basic ethical intuition, and thus the ascription of health seems to provide moral standing.
This interpretation, which I call ecological health normativism (EHN), is shared by most theorists of ecosystem health (e.g., Costanza et al. 1992), land health (Leopold 1949; Callicott 1987; Millstein 2024), and some advocates of movements such as One Welfare (Webster 2022). In its stronger version (McShane 2004), health is considered a form of well-being (i.e., health is good for an entity for its own sake) that can be literally ascribed to ecological units. If such units can be literally harmed or benefited, moral agents would have duties to protect them for their own sake. In short, to be a bearer of health is to be a bearer of well-being deserving moral considerability and thus to qualify as a moral patient (Johnson 1992). In its weaker version (Dussault 2018), health is distinguished from well-being and treated naturalistically as an entity’s normal functioning (cf. Boorse 1977), yet the intuition that health can support the moral considerability of ecosystems and other biological entities is preserved.
I argue that EHN faces two problems that undermine its viability as a normative foundation for environmental ethics and the consideration of ecological units as moral patients. First, if health is interpreted as well-being, EHN advocates face the substitution problem (Inkpen 2024): even if health is attributed to ecological wholes, in situations of decision-making, such as healthcare and conservation policies, the units we ultimately care about (i.e., the entities to which well-being is indexed) are only some parts of those wholes (e.g., the human host in a holobiont, particular living beings within an ecosystem). In short, bearers of well-being are entities we are not willing to substitute in cases of trade-offs, and ecological wholes (qua wholes) don’t meet this condition. Second, if health is interpreted naturalistically, what arises is the naturalistic fallacy or is–ought gap (Newman et al. 2017, 275–79). This occurs because ecological health is typically defined in terms of ecological performance (e.g., functions, self-renewal), but normative import, especially the generation of moral obligations, cannot be derived from ecological properties alone.
As an alternative, I sketch a conception of environmental health as a metaphor defined through a process of critical deliberation (cf. Longino 1990). On this view, environmental health is the collection of states conducive to the conservation of ecological wholes, considered units of care by a community of competent stakeholders, including scientists and experts in local knowledge. This preserves the practical utility of health-talk for conservation while avoiding the problems faced by EHN.
Celso Antonio Neto (University of Exeter): Typological Problems in Populational Science
In recent decades, philosophers have re-examined the nature and legitimacy of typological thinking. This examination focuses largely on evolutionary-developmental biology (Lewens, 2009, Witteveen, 2018). Less philosophical attention is given to typological thinking in human population genetics, yet scientists in this field often worry about it (National Academies of Science, Engineering, and Medicine, 2023). Here I evaluate the persistence of typological thinking in that field from a history and philosophy of science perspective, arguing that (i) some problems of traditional typological thinking are present in human population genetics; (ii) these problems have now a practical nature in terms of decision-making and risk management.
There is no single agreed-upon definition of “typological thinking” (Witteveen 2018). Human geneticists use this notion vaguely to capture the view of organisms as naturally divided into homogeneous groups with clear-cut boundaries (National Academies of Science, Engineering, and Medicine, 2023). For them, typological thinking systematically misunderstands how phenotypic and genotypic diversity is distributed, and this misunderstanding has pernicious consequences (e.g., reinforces racial stereotyping). In the first part of the talk, I show that the key critics of traditional typological thinking had similar scientific and ethical concerns. In different ways, Dobzhansky, Simpson, and Mayr emphasized how typological thinking overlooks variability within groups, exacerbates variability between groups, and neglects the overall “clinal” patterns of diversity. Yet, this criticism was heavily focused on matters of definition, theoretical, and metaphysical assumptions.
The statistical, model-based approaches of human population genomics can still reproduce some problems of traditional typological thinking. In the second part of the talk, I use case studies to show how those problems are now based on the practical decisions of geneticists at different research stages. For instance, genetic studies about the origins of Roma populations rely on data collection and “cleaning” strategies that reduce those populations to a genetically homogeneous sample group with a fixed genotypic profile (i.e., “statistical type”), leading to inferences that risk oversimplifying the diversity and stereotyping Roma people (Lipphardt et al. 2021). Scientists might not explicitly assume that populations are genetically homogeneous and have a fixed statistical type, but their decisions construct populations in this way. This analysis suggests that avoiding typological thinking is harder than recent scientific recommendations suggest (National Academies of Science, Engineering, and Medicine 2023).
Following Alan Love’s analysis (2009), one might claim that there is nothing intrinsically wrong with using oversimplified statistical types if scientists understand them as idealizations and communicate them as such. It might be possible to rehabilitate typological thinking in human population genetics (just like in evo-devo). However, in practice, this rehabilitation is doubtful. Scientific choices have epistemic and ethical risks (Harvard and Winsberg 2022). In human population genetics, these risks are particularly high, as statistical types have become predictive tools (e.g., in clinical settings). Strategies to overcome those risks (e.g., science communication) have mixed results, and scientists must weigh such risks throughout decision-making. Thus, in human population genetics, I conclude that problems of traditional typological thinking have morphed into decision problems and risk management problems.
Yoshinari Yoshida (University of Copenhagen): Transferring cells from zoos to laboratories: The issue of accessibility in “cellular anthropology” and its epistemic consequences
In an emerging area called “cellular anthropology,” pluripotent stem cells are artificially generated from different primate species, in particular apes (Prescott et al. 2015). Such induced pluripotent stem cells (iPSCs) of primates are used as models for investigating species-specific patterns of embryonic development, and compared with each other in order to acquire insights into human evolution. How well do primate iPSCs serve as model systems in such research? While there are various criteria for evaluating experimental organisms and model systems (Dietrich et al 2020), this presentation focuses on the criterion of accessibility. By accessibility, I mean the ease of supply of the biological system in question. For example, intensive experimental research on a specific organism relies on stable supply of the organism. This is often dependent on specific institutional settings, such as stock centers and community ethos that prioritizes collaboration. Such institutional settings have played key roles for establishment and maintenance of model organisms and had epistemic consequences (Ankeny and Leonelli 2020).
To characterize the issue of accessibility concerning primate iPSCs, this presentation discusses the results of qualitative research conducted in Japan and Germany. In vitro technologies allow stem cell researchers to bypass ethical and practical challenges concerning access to experimental organisms. However, the uniquely interdisciplinary status of primate iPSCs leads to a distinct set of challenges concerning accessibility. To explain this, I emphasize that primate iPSCs are both specimens of exotic animals and living cells capable of indefinite proliferation. As specimens of exotic animals, primate cells are often acquired from zoos. This occurs typically when a captive primate dies, which brings significant contingency and unpredictability to primate iPSC research. It also means that the moments that provide scientists with opportunities to obtain samples coincide with moments of tragedy for the zoo. Thus, stem cell researchers have to invest time to establish and maintain good relationships with zookeepers to secure their cooperation. An information platform originally aimed at promoting primate conservation plays a role in establishing such relationships in the case in Japan. As living cells capable of indefinite proliferation, primate iPSCs raise issues concerning responsible use and distribution. I describe a case in which stem cell researchers themselves arrange a special material transfer agreement that imposes unusual constraints on distribution of primate iPSCs in order to secure long-term accessibility.
These features have epistemic consequences. The contingency and unpredictability involved in acquisition of primate cells influence what questions scientists pursue by using primate iPSCs. The necessity of maintaining good relationships with zookeepers also affects orientations and scales of research. The limitations on use and distribution of primate iPSCs put constraints on specific types of collaborations while favoring others. I combine these observations and discuss how accessibility conditions shape knowledge production.
Ali Boyle (London School of Economics and Political Science): Animals, mental time travel and the harm of death
Does it harm an animal to die? Someone looking to animal welfare science for an answer to this question would conclude that it does not. It’s widely agreed within the field that death itself is not a welfare harm: welfare assessment frameworks do not include death as an instance of harm (Jensen, 2017), and core animal welfare textbooks do not discuss the possibility that death could be harmful. Killing an animal in the prime of life is considered humane, provided that it is painless. In pest control, lethally shooting an animal has been recommended on welfare grounds over scaring them with loud gas guns (Baker et al. 2016). In this talk, we identify three possible reasons why the field has adopted this position: i) a preference for concepts of welfare that justify the use of animals for human ends (Haynes, 2008); ii) a mistaken inference from the fact that there can be no welfare problem once an animal is dead (Broom, 2011), to the claim that becoming dead cannot itself constitute a welfare problem; and iii) the belief that non-human animals lack the cognitive capacities required for death to harm them (e.g. British Veterinary Association, 2016). We focus on the third explanation, since arguments of this kind have also been developed in philosophy.
The cognitive capacity most frequently singled in such arguments is mental time travel - the capacity to recall experiences from one’s past and imagine experiences one might have in the future. Mental time travel looms large in the life of the average human, but – according to these arguments – animals lack mental time travel, or have it in a comparatively impoverished form. Because of this, several philosophers have argued that death is harmful to us but not (or comparatively less so) to animals (Belshaw, 2015; McMahan, 2015; Nussbaum, 2023; Singer, 2011; Velleman, 1991). We show that such arguments are problematic in two ways. First, they rely on controversial and theoretically loaded claims about animal cognition, which do not survive contact with the empirical literature. The evidence about mental time travel in animals is subject to significant debate and uncertainty, to the extent that mental time travel cannot confidently be ruled in or out for any nonhuman animal – so, confident pronouncements the mental time travel capacities of any animal are misplaced. Whilst Singer and Nussbaum acknowledge some uncertainty and recommend precautionary reasoning – attributing mental time travel to animals whose possession of it is uncertain – we show that the application of precautionary reasoning to this question is vexed by the existence of interspecific variation (Boyle, 2022, 2024; Boyle & Brown, 2025; Schwartz & Boyle, 2025). Second, they anthropofabulate (Buckner, 2013) by exaggerating the role of mental time travel in grounding the harm of death for humans. When we adopt a more expansive account of the cognitive capacities required for death to be harmful, and the behaviours that might evidence such capacities, we find reasons for taking seriously the idea that many animals are harmed by a premature death.
Simon Brown (Ashoka University): Kinds of Semantic Memory Across Kinds of Mind
What kinds of long-term memory do different species have? Since Tulving (1983), most scholars have seen the interesting question here as whether episodic memory—seen as special, requiring dedicated neural machinery and a unique phenomenology—is unique to humans. Comparative psychologists have focused on experimentally demonstrating behaviours which could only be explained through episodic memory, ruling out explanations that appeal to ‘mere’ semantic representation or associations (For a review, see Boyle & Brown, 2025).
Less attention has been paid to what semantic memory might look like in non-human animals. This is an important gap in its own right: semantic memory is crucial to cognition. Species with different forms of semantic memory may have very different forms of mind generally as a result. Indeed, many forms of semantic memory are every bit as demanding as episodic memory, raising the possibilities that they arose after episodic memory (Healy et al., 2024), that simple forms of episodic and semantic memory co-evolved, and that semantic memory and episodic memory are both unique to humans (Murray et al., 2017). Furthermore, ignoring semantic memory distorts the debate about episodic memory itself, given how closely episodic and semantic memory are entangled (Addis & Szpunar, 2024; Aronowitz, 2022; Boyle & Brown, 2025).
It is therefore vital to consider the distribution of semantic memory, and the forms it might take in different species. However, progress on such questions is hampered by the construct ‘semantic memory’ itself, which, despite recent efforts at clarity (e.g. Addis & Szpunar, 2024), runs together very different phenomena with little in common other than not being episodic memory. I distinguish seven such phenomena:
(1) sentence-like representational format
(2) conceptual content (availability to rational inference)
(3) generalising over multiple experiences
(4) abstraction
(5) serving linguistic communication (e.g. storing information about word meanings)
(6) information organised in a structure shaped by language
(7) noetic phenomenology
Although there are important connections between some of the phenomena, they are unlikely to hang together as a natural kind. Probably, they are distributed very differently across the animal kingdom. For example, some are associated with linguistic communication, but many non-linguistic animals likely have abstraction and generalisation.
In distinguishing these phenomena, I draw on a rich literature on whether animals could have concepts and beliefs (e.g. Camp, 2009; Peacocke, 1992; Quilty-Dunn et al., 2023). Framing these traditional issues in terms of memory also brings additional insights to the traditional literature, such as focusing attention on issues of encoding, storage and consolidation, memory organisation, and retrieval and reconstruction.
Joshua Kramer (University of Pittsburgh): Does everyone know what addiction is?
The philosophy and psychology of addiction are at an impasse. For decades, researchers have focused on solving the “puzzle” of addiction: why do addicts use drugs despite negative consequences? Two solutions are common, one appealing to compulsion and another to choice; neither have proven satisfactory. But even if we find a satisfactory solution, the puzzle ignores important features of addiction: (i) relapse after extended periods of abstinence, (ii) resumed drug use on “special occasions,” (iii) curiosity as a risk factor for addiction, and (iv) the comorbidity of addiction and attention disorders, especially ADHD. These features elude compulsion- or choice-based explanations, and also suggest a central role for attention in understanding addiction. Might addiction be a problem of attention?
No one denies that attentional processes are part of addiction’s story. But we posit that addiction is a problem of attention in a stronger, more fundamental sense. We argue that compulsion- and choice-based explanations depend on more fundamental commitments regarding attention. For example, advocates of the incentive-salience account, a compulsion-centered view, refer to external cues as “capturing” the attention of the addict. Alternatively, advocates of the choice-based explanation focus on addicts’ attending to their person-level values or self-conceptions in a reflective, deliberative way. These types of attention, which we call environmental and reflective attention, respectively, are certainly part of the story. But they cannot explain the important, yet overlooked, features of addiction mentioned above (e.g., high curiosity as a risk factor).
We develop a novel account of an additional mode of attention: the organizational mode. We identify rigidity and flexibility as key dimensions along which organizational attending can vary. A flexible organizational mode structures and orients attending with an agility that is sensitive to the variably textured internal and external environment, and the agent’s shifting practical and theoretical interests. Alternatively, a rigid organizational mode of attending assumes the same structuring and orientation of attention, regardless of changes in the agent’s interests.
We argue that organizational attention is central for understanding addiction. Addiction is often described as a problem of rigidity over and above its given object; addicts often experience a general feeling of being “enslaved” by a system of routines that holistically anchor their life and that cannot be explained by a more restricted, object-focused compulsion. Our picture explains this rigidity: addicts too narrowly structure and orient attention in a way that non-addicts would acknowledge swings free from their variably textured internal and environmental interests—e.g., lacking more beneficial stress coping-mechanisms. A rigid organizational mode offers further and novel insights into the overlooked features of addiction. For example, empirical evidence suggests that “absorption,” an element of curiosity characterized by a tendency toward intense engagement, is also a risk factor for substance abuse. We argue that absorption-curiosity is a manifestation of the rigid organizational mode of attention in addiction. Finally, we draw further connections between rigidity and the attentional difficulties characteristic of ADHD, laying the foundation for an account of addiction that can make sense of its currently unexplained comorbidity with ADHD.
Rebeca Mendez (Instituto de Investigaciones Químicas, Biológicas, Biomédicas y Biofísicas, Universidad Mariano Gálvez de Guatemala): The Gut-Microbiome-Brain Heterarchy: Distributed Regulation and Explanatory Topology
Philosophical and scientific discussions of the gut-microbiome-brain axis (GMBA) often default to a hierarchical picture of control: the brain is treated as an executive regulator that modulates peripheral systems, while microbial and immune influences appear as upstream inputs or downstream effects. This paper argues that this hierarchy-first framing is not merely a stylistic metaphor but an explanatory commitment that misdescribes the causal organization of the GMBA and thereby distorts both mechanistic explanation and intervention reasoning.
I develop a topological thesis: the GMBA is more accurately modeled as a heterarchy, -an organization in which multiple regulatory subsystems (neural, endocrine, immune, metabolic, and epithelial) exert multidirectional and state-dependent influence without a fixed privileged controller. The argument proceeds in three steps. First, I characterize hierarchy and heterarchy as competing claims about control architecture. On the hierarchical picture, (i) there is a stable ordering of levels or modules, (ii) causal influence is primarily routed through designated relay points, and (iii) explanatory credit assignment tends to privilege one locus (often cortical or central). On a heterarchical view, (i) pathways can bypass intermediate relays, (ii) lateral coupling within levels is common, and (iii) which subsystem dominates is contingent on physiological state and timescale.
Second, I show how contemporary GMBA evidence supports the heterarchical architecture claim. The relevant features are not mere bidirectionality but the presence of partially independent pathways whose relative gains shift across conditions (e.g., stress, infection, circadian phase, diet, antibiotics). In acute stress, autonomic and endocrine loops can temporarily dominate; in infection, immune-to-neural signaling can reorganize behavior; in chronic dietary change, metabolic and epithelial dynamics can reweight downstream neural sensitivity. These patterns are hard to reconcile with a single enduring executive without adding ad hoc switches that effectively reintroduce heterarchy under another name.
Third, I outline the benefits out the implications for philosophy of science, specifically regarding integrative pluralism and mechanistic explanation. I argue that a heterarchical model clarifies how explanations in complex biological systems can remain non-reductive without lapsing into vague holism. It provides a framework for tracking how control shifts across interacting mechanisms, -such as microbial resource sensing or neurovisceral setpoints-, rather than searching for a uniquely fundamental executive level.
By showing that none of these layers acts as a permanent executive, this model addresses the credit assignment problem in multi-level systems: affective dispositions and regulatory decisions emerge as collective achievements of coordination rather than outputs of a single controller. Finally, I offer methodological guidance for causal inference. Because host interpretive gain and thresholds shift based on the system's state, the same intervention (e.g., a specific metabolite) may yield different outcomes across different regimes. This positions heterarchy not just as a topological description, but as a precise, testable alternative to hierarchy that accounts for the regulatory oscillations and noise often seen in dysbiosis.
Scott Partington (University of Cambridge): From Bad to Wrong: The Origins of Deontic Cognition
Many animals can represent actions as having a good or bad value, but only humans can represent them as right or wrong—as things that ought or ought not to be done. This distinction between evaluative cognition and deontic cognition raises an evolutionary puzzle (Joyce, 2006; Stanford, 2018):
The functionalist’s puzzle: If evaluative cognition already guided adaptive behavior, why did natural selection build a deontic system in our lineage, too?
Prominent theories hold that deontic cognition evolved to promote cooperation, either enabling novel forms (Bowles & Gintis, 2011; Henrich, 2015) or making existing forms more reliable (Dennett, 1995; Joyce, 2006; Kitcher, 2011). However, evaluative mechanisms can achieve such outcomes by marking cooperation as very good and going solo as very bad (Stanford, 2018). As Tomasello (2016) and Sterelny (2021) argue, such “opportunistic” mechanisms likely underpinned early forms of hominin cooperation.
By contrast, I develop an efficiency account: deontic cognition evolved because it guided adaptive behavior with greater cognitive efficiency. This account is based on the principle of cognitive efficiency (Icard, 2026): an efficient cognitive system maximizes reward, subject to constraints on resource cost (e.g., memory, time, etc.).
Consider memory costs: an evaluative system requires storing a value for each behavioral option (Spurrett, 2019). Think of labeling each option with a Post-It note: red for bad values, green for good values, with darker shades for larger magnitudes. By selecting the option with the better value, this system will always make an adaptive choice. But this system is expensive: each new distinction in payoffs has to be marked by a distinction in stored values. In general, encoding a range of values [-R, R] costs ⌈log2(2R + 1)⌉ bits per entry.
Rather than graded values, deontic systems guide behavior through a fixed collection of deontic statuses (e.g., [OUGHT, OUGHT-NOT]). If a deontic system contains K statuses, then each status assignment costs ⌈log2(K)⌉ bits per entry; for example, it costs 1 bit to encode OUGHT (= 1) vs. OUGHT-NOT ( = 0). On the margin, then, storing a new deontic status is cheaper than storing a new value representation.
Hence deontic cognition creates substantial savings whenever a precise, value-based ranking is unnecessary for adaptive decision-making. For collaborative foragers like our hominin ancestors, going solo is always worse than cooperating. Instead of storing the precise bad value of going solo, then, hominins just needed to “tag” that option with a cheap, OUGHT-NOT status. The same resource savings apply in non-cooperative contexts, too. Instead of storing a precise bad value for “eating poisonous berries,” “approaching predators,” etc. hominins could store an OUGHT-NOT status instead.
To be sure, early forms of hominin cooperation were likely supported by evaluative cognition, as Tomasello and Sterelny argue. From there, though, individuals who used deontic cognition could re-allocate surplus resources for other fitness-enhancing tasks (e.g., creating adaptive culture). This gives us a new solution to the functionalist’s puzzle: deontic cognition was favored by natural selection because it is cognitively efficient, not because it is uniquely capable of promoting cooperation, per se.
Charlotte Skye (University of Birmingham): Why the Evidence-Resistance Criterion Fails as a Scientific Marker of Delusion
Diagnostic manuals in psychiatry characterise delusions partly by reference to evidence resistance, understood as the persistence of a belief despite counterevidence. This criterion is intended to function as a demarcation tool, distinguishing pathological belief from everyday epistemic failure. However, no major diagnostic manual specifies what constitutes evidence or what it is to resist it. I argue that the evidence-resistance criterion depends on an implicit and institutionally specific conception of evidence, and that once this conception is made explicit, the criterion fails as a viable scientific marker of epistemic pathology.
I begin by reconstructing the intended diagnostic role of the evidence-resistance criterion: to identify a failure of belief revision sufficient to warrant medical classification. It therefore presupposes a benchmark of appropriate evidential reasoning. I argue that clinical diagnostic practice relies on a tacit clinical conception of evidence, according to which evidential authority is primarily attached to publicly observable, reproducible, and intersubjectively corroborated materials, while first-person experiential reports are epistemically downgraded unless supported by such materials. However, its legitimacy within clinical science does not establish its suitability as a universal standard for identifying epistemic failure.
The core of the paper is a worked case illustrating how belief persistence can arise from divergence in evidential frameworks rather than resistance to evidence. The case shows how two agents, confronted with the same perceptual stimulus, may form and maintain incompatible beliefs while remaining responsive to what each takes to be evidentially relevant. The divergence is best explained by differences in evidential bases, understood as structured commitments governing what sources are admissible as evidence, rather than resistance. From within one evidential framework, the other agent’s belief appears resistant; from within the other, it appears coherent. This demonstrates that judgements of evidence resistance are framework-relative and that persistence alone cannot ground a diagnosis of epistemic failure.
I then consider whether this difficulty can be resolved by appealing to evidential pluralism. While pluralism about evidence is theoretically plausible, it does not render the evidence-resistance criterion operational. Evidential frameworks are tacit, dynamically modulated by context and experience, and not transparently accessible to either the subject or external assessors. Consequently, clinicians lack principled access to the evidential bases relative to which a belief is formed and maintained. Without such access, it is not possible to determine whether a belief persists through resistance to evidence as understood by the believer, or through responsiveness to evidential considerations that diverge from institutional standards. The criterion therefore presupposes epistemic access that clinical practice cannot achieve.
The conclusion is methodological. A scientific criterion must be epistemically accessible, operationalisable, and reliably applicable across cases. The evidence-resistance criterion fails to meet these conditions, not because epistemic resistance never occurs, but because institutional evidential standards are extended beyond their domain of legitimate application. When this occurs, ordinary epistemic divergence is systematically misclassified as epistemic failure. I thus offer a case study in how institutionally grounded evidential standards can over-generate pathology claims when treated as universally authoritative.
Jidong Wang (Fudan University): Putting the Mind Back into Signaling Games: Modeling Unicepts within the Sender-Receiver Framework
In discussions on the origins of language and meaning, information-theoretic approaches within the sender-receiver framework are often regarded as promising candidates for a naturalised theory of meaning. However, they face seemingly inevitable inherited hurdles. Artiga highlights three:
1) Misrepresentation: content of signal varies immediately with changes in correlating states of world or actions, yielding either a purely descriptive account without normativity (and thus no room for error) or one that relies on additional, often undefended, normative assumptions;
2) The partition problem: signal content is highly model-dependent, with the modeller's prior choice of state partitions largely determining what signals mean, and no principled criterion for preferring one partition to another;
3) Contradictions among competing theories: informational content (Skyrms, Barrett), functional content (Shea, Godfrey-Smith, Cao), and related accounts each capture important aspects of meaning but lack a unifying framework. As a result, sender-receiver models are widely treated, even by sympathetic authors such as Hurford, Sterelny, and Planer, as useful heuristic tools rather than as powerful formal methods for testing and theorising about human language evolution. The core problem is that they focus primarily on communication while neglecting the internal cognitive architecture of agents.
Recent work by Wang and Zhang (2025) addresses the partition problem by introducing a cost-based perception model. In their reinforcement learning model, agents spontaneously learn to partition the states of the world, given state-features, payoff matrices, and observational costs. While this begins to unpack the cognitive black box, I argue that their account remains semantically insufficient. First, Wang and Zhang do not provide a substantive theory of meaning. They do not explain what constitutes representational content or how normative standards of correctness arise. Second, their architecture relies on internal signals that function redundantly alongside learned perceptual states.
To address these issues, I propose integrating Wang and Zhang's model with Millikan's unicept theory as a unified foundation for a naturalistic account of meaning. I develop a Millikan-inspired account on which unicepts can be modeled as structured states that integrate informational, functional, and distributional dimensions. I reconceive their internal signals as products of same-tracking mechanism, and, drawing on predictive processing, reformulate their model as a sender-receiver dialogue structure in which agents use memory, predictive-error feedback and payoffs from actions to learn whether sensory inputs constitute one object or several. I outline a toy model that implements this reinterpretation, illustrating how unicept-like internal states can be learned and stabilised. This yields a three-layered architecture comprising individual cognition (unicepts), inter-individual communication, and population-level conventions, which helps to explain how systematic misrepresentation can arise. Moreover, with Wang and Zhang's solution to the partition problem, together with a unicept-based integration of various theories on content, the resulting account offers a more direct and unified response to the three major challenges posed by Artiga. While much work remains to be done, this approach offers a promising foundation for a naturalised theory of meaning that bridges animal cognition and human language.
Yinzhu Yang (University of Cambridge): Analog and Digital Representation Reconsidered: A Network-Based Approach to Cognition-Perception Border
Where, if anywhere, is the border between perception and cognition? A family of answers appeals to representational format: perceptual states are said to be iconic, while cognitive states are discursive or language-like (Block, 2023). Yet the contemporary landscape is crowded with borderline phenomena that seem to possess both perceptual and cognitive markers (Shea, 2024). This paper undermines the expectation that a single, rigid format boundary can do the explanatory work demanded of the perception–cognition distinction. I propose that the abovementioned problems could be overcome if we view representations in the lens of network analysis.
I begin with two pressure points. In billiard-ball causation, we have experiences that display perceptual signatures like immediacy, adaptation, and known-illusion persistence, while also trainable and responsive to background expectations (Danks and Dinh, 2022). Core cognition cases in general raise similar worries: the same system can appear perceptual in its automaticity and cue-dependence, yet cognitive in its integration with reasoning and learning.
In response to these problems, I suggest we need to go back and reexamine the concept of representational format, which I believe is the culprit to blame for our assumption of a hard and fixed boundary between cognition and perception. I argue that these disputes are partly sustained by a tacit Rigid Format Assumption: analog and digital formats are mutually exclusive, sharply bounded natural kinds. Against this, I motivate a Soft Format Assumption by emphasizing hybrid analog–digital systems that require interface components for converting and coordinating representational vehicles (Maley, 2023). Once hybrid interfaces are taken seriously, it becomes less plausible that ‘analog’ and ‘digital’ neatly partition whole psychological systems such as perception and cognition.
My central proposal is that representational format should be analyzed one explanatory level down, as properties of network substructures. The concept of network plays an important role in my formulation of analog and digital representation. I argue that the distinction between formats can be flexible because their underlying network is flexible. I propose that analog and digital representations are better understood in terms of properties of network science, instead of immutable, fixed kinds. This characterization has many benefits: In particular, it would help us form a novel understanding of many problems in cognitive science, including the debates regarding cognition-perception border and cognitive ontology.
According to the network framework I describe, analog representations and digital representations could be reanalyzed in network terms. I then show how the network framework reframes the motivating borderline cases. The interactive network view handles the ‘hard cases’ like core cognition and higher-level perception with ease: It explains how we can perceive a number or a substance kind directly, why infants can reason about objects before language, and how conceptual thinking can utilize sensory-like simulations without breaking the unity of the cognitive system. This is both empirically plausible and theoretically satisfying, since it replaces an unnecessarily strict format boundary with a fluid account of how perceptual and conceptual information could interact.
Jason Guo (University of Cambridge): Explaining with Networks in the Social Sciences
Given the rise in and success of network analysis in the life sciences (e.g., ecology, systems biology, gene regulatory networks), many philosophers of biology have offered accounts of the epistemology of network analysis, with some claiming to provide a general framework that is also applicable outside of biology.
In this literature, three standard philosophical questions about scientific explanations are commonly asked about network-based explanations:
Q1: When are/should network explanations (be) invoked in place of other types of explanations?
Q2: When do networks successfully explain, and what grounds their explanatory power?
Q3: How does the explanatory strategy of network explanations compare to other types of explanations?
Despite most case studies in the philosophy of network science being drawn from the life sciences, network-based thinking arguably has an even richer history in the social sciences than the natural sciences, tracing its intellectual lineage back to Simmel’s social theory and early 20th-century sociometry. Hence, I aim to address the aforementioned questions concerning network analysis from a distinctly social scientific perspective.
To make this task manageable, I hone my analysis on one dominant usage of social network analysis (SNA), answering the question of how social scientists leverage network analysis to explain the emergence of social phenomena, understood as the macro-level attributes of social systems.
I formulate my thesis in the form of responses to each of the three questions, situated in the context of the social sciences.
A1: Network-based explanations are/should (be) invoked when, with reference to the target social phenomenon, our system behaves as a non-decomposable system where agential states/behaviours, even in the short run, exhibit a non-negligible dependency on other agential states/behaviours.
A2: Social network explanations are successful when they identify relational processes/mechanisms that produce the observed dependencies between agential states/behaviours, thereby increasing the expectability of the target attribute of the non-decomposable system.
A3: Social network explanations of social phenomena are causal explanations that rely on essentially the same explanatory strategy as other types of causal explanations in the social sciences (e.g., structural and individualist explanations), where explanatory success depends on the identified causal factor contributing to an intuitive or otherwise compelling causal story of how the phenomenon emerged.
I support my argument through an extended case study of how social scientists use SNA to investigate the relationship between various kinds of interpersonal ties and social outcomes (e.g., success in the labour market and the diffusion of innovations). Through this case study, I demonstrate how the explanatory relevance of network concepts depends on their ability to identify causally relevant relational mechanisms, and how their relevance can be derailed by behavioural and cultural factors.
Finally, I end by: (1) contrasting my account of network explanations to a dominant perspective which claims that network analysis provides a non-mechanistic explanatory strategy; and (2) exploring the implications my account has for actual social scientific practice.
Teemu Lari (Stockholm University): Disagreement on the feasibility of sustainable economic growth: Where does it stem from?
Researchers across disciplines and fields have differing and even opposite views on the relationship between environmental sustainability and economic growth (Drews & van den Bergh, 2017). For those convinced of the ideal of "green growth", economic growth can be separated from harmful environmental impacts like carbon emissions, and thus governments can and should continue aiming for growth even in the face of serious environmental problems. For researchers subscribing to ideals like "degrowth" or "post-growth", decoupling economic growth from environmental problems does not seem like a realistic prospect. Accordingly, they argue that countries with a high standard of living should give up the pursuit of further growth to protect the environment.
The disagreements on sustainability and economic growth have existed for decades and show no signs of being resolved. They seem to have even intensified as the evidence on environmental problems has been getting more alarming (Richardson et al., 2023). Attempts at constructive interaction and dialogue are often perceived as unproductive and marred by misunderstandings. However, persistent scientific disagreement on such a pressing issue is problematic, not least because the audiences of scientific advice, such as policymakers, may not have a well-justified way of choosing which side to believe.
I aim at clarifying what makes the disagreements so difficult to resolve, since a better understanding of where the disagreements stem from could help resolve them. I will analyse argumentative texts (both academic and policy-oriented) by researchers on both sides of the divide. My central references are recent publications by prominent and influential authors advocating sustainable or "green" growth (Stern, 2025; Susskind, 2025), "degrowth" or "post-growth" (Parrique, 2025; Schmelzer et al., 2022) as well as those taking intermediate positions (Van den Bergh, 2018). Previous research in philosophy, especially philosophy of science, suggests several factors that can contribute to the persistence of scientific disagreements. One possibility is that central concepts such as "economic growth" are understood in different ways in different research fields. A related possibility is that modal concepts such as "feasible" and "possible" are used differently. In addition, it can be that values play a role, such as when evidence is inconclusive and ethical values give rise to differing responses to inductive risk (Douglas, 2000), or when theory or model choice is affected by differently emphasized epistemic values (Kuhn, 1977). Yet another possibility is that disagreements stem from differences in methodological norms and conventions, for example from different conceptions of how models should and should not idealise their targets (Mäki, 2004). Finally, all these considerations can be intertwined, so that for example normative commitments affect preferences over definitions and methodological choices.
While those involved in the growth vs. sustainability debates have their views on what causes the disagreements (Drews & van den Bergh, 2017), this paper will be the first one to examine the issue by applying tools from philosophy of science to analyse scientists’ argumentation.
Luna De Souter (Witten/Herdecke University): Machine learning for philosophy-based causal discovery
Configurational Comparative Methods (CCMs) are a family of methods for automatically discovering complex causal structures from datasets. When applied to a dataset, these methods algorithmically build models that approximately satisfy the criteria of causation spelled out by philosophical theories within the INUS framework (where INUS stands for "Insufficient but Non-redundant part of an Unnecessary but Sufficient condition") and therefore are good candidates for representing the causal structure of that dataset. CCMs have been frequently used for INUS causal discovery in social science and healthcare research (Baumgartner and Falk 2023).
A central limitation to the applicability of CCMs is that they are unable to analyze datasets comprising more than 30 variables, while many scientific datasets comprise considerably more variables than this. Therefore, until now, several variables must often be eliminated from a dataset before applying CCMs, which may jeopardize the reliability of the results of CCM analyses.
In this presentation, I argue that the Boolean Rule Column Generation (BRCG) algorithm (Dash, Günlük, and Wei 2018), which was originally developed by machine learning researchers as an algorithm for building interpretable binary classification models for datasets comprising hundreds of variables, can and should also be used as an algorithm for INUS causal discovery.
My main argument is conceptual: I demonstrate that, if the BRCG algorithm succeeds at optimizing its objective function, it returns a model that best satisfies the criteria of the most current INUS theory for the analyzed data. Additionally, since the BRCG algorithm is not guaranteed to optimize its objective function, I complement this conceptual argument with findings from simulation studies. These studies show that the BRCG algorithm performs similarly to the popular CCM algorithm of Coincidence Analysis (CNA) for datasets small enough for CNA to handle and that the BRCG algorithm reliably recovers causal structures from datasets with more variables than CNA can handle. These findings indicate that replacing current CCM algorithms by the BRCG algorithm will result in more reliable INUS causal discovery.
Julien Dutant (King's College London): Measuring Reliability
“Reliability” (in epistemology) a.k.a “accuracy” (medical diagnosis, machine learning, statistics) is the overall correctness of an investigative procedure, a belief-forming process, a classifier, a measuring instrument, or the like. It is standardly characterised it as a ratio of truths among the method's outputs. In the simple case (a method that only ever outputs one proposition, p), this is the same as a positive predictive value (the probability of p conditional on the method outputting p, for some relevant measure of probability). We put forward an intuitive constraint that any such measure should satisfy, analogous to a single-premise closure principle: namely, that replacing a method’s outputs by logically weaker ones should preserve its reliability. We show that when we move to propositions with multiple outputs, the truth ratio measure and variants of that idea violate the constraint. In a nutshell: when the method's outputs are weakened, two or more true outputs may merge into one (‘p’, ‘q’ both weaken to ‘p or q’) and one false output may split into two or more (‘p and q’ weakens to ‘p’ but also to ‘q’), thereby lowering the method's truth ratio. We argue that, rather than rejecting the constraint, we should adopt a reliability measure consistent with it. We put forward one such measure, according to which a method’s reliability is the subject-matter weighted average truth-value of its outputs: we fix the total weight of the set of its outputs on a given subject-matter irrespective of their number. We suggest that the way the notion of reliability is deployed in empirical applications can consistently be interpreted as endorsing such a measure. While the measure preserves the principle, it comes with its own theoretical costs.
Tobia Fogarin (IMT School for Advanced Studies Lucca): Reconciling Epistemic Utility Theory and Truthlikeness: An Alternative to Proximity
Epistemic utility theory provides arguments for norms governing degrees of belief, such as probabilism, conditionalization, and the principal principle. The standard strategy is to show that if a credence function violates a given norm, there exists an alternative credence function that satisfies the norm and is epistemically more valuable independently of what turns out to be true (Pettigrew 2016). Epistemic value is understood as depending solely on how well a doxastic state represents the world, independently of the consequences of acting on it. On this approach, all epistemically relevant considerations in a given context are captured by an epistemic utility measure, which is then required to satisfy certain general conditions. Some authors have highlighted the important role that truthlikeness can play in legitimate epistemic utility measures (Douven 2021). A proposition is said to be more truthlike the better it approximates the most informative true proposition in a given set. However, a tension between truthlikeness and epistemic utility theory has been observed by Oddie 2019. Oddie proposes a condition on measures of epistemic value, Proximity, intended to capture considerations about truthlikeness, and argues that Proximity is incompatible with Strict Propriety, a condition assumed in most epistemic utility arguments. Other authors have tried to solve this incompatibility by proposing weaker versions of Proximity (Schoenfield 2019, McCutcheon 2023).
The aim of this talk is to challenge the alleged incompatibility and the tentative solutions proposed in the literature. We argue that Proximity-like conditions do not capture the way in which considerations about truthlikeness can impact the epistemic life of an agent according to the truthlikeness literature. An influential approach in this literature for assessing closeness to the truth without knowing the truth is to consider the expected truthlikeness of the most informative proposition entailed by a theory. Accordingly, an agent can choose a set of qualitative beliefs by selecting the proposition that maximizes expected truthlikeness relative to a credence function, and then believing its consequences. Epistemic utility measures satisfying Proximity-like conditions do not assign higher scores to agents who make better judgments in this sense. Instead, we encode this idea of the value of truthlikeness in what we call the Truthlikeness Property. The Truthlikeness Property requires that, at a world w, if adopting credence function p leads an agent, via this procedure, to select a proposition that is strictly more truthlike at w than the proposition selected under credence function q, then p must be epistemically more valuable than q at w.
We argue that this move radically reshapes the debate over the alleged incompatibility. Our argument proceeds in three steps. First, we show that Proximity-like conditions cannot capture our conception of truthlikeness. Second, we propose a method for constructing epistemic utility measures that are both strictly proper and satisfy our property. Finally, we discuss the broader implications of our result for epistemic utility theory. We argue that these results can help to foster interaction between the epistemic utility and truthlikeness literatures and deepen our understanding of epistemic value.
Rafael Fuchs (LMU): Statistics Wars: Return of the Bayesians
Turmoil has engulfed the scientific community. The 'statistics wars' (Mayo 2018) between Bayesian and frequentist schools are still raging on (Radzvilas et al. 2021). In this contribution, I present a framework of epistemic accuracy (Pettigrew 2016) and efficient experimentation (Crupi et al. 2018) that aims at shedding new light on the most fundamental issues of the dispute. One major point of contention is the likelihood principle (LP), and related issues concerning admissible stopping rules. Intuitively, the LP says that all information an experiment conveys about a set of statistical hypotheses is contained in the likelihood function (i.e. the conditional probability of the data given any particular hypothesis). The LP entails that, standardly, the 'stopping rule' (the decision criterion by which the experimenter decides when to stop collecting data) doesn’t matter for inferences from the data about the hypothesis.
While standard Bayesians endorse the LP, there are strong reservations on the frequentist side: in particular, accepting the LP seems to permit a 'persistent experimenter' to keep collecting data until they establish any conclusion of their liking, independently of its truth (Mayo & Kruse 2001, Fletcher 2024). At the same time, however, it has been shown that Bayesian inference allows for 'optional stopping' (i.e. keep experimenting until the posterior reaches a target decision threshold) while being calibrated to desired frequentist error rates, at least in some scenarios (Rouder 2014). Optional stopping can make data collection much more efficient than following standard (a priori) power analyses, provided it comes with no loss of accuracy — but when does this hold, and what does it tell us about the underlying issue concerning the LP?
Based on the framework of epistemic accuracy, we can derive the following argument. First, minimising expected posterior inaccuracy for a given set of observations of random variables together with some weak background conditions entails Bayesian conditionalisation (e.g. Rosenkrantz 1992). Simultaneously, for any strictly proper scoring rule, minimising expected inaccuracy entails convergence to actual relative frequencies (and relatedly, being able to achieve target error rates). Hence, whenever the evidence is such that conditionalisation applies (i.e. results from minimising expected inaccuracy), following the LP is normatively justified. Yet, even in cases where the stopping rule is non-informative, the stopping time can still be highly informative — which resolves the apparent puzzle discussed in Mayo & Kruse (2001). More generally, however, in cases with more complex evidential constraints, the accuracy-based framework produces updating- and inference rules that are more general than standard conditionalisation, and allows for a unified perspective on Bayesian and frequentist estimation techniques. Relatedly, epistemic accuracy also integrates with alternative Bayesian approaches that may violate the LP (like Bernardo 1979). In summary, the more fundamental principles of epistemic accuracy provide a criterion that can sharply determine under which conditions the LP follows (and thus, normatively, the stopping rule doesn’t matter), and when it does not. This will help us to unify Bayesian and frequentist perspectives and eventually take steps towards resolving a deeply entrenched philosophical conflict.
Ina Jäntgen (Ludwig Maximilian University): When broken is not worth fixing: a decision-theoretic approach to adjusting effect sizes for meta-biases
In the biomedical and social sciences, researchers often quantify the effectiveness of tested interventions using effect size measures, and these effect sizes are often amalgamated in meta-analysis. Such amalgamated effect sizes widely inform evidence-based decision-making, from clinical choices to policymaking. However, such amalgamated effect sizes can be, and often are, subject to meta-biases—biases affecting the body of evidence on which an amalgamated effect size is estimated as a whole, such as publication bias and funding bias (e.g., Murad et al. 2018).
In response, several philosophers suggest that researchers should, whenever possible, adjust amalgamated effect sizes to account for meta-biases (e.g., Erasmus 2023; Stegenga 2018). Indeed, meta-analysts have developed methods to adjust effect sizes for meta-biases, mostly for publication bias (e.g., Sladekova et al. 2023), and such methods are increasingly used. Many of these methods, however, require strong statistical assumptions, expertise beyond standard meta-analytic techniques and can themselves be used in biased ways. As a result, Cochrane—the main network for meta-analyses in medicine—even advises against adjustments of effect sizes in their meta-analyses.
In this talk, I first provide a reason to doubt the prevalent philosophical recommendation: I argue that a biased and a correctly adjusted effect size can, in some cases, provide a decision-maker with equally valuable information to rationally choose between interventions. Whether adjusting an effect size adds information valuable to a rational decision-maker depends on (i) features of the decision-maker, (ii) the effect size measure and (iii) the degree and type of meta-bias. As a result, to inform decision-makers, adjusting effect sizes for meta-biases is sometimes beneficial, and sometimes it is not. Paired with the mentioned challenges of adjusting, this observation sheds doubt on the idea that researchers should, whenever possible, adjust. To establish this result, I develop an expected utility model for treatment choices informed by biased effect sizes, expanding decision-theoretic work on effect sizes (e.g., Sprenger and Stegenga 2017) to the context of meta-biases.
Using my decision model, I then identify conditions under which adjusting is valuable for decision-makers. More precisely, I specify for which groups of rational decision-makers merely ruling out that an amalgamated effect size is biased beyond some specific degree is just as valuable as learning the precisely adjusted effect size would be, and for which groups learning the precisely adjusted effect size would be more valuable. Here, I focus on three common effect size measures, the mean difference, the relative risk and the risk difference. Such results can guide meta-analysts as to when adjusting effect sizes is valuable.
I conclude by discussing an objection against the feasibility of targeting adjustments towards rational decision-makers’ needs, and another one concerning the value of accurate estimates for patients’ informed consent and scientific progress. Overall, this paper draws on the perspective of rational decision-making to develop a context-sensitive approach to adjusting effect sizes for meta-biases.
Zhitao Zhang (Marche Polytechnic University): A causal account for estimating effects across populations when no causal information is available
Uncovering causal structure is often thought to involve two steps: ‘the inference from sample correlations to the probabilities of the underlying populations, and that from probabilities to causes’ (Reiss, 2005). Currently available causal search algorithms operate on population probabilities (Pearl, 2009; Sprites et al., 2000). Little attention has been paid on obtaining the population distributions. This paper argues that the stochastic nature of data raises problems in the first step and develops possible solutions by drawing on insights from statistics.
In the first part of the paper I spell out the problem in more detail. Data often fail to reflect population probabilities, but rather fluctuate around them within a certain range. Data sampled from uncorrelated events can show correlations and vice versa. This problem persists even as the sample size grows larger, and does not arise due to disturbing causal factors, but stems from the stochastic nature of data. It is well-known and has been thoroughly investigated by statisticians who developed significance tests to address it. If, for example, the dependence of cancer on smoking should be tested and the data display a correlation, the general idea is to reject the null hypothesis that ‘smoking does not make any difference for cancer’. Next, the likelihood of observing the data we collected under the assumption of the null hypothesis is calculated, where the ‘rareness’ of the actual data under the null hypothesis corresponds to the confidence level of the observed correlation. Setting up the null hypothesis and a pathway to obtain the target likelihood is essential for the test, but how to do so for an inference of causal structures remains unclear.
In the second part I explain Fisher’s original significance test and the strategy developed within the Rubin Causal Model and explore how they may help to address the problem for causal inference. Fisher’s approach takes the outcome variable as independently distributed, which is decided by statistical conditions (random sampling) and assumptions (normal distribution). The test based on this univariant distribution is homogeneous for different causal structures between the observed variables, and thus not suitable for causal inference. Rubin (1972) developed his potential outcome framework for causal inference and a modified version of the significance test. A key advancement of this approach is that the distribution of the outcome under the null hypothesis is derived from the underlying causal structure. The outcome of this significance test would thus depend on the specific causal hypothesis at hand.
However, Rubin did not yet have the resources of modern causal modelling approaches available and, thus, his approach rests more on causal intuitions than on a full-fledged account of causal inference. The third part of the paper will bridge that gap by connecting Rubin’s ideas to modern causal inference resulting in the development of a causally warranted prototype of a significance test for specific causal structures. It will also provide an outlook on the extent to which underdetermination issues from the causal discovery literature may be overcome by such a causal significance test.
Srijan Butola (University of Copenhagen): Personal Identifiers - How the Social and the Political Matter for Data-Intensive Health Research
Science journalist Lone Frank, in her 2000 Science article, titled “When an Entire Country is a Cohort”, posits that Denmark provides a unique case for health researchers to study medical phenomena at the level of the population. This, she argues, is due to the scope and robustness of medical registries and socioeconomic databases, and the possibility of linking them near-instantaneously to render relevant insights. For Frank, as well as several Danish epidemiologists (Schmidt et al., 2010; Schmidt et al., 2014; Arendt et al., 2020), what makes such research possible is a unique 10-digit personal identification number, the CPR-number, that every Danish resident carries with themselves. This identifier is used for a plurality of purposes – from personal banking to accessing public healthcare to securing university and college admissions within Denmark, among other uses. It is this number that helps various authorities within Denmark to collect data about the Danish population, and, in Frank’s perspective, render the nation as a cohort.
Personal identifiers such as the Danish CPR-number must thus be treated as key objects of analysis in our understanding of data-intensive health research. In the wake of the COVID-19 pandemic, the population has been revitalized as a site of health research, primarily from the perspective of designing disease surveillance infrastructures (implemented at national and supranational levels, such as the EU’s OneHealth4Surveillance project) that seek to integrate non-medical databases with medical ones to predict and prepare for future global health events. The question that then arises is: how do personal identifiers make such data integrations possible?
The aforementioned literature as well as the that from the field of bioinformatics (McMurry et al., 2017) centre the role of identifiers in enabling the mobility of data across different sites of use. From this perspective, these identifiers are akin to the keys in relational database systems that help link data across different databases. But this only covers the computational aspect of these identifiers. Given that individuals use these identifiers in a variety of contexts, that governments rely upon these identifiers to produce knowledge about their subjects, that the data collected happens over particular media and through the use of particular markers of identification (such as names, addresses, biometrics, and so on), personal identifiers present a much more complicated story of how data-intensive health research is actually made possible.
This paper utilizes a comparative approach between two systems of person identification – Denmark's CPR-number and India’s Aadhaar number, two vastly different political and social contexts, to bring forth the role of various social forces, practices and rationalities that coalesce around the personal-identifier that make particular forms of data mobility possible. In this way, this paper aims to contribute to the existing literature on data journeys (Leonelli, 2016) which have thus far focused on practices that make data available for re-use across different contexts. This paper will as a result argue for the centrality of the social and the governmental in conceptualizing data mobility that make contemporary health research possible.
Benjamin Conover (Saint Louis University): Examining Newton’s Geometric Strategy for Certainty in Natural Philosophy
The central role of geometry in Sir Isaac Newton’s natural philosophy has long been emphasized by scholars such as Mary Domski, Niccolò Guicciardini, and D.T. Whiteside. Newton himself makes this importance clear in his First Preface to The Mathematical Principles of Natural Philosophy (hereafter Principia), where he famously tells readers that geometry “appertains to mechanics” and is “nothing other than that part of universal mechanics which reduces the art of measuring to exact propositions and demonstrations (C&W 381-382). Similarly in his early Optical Lectures, Newton urges geometers and natural philosophers to study each other’s subjects: “[T]ruly with the help of philosophical geometers and geometrical philosophers, instead of the conjectures and probabilities that are being blazoned about everywhere, we shall finally achieve a science of nature supported by the highest evidence” (Newton 1670-72, 437-439). For Newton, the “highest evidence” of certainty in geometry comes from synthesis — the construction of, e.g., a figure from better-known postulates such as, e.g., angles and lines — following a geometric analysis — working back from the postulated target, e.g., a figure, via better-known postulates such as, e.g., angles and lines. To inject that certainty into natural philosophy, Newton proposes a transference of the method of analysis and, most importantly, synthesis to natural philosophy (see Guicciardini 2009, Chapter 14). In his natural philosophy, analysis sets out “to discover the forces of nature from the phenomena of motions” while synthesis aims “to demonstrate the other phenomena from these forces” (C&W 382).
In what follows, I contend that Newton’s geometric strategy for certainty in natural philosophy merits additional focus from contemporary philosophers, even though it suffers from significant difficulties. I begin by briefly addressing his understanding of geometry and how its method of analysis and, especially, synthesis brings certainty to the mathematical discipline. In line with both Domski and Guiccardini, I highlight how Newton’s commitments in (what we would now call) the philosophy of mathematics concerning geometry underpin the success in the Principia of describing previously unaccounted for celestial phenomena. Next, I explicate how Newton transfers the method of analysis and synthesis into natural philosophy. Importantly, I show that he adds an extra step not found in the geometric version: the application of analyzed forces to new phenomena. Although Guiccardini raises three difficulties for Newton’s methodological transfer even raised by his contemporaries — interpretative, epistemic, and causal non-uniqueness in physics contra geometry (see 2009, Chp. 14.3) — I argue that this new difficulty proves more challenging for Newton’s aspirations of certainty. Notwithstanding the difficulties, I conclude with some provisional remarks on the positive lessons regarding the applicability of mathematics in physics that contemporary philosophers can take from Newton’s geometric strategy for natural philosophy.
Paric Harting (Lingnan University): Blind Peer Review with Accountability: A Game-Theoretic Approach
Academic peer review plays a central role in modern scientific publishing (Drozdz and Ladomery, 2024). But there are also growing concerns over the quality, effectiveness, and sustainability of the practice. Heesen and Bright (2021) have argued recently that (pre-publication) peer review should be abolished, while Rowbottom (2022) offered a defence. In this paper, I examine the tension between the benefits of blind review and reviewer accountability and suggest a practical way to get the best of both worlds.
Double- or even triple-blind peer review is held by some researchers to be the gold standard for academic publishing (Brodie et al., 2021). Blind review is thought to offer a number of benefits, such as reducing unfair bias against certain groups of researchers, mitigating social phenomena such as the Matthew effect, and also protecting reviewers from unfair retaliation. But anonymity also has potential downsides. Anonymous reviewers face little to no consequences for producing low-effort or low-quality reports. And, to make matters worse, they also have no extrinsic incentives to produce high-quality work, as doing so results in no professional credit. Taken together this leads to a range of concerns from unreliability to subversion, against which open review has been suggested as a potential solution (Ross-Hellauer, 2017).
I propose and defend a hybrid approach that preserves that strives to maintain most of the benefits of both blind and open review practices. All submissions are handled under the usual double- (or triple-) blind procedures until an editorial decision has been reached. After that, a fixed small proportion of completed reviews and reviewer identities are randomly selected for public release. Crucially, neither editors, reviewers, nor authors know in advance whether any particular review will be revealed.
My central argument is game-theoretic. Ordinary blind review has a weak incentive structure: there are no costs for poor work, and no rewards for good work. While a fully open review process results in strong accountability, it also comes with the mentioned downsides. My proposal changes the incentive structure for reviewers. It increases the expected cost of poor conduct, due to the chance of discovery and it also increases the expected benefit of providing high-quality, well-argued reports for the same reason. Finally, my proposal maintains any existing protections for authors, while still retaining most of the protection for reviewers. I formalize this change to the incentive structure with a simple mixed-strategy model to explore the range of reasonable reveal rates which move the equilibrium behaviour towards higher effort without reproducing the incentive problems of either extreme.
I then address the key design question of how to set an appropriate reveal rate. I suggest that there are several plausible approaches, like using anonymous author satisfaction surveys as a feedback signal. Finally, I consider whether random-reveal might make it harder to recruit reviewers and consider mitigating incentives. Overall, the random-reveal review process is a practical institutional improvement that can increase review quality through accountability while retaining the core fairness benefits of blind review.
Anastasiia Lazutkina (University of Wuppertaö): How evidence is made: measuring the cosmic microwave background radiation before discovering it
In philosophy of science, empirical evidence is often treated either abstractly through formal accounts of confirmation and inference (Earman 1992; Howson and Urbach 2006; Mayo 1996), or else through detailed accounts of how data are stabilized and structured so as to play an evidential role (Leonelli 2016; Boyd 2018). What is less often spelled out is the mechanism that allows structured data to become comparable with theoretical models in the first place so as to function as evidence for or against a theory. I aim to spell out how this mechanism works through a reconstruction of the historical context of one of the most significant discoveries of the twentieth century: the detection of the cosmic microwave background (CMB). The CMB is a nearly isotropic blackbody radiation that was predicted by Big Bang cosmology already in the 1940s but only discovered in 1965. I show that the same radiation was in fact measured several times between the 1940s and 1950s but was not interpreted cosmologically (see Kragh 1996; Peebles, Page and Partridge 2009). I thus ask: what changed in 1965 that allowed the same kind of measurement to count as evidence for the Big Bang theory?
I argue that this transformation was twofold in a way that previous accounts have not completely explained. First, it involved what Boyd (2018) calls enrichment: the careful work by Penzias and Wilson (1965) that made their residual antenna temperature a stable, reproducible result. Second, it required what a paper by Dicke, Peebles, Roll, and Wilkinson (1965) supplied: a “schematization of the observer” (Stein 1994, Curiel 2020) that defined where and how to look for the CMB, what kind of apparatus would register it, and what would count as contamination or noise. Only when these two layers, empirical enrichment and observer schematization, were combined did the result become comparable with the theoretical model of relic radiation from the early universe. The CMB case illustrates how a representation of the observer or laboratory instrument must be integrated into the theoretical framework for data and theory to become genuinely comparable, so that evidential relations can be established at all. I argue that practices related to instrument calibration, background subtraction, and rules for treating residuals help to constitute evidential relations rather than merely refine them.
Therefore, while I consider previous accounts to provide valuable pieces for solving this puzzle, none of them alone suffices. Evidence is not merely stabilized data nor merely interpreted theory: it is the product of a coupling between enriched evidence and observer schematization, which jointly establish theory–data comparability. The CMB case shows that enrichment and schematization can come apart and that discovery only happens when they are coupled, but not when one is present without the other.
Arlene Lo (London School of Economics and Political Science): Fengshui and the Spectre of Scientific Demarcation
Fengshui, historically and contemporarily, is often taken as a paradigmatic example of pseudoscience. In line with this received conception, philosophers evaluated fengshui with respect to the criteria of scientificity, argued that it would not fall within the bounds of science so demarcated, and concluded that it is epistemically corrupt (e.g. Matthews 2019, Bhakthavatsalam and Sun 2021, Fernandez-Beanato 2025). By examining the actual history of fengshui as a socio-legal practice and mode of reasoning, I argue that analysing fengshui – by implication many other non-Western epistemic traditions – with respect to their (pseudo)scientificity yields these inadequacies:
1. Category error: demarcation is misguided as the aims of fengshui are not scientific in nature.
2. Epistemic distortions: demarcation distorts the epistemic products and structures of fengshui.
3. Value of transdisciplinarity: demarcation is incapable of accounting for the (social) epistemic and instrumental value derived from transdisciplinarity.
Using Brown’s (2023) Qing historical study, I contend that fengshui, as an epistemic tradition, aims at ‘auspiciousness’ of socio-technical interventions with the land (e.g. graves, residences, infrastructure). Although there are naturalistic components to fengshui (e.g. economy, environmental conservation, health), the ends are often social in nature (e.g. social security, political stability). The epistemic practices, in turn, concern the creation and application of practical knowledge re designing and building structures given local environmental and social conditions.
On (1): The felicitous ascription of pseudoscientificity requires fengshui’s aims to resemble those of science. However, fengshui appears to be transdisciplinary in resembling a variety of aims of scientific and non-scientific disciplines: environmental science, economics, geography, architecture, law, politics, etc. The demarcation analysis is thus committing a category error of assuming that fengshui has solely scientific aims (thus should be compared to scientific practice) when it doesn’t. Therefore, the resulting epistemic evaluation from demarcation is irrelevant.
On (2): Felicitously evaluating only the ‘scientific parts’ of fengshui in demarcation analysis introduces two types of epistemic distortions, making the epistemic evaluation inaccurate:
1. Artificial deflation: the demarcation analysis misses out on substantive claims relating the empirical to the socio-legal. The apparent deficiencies in fruitfulness and explanatory power, therefore, are mere artefacts of applying a disciplinary perspective to transdisciplinary inquiries.
2. Imperfect severing: fengshui concepts often have intertwined meanings across different domains (empirical-physical, socio-legal, cosmological-spiritual). It is difficult to completely sever the extra-scientific connotations in fengshui concepts. Therefore, the apparent pseudoscientificity is often an artefact of such imperfect severing.
On (3): Transdisciplinary inquiries like fengshui yield important epistemic benefits through synthesising multiple domains of knowledge. There are also instrumental benefits of fostering knowledge sharing and productivity gains. The disciplinary perspective of science in the demarcation analysis inherently cannot capture these transdisciplinary gains in epistemic evaluations. Therefore, the epistemic evaluation will always remain incomplete.
Since its first Western encounter, discourse is always centred on fengshui’s scientificity as a reproach of the Oriental Other. However, scientific demarcation only misleads, distorts, and obscures our epistemic evaluation of these non-Western traditions. If our interest is to understand their epistemic value, we must treat them as epistemic traditions in their own right.
Kirsten Walsh (University of Exeter): Navigating Agential Distance: Margaret Cavendish on the Pitfalls of Experimentation
Margaret Cavendish was critical of the experimental practices of the early Royal Society, likening the early modern experimental philosophers, such as Robert Hooke and Robert Boyle, to “boys that play with watery bubbles or fling dust into each other’s eyes, or make a hobbyhorse of snow” (Cavendish 1666/2001, 52), calling them “worthy of reproof rather than praise, for wasting their time with useless sports” (ibid.). But her objection was more than simply that experimental philosophy was a childish pursuit. She had serious concerns about the efficacy of experimentation to achieve the very goals she shared with them, namely, to understand the workings of nature so that we might use that knowledge to improve human lives.
In this paper, I use Cavendish’s objections to experimental philosophy to explore a notion of ‘agential distance’ (developed from Murphy et. al. Forthcoming), both developing that concept and showing that it generates a deeper understanding of Cavendish’s work. ‘Agential distance’ refers to the relationship between a scientific agent’s capacity and their target of inquiry. In Newton’s optical experiments, for example, when he uses a prism to study the refraction of light, we take that distance to be very small, due the immediacy between Newton’s interventions and the phenomenon he examined. In contrast, when a team of scientists use data produced by the LHC to study dark matter, distance is much greater: there is significant delay between scientific intervention and phenomenon identification, moreover, the phenomenon operates at a wildly different scale from the agents themselves. Agential distance presents an epistemic challenge; one that might be mitigated with the help of various instruments, interventions and computational techniques. Such techniques might be said to expand ‘agential reach’ and so mitigate agential distance.
Early modern experimental philosophers recognised that human capacity is limited by their senses. This was potentially a big problem, given that they also believed natural philosophy should be based on matters of fact (information drawn from the senses) rather than speculation. Viewing (a) the senses as assets that could be aided and enhanced and (b) the natural world as something that could be coaxed and manipulated into revealing its underlying processes, experimental philosophers employed all sorts of instruments and interventions to mitigate agential distance. Hooke, for example, used microscopes to drill down into minutiae and study the inner workings of things (Hooke 1966/1665), and Boyle used the air pump to study the ‘spring’ of the air (Boyle 1999/1660). Cavendish agreed that human sensitive capacity is limited. However, for her, using instruments and other techniques would hinder, rather than help. This was because (a’) instrumentation would distort rather than enhance and (b’) experimental interventions would generate artificial rather than natural phenomena. She agreed that humans could extend their reach. But this could be done by exercising their rational capacities, rather than attempting to enhance their sensitive capacities: “our exterior senses can go no further than the exterior figures of creatures, and their exterior actions: but our reason may pierce deeper, and consider their inherent natures, and interior actions” (Cavendish 1666/2001, 100).
I demonstrate how we can use the notion of ‘agential distance’, and the differing attitudes of Cavendish and Hooke towards distance and its mitigation, to make sense of their differing attitudes towards experimental practices.
Anjan Chakravartty (University of Miami): Philosophical Imperatives for Science Education and Public Understanding
It is often suggested that improving the public understanding of science could play an important role in tackling crises of trust in science. Discussions of this in the philosophy of science, however, rarely engage with a crucial part of shaping such understanding: what we might call general science education – the extent of formal training that the vast majority of citizens will receive, running its course in secondary or early higher education. This talk explores the issue of what the aims of this education should be if it is to serve the purpose of improving public understanding, arguing that this purpose is best served by framing the teaching of science in terms of an emphasis on the doxastic attitude of acceptance, which is compatible with but does not require belief.
After motivating the idea that general science education is a significant determinant of later public understanding, I focus on the most influential proposal to emerge in recent decades for reforming science education, which is to incorporate in teaching science a more substantial emphasis on the nature of science, comprising a study of methodological, socio-historical, and epistemic features of scientific practice (e.g., Matthews 2015) – commonly abbreviated as ‘NOS’. While this proposal has garnered substantial approval in education circles, I contend that its typical formulations are unlikely to achieve the intended result of promoting better public understandings of science.
It seems clear, for instance, that the extraordinary methodological breadth of the sciences cannot be conveyed in a general science education, which at best encompasses an introductory exposure to certain highly specific domains of science. This is furthermore unlikely to promote an understanding of what may later manifest as pressing scientific issues facing society, which (today) includes things like climate modeling and artificial intelligence. The “nature of science” is also extremely difficult to characterize in any very determinate way. Some advocates of NOS (e.g., Ennis 1979; McComas et al. 1998) attempt to do this by offering “consensus lists” of descriptive features of science, but as is widely appreciated in the philosophy of science (if not in science education and communication), the precise epistemic status of our best science is subject longstanding and ongoing debate (cf. Alters 1997; Chakravartty 2023).
I suggest a much more specific proposal for the framing of general science education, emphasizing a genuinely consensus feature of otherwise conflicting epistemologies of science: the idea of empirical success (predictive, interventional, manipulative, etc.). This conception is properly associated with an epistemic attitude of acceptance: a commitment to using something as though it were true for some purpose – for example, the use of biochemical models to develop promising medical treatments, or the use of physical models to improve vehicle safety – as a basis for making decisions about how to act. While acceptance is compatible with belief, it is less demanding in that it does not require belief. This, I argue, while leaving the door open to belief, also leaves the door open to those susceptible to forces of science skepticism.
Alexander Christian (Heinrich Heine University): Silent moral complicity in science governance
Moral complicity is traditionally understood as the status of a secondary agent who becomes a wrongdoer by virtue of their relationship to a primary wrong. In contemporary moral philosophy, this relationship is typically explained through two dominant accounts: (i) the causal contribution account, which grounds blameworthiness in the actual difference an agent’s actions make to a harmful outcome (Gardner, 2007), and (ii) the shared intention or participatory account, which focuses on an agent’s participatory intention to contribute to a collective goal, regardless of their specific causal efficacy (Kutz, 2007; Lawson, 2013).
However, as (Donohue, 2024) has recently argued, these accounts struggle to handle silence as complicity. Causal accounts often fail because silence rarely makes a demonstrable causal difference in large-scale injustices, while intentional accounts fail because silent bystanders often lack a positive intention to participate in the wrongdoing. To bridge this gap, Donohue proposes a third account: deliberative complicity. Rooted in recent work in speech act theory and social epistemology (Lackey, 2020; also McCain & Stapleford, 2020), this account posits that agents have interpersonal deliberative obligations to object to or speak out against wrongdoing. Complicity, on this view, arises when an agent fails to fulfill a duty of due care regarding the moral deliberation and beliefs of others.
While moral complicity is well-documented in research ethics — such as the outsourcing of controversial human embryonic stem cell research to less-regulated jurisdictions (Devolder, 2015) — a significant research gap exists regarding complicity within the deliberative processes of science governance. This paper applies the deliberative complicity account to analyze various forms of deliberative complicity in science governance, focusing on examples from regulatory debates concerning notoriously controversial biomedical research, e.g. animal research, gain of function research, xenotransplantation.
The paper defends three theses: First, I argue that scientists possess vastly more (strong) deliberative obligations in these contexts than laypersons due to their professional roles in research processes, specialized expertise, and unique epistemic access to potential harms. Second, I contend there is no “opting out” of moral discourse in science governance. This claim is grounded in a re-evaluation of the ideal of scientific freedom (Wilholt, 2010): if scientific freedom grants (relative) negative rights against external intervention, it necessitates a reciprocal positive duty to engage in robust deliberative self-regulation. Silence in the face of flawed science governance is consequently not a neutral stance but a failure of this professional duty. I will discuss various objections to this claim, e.g. moral overdemandingness (cf. Stern, 2020) and claims to a right to moral deference (cf. Davia & Palmira, 2015). Finally, I show that current science governance structures often cultivate deliberative failures by organizing deliberative processes in ways that enable silence and marginalize dissent (author). By acting-as-though certain practices are morally settled through their silence (Donohue, 2024), scientific communities may vastly underestimate internal moral dissent by enabling underreporting of moral viewpoints. I conclude that the persistent cultivation of such “silent complicity” provides a valid normative reason to narrow the scope of negative rights traditionally seen as implied by the ideal of scientific freedom, subordinating autonomy to the demands of deliberative responsibility.
Hanna Metzen (University of Konstanz): Science PR, communicative goals, and trust in science
Properly communicating science is crucial for trust in science and can guide decisions by laypersons as well as policymakers (Intemann 2023). I will focus on a type of science communication that is clearly impactful yet highly understudied by philosophers of science: science PR done by universities or other research organisations. Typically, this is carried out by professionally trained communicators working in press offices and communication departments that belong to the university administration or answer to the head of organisations and departments. They often have considerable influence in deciding which researchers as well as research results get featured in public debates and how this is framed. As such, science PR plays an important role in gatekeeping or curating scientific information (Winsberg 2026), however, when done wrong, it can distort public understanding of scientific issues and undermine trust in science.
In my talk, I will argue that science PR faces two particular challenges and make suggestions for dealing with them. My theoretical analysis draws on a case study from COVID-19 communication: In 2021, the University of Hamburg issued a press release on the origin of SARS-COV-2 that featured a non-peer-reviewed analysis by one of their physics professors. While the University’s press office maintained that is not their job to evaluate the accuracy of researchers’ claims, they were criticised for violating guidelines for good science PR (Weißschädel 2021).
The central challenge that I want to analyse concerns the communicative goals of science PR. Science PR is often part of publicly funded research organisations and in this sense committed to the public interest. Yet, science PR also follows strategic goals (Priest 2018), which includes both communication in the interest of science and communication in the interest of the specific scientific organisation. Especially for the latter goal, it is unclear how far this legitimately can go, for example in determining standards for featuring scientific results. Furthermore, the different communicative goals often conflict with each other. Navigating these conflicting goals can be a complicated task that requires assessments on a case-by-case basis.
However, another challenge arises from institutional background factors. The marketisation of science, organisational hierarchies and epistemic asymmetries between scientists and communicators can distort the proper balance between communicative goals. I will argue that because of this, good science PR requires ongoing reflections of their communicative goals, norms and standards, as well as a plurality of communicative actors who can scrutinise science PR communication. This includes scientists, science journalists, and other intermediaries.
My talk will contribute to recent debates on gatekeeping in science as well as the growing philosophical literature on science communication, especially attempts to expand the field beyond communication done by scientists themselves (e.g., Elliott 2023).
Charlotte Constanze Poller (Bergische Universität Wuppertal): Between Thickness and Comparability: Measurement Drivers in Global Well-Being Metrics
Global measurements of well-being shape how individuals, institutions, and governments understand and pursue the “good life”. Given persistent value disagreement about what well-being consists in, they face a core dilemma: How thick can measures be while remaining globally comparable? The World Happiness Report (WHR) is a paradigmatic case for examining how complex, value-laden concepts such as well-being are rendered measurable at a global scale. It publishes annual cross-national rankings based on survey data, positioning itself as an empirical basis for public-policy debate (Helliwell et al. 2025). To analyze the WHR’s measurement strategy, I draw on Basso and Alexandrova’s (2025) framework of measurement “drivers,” understood as competing epistemic, ethical, pragmatic, and metrological considerations that guide design choices and inevitably force trade-offs. I argue that in the WHR pragmatic and metrological drivers are systematically privileged over epistemic and ethical drivers, yielding a deliberately "thin" operationalization of well-being, and thereby constraining the range of value commitments that can be publicly articulated and contested as legitimate reasons in policy debates.
This “thinness” is particularly visible in the WHR’s core measurement of well-being, which relies on a single life-evaluation item asking respondents to place their current life on a 0–10 ladder. Presented as value-neutral and universally applicable, it facilitates global comparison, yet leaves unspecified which aspects of life respondents take to matter for their well-being. Although the WHR supplements its core life-evaluation measure with six explanatory factors, these function as predictors rather than constitutive dimensions, and thus do not resolve the underlying normative indeterminacy.
Recent work in the philosophy of social science emphasizes that policy-relevant indicators do not merely inform decisions but also shape how those decisions can be publicly justified and criticized. In this vein, Thoma (2024) argues that, under conditions of persistent value disagreement, policy-relevant indicators should not restrict the public articulation and contestation of competing value commitments. This requirement becomes especially demanding in the case of well-being. Well-being is a thick, context-dependent concept: different societies endorse different considerations as constitutive of a good life (Alexandrova 2017). Hence, global well-being measurement inevitably confronts disagreement over which aspects are constitutive of well-being. This generates a conceptual dilemma: adequate measurement requires some degree of thickness to avoid arbitrariness and to legitimize policy guidance, yet substantial value disagreement undermines the abstraction and standardization on which global comparison depends.
The WHR resolves this tension by effectively standardizing away substantively relevant differences in what well-being is taken to be. Drawing on Thoma’s (2024) argument that policy-relevant indicators should not restrict the public articulation of plural value commitments, I defend a minimum-thickness requirement for global well-being measurement. The core idea is not to replace thin global metrics with a comprehensive theory of well-being, but to rebalance competing measurement drivers so as to make at least some normative commitments explicit and contestable.
I suggest that this can be pursued by modestly enriching the WHR’s life-evaluation measure with a small set of explicitly normative dimensions, and by presenting these dimensions in a pluralistic, dashboard-like format. Together, these moves shift priority toward epistemic and ethical drivers in well-being measurement, at the expense of maximal comparability and communicative simplicity.
Lorenzo Sartori (The University of Sheffield): The Meagre Epistemic Value of Aesthetic Values in Science
In this paper, I critically assess three recent attempts to defend the claim that aesthetic values have genuine epistemic significance in scientific practice. More precisely, I examine what I call the Epistemic Worth of Aesthetic values in science thesis (EWA), according to which aesthetic considerations can rationally guide theory choice and justify trust in scientific products in a way that is not reducible to familiar epistemic or pragmatic standards. My analysis focuses on three independent, recent proposals, developed by Elgin (2020), Ivanova (2020), and Murphy (2023), which seek to rehabilitate the epistemic role of aesthetic values while avoiding earlier, well-known objections to aesthetic accounts of scientific rationality (e.g., Todd 2008).
I first clarify the normative character of EWA. The central issue is not whether scientists do in fact appeal to notions such as elegance, simplicity, harmony, or beauty, but whether such considerations can legitimately function as reasons for accepting, preferring, or trusting scientific theories, models, and practices. In order for EWA to constitute a non-trivial philosophical claim, I further suggest aesthetic considerations must not merely be traditional epistemic or practical virtues in disguise – e.g., explanatory power, unification, tractability, predictive success, or ease of application.
Against this background, I argue that the aforementioned defences of EWA fail to establish an epistemically normative role for the aesthetic. In the case of Elgin’s proposal, according to which aesthetic factors operate as “gatekeepers” of scientific acceptability, I show that the alleged aesthetic dimensions—symmetry, simplicity, systematicity and elegance—derive their normative force entirely from familiar epistemic and pragmatic considerations. Once these underlying considerations are made explicit, the aesthetic vocabulary becomes explanatorily idle.
I then turn to Ivanova’s attempt to ground the epistemic relevance of beauty in its contribution to scientific understanding, rather than to truth. I argue that this strategy either presupposes a factive conception of understanding, in which case the appeal to the aesthetic ultimately collapses back into standard epistemic reasoning, or relies on a non-factive conception of understanding whose defining features substantially overlap with the very aesthetic properties it is supposed to justify. As a result, the connection between aesthetic value and understanding is either insufficiently motivated or becomes circular.
Finally, I analyse Murphy’s proposal, which identifies the epistemic value of the aesthetic with the relation between form and content in scientific representations, especially in thought experiments. I contend that the form-content relation cannot plausibly be characterised as a distinctively aesthetic feature in scientific contexts, and that assessments of whether a given form successfully conveys a given content must, again, ultimately appeal to epistemic and pragmatic criteria such as intelligibility, relevance, and exemplificatory power.
Taken together, these analyses support a general conclusion: whenever aesthetic considerations appear to play a rational role in scientific practice, that role can be fully accounted for by underlying epistemic and practical factors. While aesthetic language may have descriptive or rhetorical significance in scientific discourse, it does not provide an independent source of epistemic normativity. Consequently, the epistemic contribution of aesthetic values in science, if any, remains meagre.
Mona-Marie Wandrey (University of Cambridge): Distributive epistemic justice for nonhuman animals
Assessing which nonhuman animals are conscious is morally significant, but fraught with uncertainty, due to the inherent subjectivity of conscious experience. Scientists rely on evidence-based methods to infer the likelihood of consciousness in nonhuman animals, for example by identifying behavioral and neural markers closely linked to consciousness (Andrews et al., 2025). Yet, determining when evidence becomes substantial enough to warrant attributions of consciousness inevitably involves non-epistemic value judgments about the risks of over- and underattributing consciousness (Birch, 2018). When it comes to the question of whose values should determine evidential standards, the presumed values of animals themselves are not usually considered in animal consciousness science. In this paper, I argue that setting evidential standards without representing animal interests constitutes a form of distributive epistemic injustice, or an unfair distribution of knowledge, that has not been theorized so far (Irzik & Kurtulmus, 2024).
To my knowledge, the framework of distributive epistemic injustice has only been applied to the human context. Yet, nonhuman animals are the ones most affected by evidential standards in animal consciousness science, as their very capacity to be recognized as bearers of interests often hinges on these standards. Adopting high evidential standards risks neglecting the interests of many animals by treating them as incapable of consciousness. This leaves animal advocates without the epistemic resources to make a case for the juster treatment of animals in public debate. Therefore, I argue that evidential standards in animal consciousness science should be responsive to the presumed values of animals themselves. In many cases, this entails setting evidential standards as low as possible while respecting epistemic values.
The argument proceeds as follows: First, I show that setting evidential standards in animal consciousness science is inevitably value-laden. Then, I introduce the debate on values in science to analyze whose values matter when setting evidential standards. I draw on political theories of animal rights to argue that evidential standards need to represent the values of animals themselves to achieve distributive epistemic justice. My argument is based on three premises: First, animals should be regarded as moral subjects whose interests matter for their own sake. Second, animals should be regarded as epistemic agents who communicate their needs and preferences in ways that are often marginalized in discussions of epistemic injustice. Third, animals should be regarded as political subjects whose interests matter when distributing resources, including scientific knowledge (Donaldson & Kymlicka, 2011). Recognizing that animal interests matter for distributive epistemic justice helps to move the debate around where evidential standards in animal consciousness science should lie beyond value disagreements between human members of society.
Antoni Antoszek (Jagiellonian University): The Entropic Self: A Stochastic-Selective Architecture for Artificial Agency
Current progress in artificial intelligence (AI) development poses natural questions regarding agency of artificial intelligence. Both the vast literature on agency and the growing number of texts on artificial agency seem to agree that agents must exhibit states such as “trying, wanting, perceiving” (Steward, 2012); it should “want certain things for itself” (Briegel & Müller, 2025) and in that sense “[be] able to act for reasons” (Dung, 2025). Moreover, we tend to require that an agent is “a settler of matters concerning certain of the movements of its own body”, and such settling “cannot be regarded merely as the inevitable consequences of what has gone before” (Steward, 2012; Floridi describes this point as “capacity for some self-initiated and directed action within environmental constraints”, see Floridi, 2025).
We argue that current neural networks, including Large Language Models (LLMs), regardless of their architectural depth, fail to meet the conjunction of those criteria.
To illustrate our argument, we propose a Buridan’s LLM thought experiment in which the model has an opportunity to demonstrate agency but cannot exhibit the capacity to resolve underdetermination in a way attributable to the system itself. We consider a Large Language Model with temperature set to zero that faces two equiprobable tokens during generation. We identify three possible scenarios (of which the first one is a theoretical possibility and the other two are being implemented in practice):
1. The model becomes paralyzed and ceases the generation process.
2. The model is programmed to use deterministic tie-breaker (see Welleck et al., 2024) such as choosing the token with the lowest ID.
3. The model is programmed to use an indeterministic procedure to select the token randomly (Minh et al., 2025).
We believe that neither of the scenarios can account for agency: while the first one annihilates the agent and the second one is fully deterministic (we share here the incompatibilist account of Steward and Briegel), the third one forces the model to use a coin-toss procedure rather than take ownership of the stochasticity. In this scenario, an agentic LLM recognizes the stalemate and generates a meta-goal ("I need to save time") to justify the random choice it is about to make. The Buridan case reveals that agency is not about stochasticity or determinism per se, but about ownership of the resolution of indeterminacy.
To resolve this, we introduce the concept of the Entropic Self: a dual-process architecture comprising a High-Temperature Proposer and a Low-Temperature Censor. The Proposer ensures that the system’s action space is not reducible to a static mapping from prior inputs. The Censor then evaluates candidates for action against a rigid, self-referential value function, collapsing the probability cloud into a single, intentional trajectory.
The resulting behavior is neither mechanically fixed by prior states nor reducible to an arbitrary random draw, but becomes a selected event. We conclude by requiring artificial agency to have the capacity to constrain its own entropy – utilizing noise as a fuel for novelty while maintaining the structural integrity of the self.
Daniel Grimmer (Yale University): Evolutionary Meta-Learning in Neural Networks as a\\ Neutral Testing Ground for Nativism and Empiricism
The recent engineering success of neural networks is often seen as favoring Empiricism over Nativism. Indeed, Empiricist accounts of epistemology (i.e., tabula rasa initialization followed by general-purpose learning from massive amounts of sensory data) align closely with the standard way of training neural networks. Crucially, however, this training regime imposes a broadly Empiricist epistemology by engineering fiat, rather than fairly testing it against a Nativist alternative. We argue that this issue can be remedied with a change of training regime; our Evolutionary Meta-Learning framework allows one to use neural networks to efficiently simulate the evolutionary pressures that shaped the human mind. We provide a first-principles derivation showing that the dynamics of a Darwinian Lineage Simulation (DLS) are formally equivalent to a noisy Stochastic Gradient Ascent (SGA) up a log-fitness landscape. Furthermore, we demonstrate that the Baldwin Effect—modeled here as an evolutionary pressure towards rapid learning so as to minimize childhood mortality—is well captured by the MAML++ meta-objective function. This framework offers a neutral testing ground for Nativist and Empiricist adaptation strategies while also aligning both approaches with Sutton's Bitter Lesson. Without hand-coding or brute stipulation, adaptations are chosen by the evolutionary process itself as it navigates its way up an ever-changing fitness landscape.
Katia Parshina (LMU, MCMP): Surveyability in mathematical discovery with neural networks
Neural networks (NN) are used for a great variety of tasks. Among other things, they are widely applied in mathematical research. The majority of NN-based tools being integrated into mathematical practice are based on an architecture called a transformer. Transformers can process complex tasks by selectively focusing on relevant parts of the data, making this class of neural networks a central focus of recent research and development.
Transformer models trained on an extensive amount of natural language data are called large language models (LLMs). The most direct way a mathematician can use a transformer model in their research is by interacting with LLM-based chatbots in natural language. However, the application of LLMs or other transformer models is not restricted to LLM-based conversational agents. Recently, AI-assisted mathematics has seen a significant advance as researchers started using transformers trained on symbolic mathematics or LLM-powered code-evolving systems to discover representative mathematical constructions (Georgiev 2025). The nature of integration of these systems into mathematical research is quite different from LLM chatbot conversations: they do not aim to improve a line of reasoning or verify an already formulated conjecture. The two most recently developed transformer-based discovery frameworks, PatternBoost (a methodology) and AlphaEvolve (a system combining several code-based LLMs), rather help by generating symbolic or code representations of mathematical objects and evolving these representations in the direction mathematicians are interested in. The final goal of the application of such systems is to be able to consider more mathematical constructions than would be feasible for a human mathematician alone.
Several epistemological challenges arise from integrating neural networks into mathematical research, and an extensive philosophical literature addresses these challenges. However, the classical discussions rooted in Tymoczko’s criteria of reliability and surveyability tend to focus on mathematical proofs that include a computer-based calculation as an inevitable structural step (Mainzer 2021). Transformer-based NNs are not used in mathematics in the same capacity: they assist primarily by generating and evolving mathematical objects, rather than verifying proofs or performing routine calculations. As a result, the epistemic questions they raise require a different critical focus. Much of the existing discussion of neural networks in mathematical practice concentrates either on LLMs used as conversational agents (Buzzard 2024; Tanswell & Berg 2025) or on their potential integration into formal theorem provers (Pantsar 2024). With the recent rapid development of transformer-based methods for mathematical discovery, it is worth asking whether these earlier concerns remain applicable, since they were initially aimed at different types of AI-based tools. I argue that most of the previously discussed philosophical concerns related to neural networks in mathematical practice are grounded in the unreliability of natural language. By contrast, the recently introduced transformer-based methods for mathematical discovery largely evade these issues, unless they are coupled with LLM-based conversational interfaces. Thus, the concerns about the trustworthiness of NN-based assistants primarily arise when transformer-based tools engage with a natural language component. When transformers are trained on code of symbolic mathematics, many previously raised epistemic worries can be evaded.
Molly Powell (Aarhus University): Norm Entrenchment Through Algorithmic Transparency and Radical Republicanism
Algorithmic transparency is widely seen as a way to promote fairness, accountability, and citizen trust in AI. Yet transparency does not simply inform; it also shapes citizen perceptions, constrains political imagination, and subtly reinforces particular normative orders. We explore how transparency can contribute to norm-objectification, the process by which specific values embedded in AI systems come to appear neutral, justified, and unavoidable, and thereby contribute in some instances to systems of domination.
Drawing on the example of FICO credit scoring from recent literature, we show how transparency mechanisms aimed at ‘helping’ users often individualize responsibility, guiding citizens on how to adapt to system rules, while obscuring collective avenues for resistance or reform. When framed as neutral or merely informative, such disclosures can normalize the underlying norms, e.g., the prioritization of debt repayment and individual responsibility for creditworthiness, and make them appear natural or legitimate, even when they are contested or unjust (Wang 2023).
While recent debates focus on whether algorithmic transparency is manipulative (Franke, 2022; Klenk, 2023; Wang, 2022, 2023), we argue that the deeper political concern is about legitimacy and domination. We suggest that many of these papers present at least some part of the story correctly, or point us in the right direction. While Wang’s account is not constructivist all the way down, we suggest that the underlying normative account would nonetheless benefit from further development, and that there exist accounts from analytic political philosophy that take power asymmetries and norm entrenchment seriously, and which could prove enlightening here. Similarly, Franke’s analyses of positive and negative liberty as relevantly analogous to types of manipulation ignores a further understanding of freedom which we will argue is more appropriate here, namely one from the republican tradition, where freedom is understood as non-domination. While manipulation may indeed be a part of what is going wrong in cases of norm objectification, it seems that the accounts offered are either overly broad, or otherwise still leave open the question of if and why manipulation is actually wrong in any given instance, thus kicking the normative can further down the road.
We begin by laying out descriptively what norm entrenchment through algorithmic transparency is, and how it operates in the case of FICO. We propose that three kinds of norms are entrenched, namely person-directed, system-directed, and recourse-directed. Although these concerns may (and indeed do) arise in other contexts, various features of contemporary algorithms do seem to intensify their effects or render them more troubling. We then suggest that a radical republican approach can explain what is troubling about norm entrenchment through algorithmic transparency (Thompson, 2018). This fits well with Wang’s explication of how power operates in this context, while giving a much more thoroughly developed account of what an unjust system looks like, and why and how we might have an obligation to reform it.
Henry Weatherburn (University of Aberdeen): From Radiological Protection to Artificial Intelligence: A Transferable Paired-Value Ethical Framework for Governing High-Stakes Technologies
This paper argues that the ethical governance of Artificial Intelligence (AI) need not begin from first principles, nor rely upon synthesis of fragmented guidelines, but, instead. can be grounded in an established and operationalised ethical framework drawn from radiological protection in medicine. Specifically, it advances the claim that the value architecture developed by the International Commission on Radiological Protection (ICRP), and elaborated in Publication 157, constitutes a philosophically coherent and practically tested model transferable to other high-stakes technological domains characterised by invisible risk, epistemic asymmetry, and probabilistic harm.
The analysis begins by tracing the philosophical genealogy of ICRP’s core and procedural value of: dignity; autonomy; beneficence; non-maleficence; justice; solidarity; prudence; precaution; transparency; accountability; and inclusiveness, situating them within plural ethical traditions spanning deontological, consequentialist, virtue ethical, and deliberative democratic thought. This pluralist grounding is presented not as theoretical eclecticism but as a strength enabling cross-cultural and interdisciplinary applicability.
The paper then examines the methodological innovation of ICRP Publication 157: the operational pairing of ethical values to guide professional judgement under conditions of uncertainty. Through reflection based on scenarios in medical imaging and radiotherapy, these pairings function as diagnostic tools for identifying ethical tensions in practice, moving ethics from aspirational declaration to structured reasoning.
Building on this foundation, the paper develops a methodological transition from radiological ethics to AI ethics. It argues that both domains confront structurally analogous challenges: invisible mechanisms of action; complex risk distributions; opacity to affected populations; and high consequence failure modes. The paired-value model is therefore applied analytically across AI deployment contexts, including clinical decision support, recruitment filtering, financial lending, predictive policing, educational surveillance, autonomous vehicles, robotic care, algorithmic trading, and autonomous weapons systems. In each case, the framework demonstrates capacity to interrogate issues of autonomy erosion, embedded bias, premature deployment, accountability gaps, and harms to dignity.
To move beyond conceptual analysis, the paper proposes a domain-independent Competency Framework for Ethics in AI. This translates value commitments into trainable professional capacities spanning ethical literacy, data stewardship; risk evaluation; stakeholder engagement; communication; accountability; reflexivity; governance awareness; inclusiveness; and lifelong ethical learning. The framework is designed for application across developers, deployers, regulators, and institutional oversight bodies.
The paper concludes with an implementation argument: that AI ethics must become operational rather than rhetorical. Radiological protection is presented as a historical precedent in which ethical governance successfully matured alongside technological power. By adopting its paired-value reasoning model and competency structures, AI governance can similarly anchor innovation within accountable moral architecture.
The central thesis is therefore demonstrative rather than analogical: the ethical reasoning system developed to manage invisible radiation risk is structurally and philosophically portable to the governance of invisible computational power.
Iwan Williams (University of Copenhagen): Doing without etiological functions in AI metasemantics
What do the internal states of AI models represent? Do activation patterns in large language models represent familiar entities like grapefruits and governments, or merely linguistic objects like words and syntactic structures? In specifying the conditions under which AI internal states bear particular representational contents, several philosophers appeal to etiological functions – functions deriving from a system’s causal history (Butlin 2023; Coelho Mollo & Millierè forthcoming; Goldstein & Levinstein 2024). In this paper, I argue that an alternative kind of function – one deriving from the intentions of designers, deployers and users – is better suited for this task.
On standard teleosemantic accounts of biological representation, content is fixed by the conditions that an internal state has the function of carrying information about, where functions derive from selection or stabilisation processes like natural selection or feedback-based learning (Millikan 1984; Garson 2016; Shea 2018). Applied to AI, advocates argue that machine learning constitutes such a function-conferring process: internal states selected during training for carrying information about external features thereby acquire the function of representing those features. However, the assumption of etiological functions in this context has not been sufficiently justified. I develop an account of “deployment functions” and show it better satisfies three desiderata for functions for AI metasemantics.
First, I define system-level deployment functions: an AI system has a deployment function to X if agents intentionally design, deploy, or use it to X. Thus, a chatbot deployed for answering factual questions has that deployment function regardless of whether it was trained on a next-token prediction task.
Second, I show how system-level functions confer functions on internal components. A challenge is that the specific internal workings of deep learning systems are not hand coded, but emerge through training, thus it is unclear how human intentions could confer component-level functions (Hurshman 2024). However, I argue intentions can ground component functions indirectly: a component has a deployment function F if it contributes to the containing system’s (intentionally derived) deployment function X by F-ing. This requires no designer beliefs or intentions about the component itself.
I show that deployment functions satisfy three desiderata: (i) normativity: components failing to make their characteristic contribution to the intended purpose of the system are malfunctioning, not merely playing a different causal role. (ii) misrepresentation: a component with the deployment function of carrying information about P, but failing to do so, thereby misrepresents. While etiological functions also satisfy these first two desiderata, deployment functions uniquely satisfy (iii) explanatory relevance: Interpretability researchers attribute representational contents to model-internals to serve pragmatic goals, e.g., predicting failure modes, mitigating dysfunctional behaviour, and building models that better serve our interests (Sharkey et al. 2025). These concerns are largely ahistorical and impose use-centred success criteria. By positing internal representations, researchers want to explain how components contribute to tasks we currently use models for, rather than roles stabilised through training. Deployment functions are thus more faithful to the explanatory goals of interpretability research, and should be preferred over etiological accounts.
Alex Broadbent (Durham University): Evidence and Models in Epidemiology: Towards Evidence-Based Modelling
During Covid-19, mathematical models rapidly moved from a supporting role in public health to an authoritative role in justifying unprecedented non-pharmaceutical interventions (NPIs). Yet prior to 2020, pandemic planning was strongly shaped by evidence-based policy frameworks that tended to treat mathematical modelling as uncertain, indirect, or low-quality evidence (Fuller 2020). In this talk we seek to explain how models became suddenly so prominent, and argue that they must adopt some of the rigor of the evidence-based tradition if they are to play a justificatory, and not merely exploratory, role in public health.
Following seminal publications in the mid-2000s (e.g.: Hollingsworth et al. 2006), mathematical modelling of infectious diseases developed rapidly, and was used to evaluate outcomes of possible interventions on pandemic influenza. However, pre-2020 pandemic plans are explicitly evidence based, and adopt an ambivalent attitude towards modelling. They advocate waiting for data and making proportionate responses, and they see the goal as mitigation rather than suppression. Influential figures even disapproved of mathematical modellers for speculating about stringent NPIs (Inglesby et al. 2006).
However, on 23 January 2020, Chinese authorities implemented stringent NPIs in Wuhan in response to Covid-19. Reported cases fell soon afterwards. A few weeks later, Italy showed the world what could happen if policy reacted too slowly. We argue that these episodes together functioned as a kind of crucial experiment. Many scientists (including Inglesby, cited above) became convinced of the effectiveness of NPIs, the viability of reversing rather than merely delaying growth, the necessity of doing so, and the value of the modelling approaches that had become slightly notorious for speculating about such matters. Suddenly, models were dominating public health policy, a domain where they had previously sought to gain a foothold.
Yet something was lost in this transfer of power from evidence to models: the rigorous evaluation of NPIs in Wuhan was forgotten. Drawing on our audit of the most cited papers – both modelling and empirical – on Wuhan, we argue that temporality was never subject to rigorous evaluation, despite being identified as an issue early on (Lipsitch et al. 2020). The literature did not seriously investigate the possibility that infections peaked before NPIs were introduced on 23 January. As a consequence, NPIs were widely perceived to be justified putative evidence – modelling results – which had not been subject to the kind of scrutiny usually required of evidence cited to justify policy.
We contend that the problem was not with modelling per se but with the methodological ethos around it. Mathematical disease modelling has operated in the context of discovery, and is ill-equipped for the context of justification. A number of philosophers have already argued that purpose-specification is central for modelling (Winsberg and Harvard 2024). We extend this line of argument, and propose a set of principles that enable epidemiological modelling to successfully transition from the context of discovery to the context of justification. We call this framework Evidence Based Modelling (EBMod).
Sam Fellowes (Lancaster University): Self-diagnosis in psychiatry, modelling and the Duhem-Quine thesis
The primary argument for the accuracy of self-diagnosis in psychiatry is lived experience. Someone who is autistic has lived experience of autism, so this gives them better knowledge of themselves and the diagnosis than an outside observer, including professional diagnosticians. For self-diagnosis to be generally accurate, we need also show that people without the condition, and therefore without lived experience of the condition, are able to accurately work out that they lack the condition.
I portray psychiatric diagnoses as models, whereby they do not fully reflect people since instances of the diagnosis do not have all attributes of the diagnosis and they have many relevant attributes which are not covered by the diagnosis. I also portray the act of diagnosing as one of modelling whereby an individual is modelled as an instance of the diagnosis and the diagnosis is given or rejected based upon level of fit. Self-diagnosing individuals need take data, their lived experience, and model that data as an instance of a diagnosis, then make a judgement about fit.
The Duhem-Quine problem posits that it is unclear where the problem lies when evidence goes against the hypothesis. The problem may be with data production, the modelling of the data (via intermediaries) to the theory or the theory being tested. Similarly, when someone who is self-diagnosing realises that they lack sufficient data to justify a self-diagnosis, they face a choice over where the problem lies. They might simply not have the condition, but it could also mean they have misinterpreted the data, misinterpreted modelling the data onto the diagnosis, or the diagnosis itself could be flawed.
This means that someone who does not have sufficient relevant data has multiple ways that they can make a reinterpretation which leads them to self-diagnose. Two factors lead to over-diagnosis. Firstly, advocates of self-diagnosis emphasise the importance of lived experience, which means that someone who falsely assumes they have lived experience of the condition has a significant inclination to make incorrect reinterpretations when faced with potentially problematic data. If the data does not fit the diagnosis then they can (falsely) appeal to lived experience to reinterpret the data or remodel the data to the diagnosis to make the data fit the diagnosis. Secondly, advocates of self-diagnosis emphasise the acceptability of self-diagnosing using alternative symptoms and diagnostic criteria than that used by official diagnosticians because those psychiatrists who formulated those symptoms and diagnostic criteria lacked lived experience whereas people who self-diagnose (are presumed to) have lived experience. Someone can (falsely) appeal to lived experience if the data does not fit the diagnosis to assume that their alternative notion of the diagnosis is better than the notion employed by psychiatrists, making the data fit the diagnosis.
Someone who lacks lived experience of a condition but suspects they have lived experience of the condition can appeal to lived experience as justification for incorrectly, via a Duhem-Quine process, making a reinterpretation that makes the data fit the diagnosis. This shows the major risk of inaccurate self-diagnosis occurring.
Matthieu Fontaine & Cristina Barés Gómez (University of Seville): Rethinking Mechanisms: An Inferential Approach
Philosophical discussions of mechanisms in medicine and sciences often treat them as a distinct kind of evidence, complementing probabilistic and statistical correlations. (Russo & Williamson 2007; Thagard 1998). Yet, it remains unclear what, if anything, is epistemically distinctive about so-called “mechanistic evidence”. (Illari 2011) This paper proposes an alternative approach grounded in an inferentialist account of scientific reasoning: mechanistic hypotheses are best understood in terms of the inferential role they play within inquiry, in which different kinds of reasoning are intertwined (i.e. abduction, deduction, induction). In other words, mechanisms are studied from the perspective of the articulation of patterns of commitment and entitlement involved in a more general inferential framework.
On this view, mechanisms are not solely assessed by their direct confirmability, not do they introduce a sui generis evidential category. Instead, they function as nodes in networks of inference that articulate patterns of commitment and entitlement. To endorse a mechanistic hypothesis is to undertake specific inferential commitments – licensing certain explanatory, predictive, and interventional inferences – while also constraining which further inferences are defeasible, warranted, or excluded. In this sense, mechanisms structure bio-medical reasoning not by adding evidence, but by reorganizing the inferential space in which hypotheses are generated, evaluated, and acted upon.
The paper develops this claim by integrating a Peircean account of abduction with an inferential account of justification. (Peirce 1931/1958) Abductive inferences to mechanisms are understood as moves that introduce new inferential commitment in response to explanatory breakdowns, without thereby providing evidential support for their conclusion. Conversely, inferences from mechanisms are analysed as transitions licensed by previously accepted inferential roles, even in the absence of independent confirmation of the mechanistic hypothesis itself. What is often called “mechanistic evidence” is thus reinterpreted as the normative force of inferential transitions enabled by mechanistic commitments.
This framework allows for a principled distinction between inductive confirmation, which supplies factual evidence in the traditional sense, and inferential articulation, which determines how hypotheses may legitimately be used in reasoning and action. It also explains how mechanistic reasoning can be action-guided and indispensable in scientific practice without presupposing a realist ontology of mechanism, or inflating their epistemic status. By recasting mechanisms in inferential terms, we offer a novel contribution to debates on abduction, causality, and evidence. More generally, we suggest that understanding reasoning requires attention from evidential hierarchies to the normative structure of inference itself.
Liyuan Jiao (Durham University): Gender, Intervention, and Causal Identification in Clinical Judgment
Intervention-based approaches to causal inference, such as the Rubin Causal Model (RCM), ground causal claims in well-defined treatments and counterfactual contrasts (Rubin 1974). In medical and social research, investigators frequently appeal to socially structured factors such as gender to explain systematic disparities in evaluation and healthcare delivery. The difficulty is not that gender lacks causal significance, but that it does not readily correspond to a well-defined intervention in the sense required by the RCM. Whereas drug administration can be clearly represented as a discrete treatment variable, construing the same individual as counterfactually assigned a different gender does not yield a comparably well-defined treatment contrast. This tension generates a foundational methodological question: how can causal claims about gender be formulated and justified within an intervention-based framework?
This difficulty has motivated influential arguments, most notably by Holland (1986), that gender cannot function as a legitimate treatment variable within the RCM. At the same time, persistent gender-based disparities in medicine demand causal explanation. A prominent empirical strategy responds by intervening not on gender itself, but on gender cues such as names, pronouns, and other perceptual markers under controlled conditions (Goldin and Rouse 2000; Moss-Racusin et al. 2012). In clinical studies using standardized or virtual patient descriptions, patients presenting identical symptoms receive systematically different evaluations and treatment recommendations depending on perceived gender cues (Stutts et al., 2010). By holding symptom presentation fixed while varying perceived gender, such studies abstract from biological differences and isolate the role of gendered interpretation in diagnostic reasoning.
The paper argues for three claims. First, in many diagnostic contexts, gender is causally operative as a socially structured category through a network of mediated mechanisms, one of which is perceived gender in diagnostic categorization. Clinical judgment often proceeds through processes of categorization and stereotype-mediated evaluation, in which patients classified as women may be regarded as more emotional or less credible reporters of pain. Second, cue-based interventions manipulate this causally operative node, namely perceived gender, thereby identifying a path-specific causal effect within a broader network of gender-related mechanisms. Such interventions do not replace gender with perceived gender; rather, they identify a mechanism through which gender exerts causal influence in diagnostic reasoning. Third, extrapolation from cue experiments to claims about gender’s causal relevance is justified only if the experimentally isolated categorization mechanism also operates in ordinary clinical practice and is not systematically counteracted.
Under these conditions, cue-based evidence supports qualitative claims about gender, specifically that gender is positively causally relevant to diagnostic evaluation, without requiring that gender itself be represented as a directly manipulable treatment within the RCM. The analysis clarifies the scope and limits of cue-based evidence and shows how intervention-based frameworks can accommodate causal claims about socially structured factors that resist direct manipulation by identifying the mechanisms through which they operate.
Jolie Zhou (University of Cambridge): What Counts as Health? An Integrative Functioning–Embodiment Framework
This paper aims to establish a new framework for assessing health and examine its normative significance particularly in biotechnological contexts. For example, should a person be considered healthier with a computerised prosthesis or with maximally restored biology? Existing debates capture important aspects of health but offer no unified framework for such comparisons.
I propose a two-dimensional framework: Integrative Functioning and Embodiment (IFE).
For naturalism, health-related functioning often refers to the contribution of internal bodily parts or processes to individual survival or reproduction (e.g. Boorse 1977). Inspired by pluralistic naturalism (Lewens 2015), Integrative Functioning extends beyond the biological to include technological and environmental supports. Unlike existing naturalist accounts, however, it excludes reproductive functioning as essential to health and therefore concerns only mechanisms that sustain individual survival.
Embodiment concerns the experience of being and having a body. A healthy body is typically 'transparent', whereas illness renders it hyper-present or alienating. Technologies can likewise become transparent once incorporated into one's body schema (Merleau-Ponty 2012). Meanwhile, perceptual factors, such as the feel of a prosthetic limb, and social roles or prejudice can heighten bodily hyper-presence: blindness may be more salient for a painter, and hostile stares may make a wheelchair user acutely aware of their body and equipment.
Health obtains when neither Integrative Functioning nor Embodiment falls significantly below the average; ill health is a holistic condition rather than a biological trait per se.
Three key elements within IFE:
(1) Reproductive functioning may affect Embodiment but not Integrative Functioning. Even if reproduction is regarded as an organism's ultimate goal, this does not entail balancing reproduction and survival when judging health: treating species-typical conflicts between survival and reproduction as 'healthy' risks the absurd conclusion that death is healthy, as illustrated by semelparous North American salmonids.
Furthermore, health is an individual good, while reproduction is not necessarily an individual goal. Reduced reproductive functioning does not necessarily diminish health: a person with infertility may remain healthy in an inclusive society. Nor does high reproductive capacity guarantee health; for example, it is compromised when the pregnant body is stigmatised.
(2) Social prejudice (SP), including internalised SP (self-blame), damages Embodiment and thus health, and calls for structural redress. Impaired functioning and embodied responses (e.g. discomfort with technological intrusion) can worsen prospects and might reinforce SP. This does not justify SP; rather, it implies that with adequate technological support, atypical bodies can sustain creative forms of life that serve as ‘counter-stereotypical exemplars’ (Johnson 2024), helping erode unjust norms alongside other anti-prejudice strategies.
(3) Both dimensions are measured against statistical typicality. Severe deficits in Integrative Functioning are unhealthy even if embodied affirmatively (e.g. embraced life-threatening pregnancies), while profound disembodiment is unhealthy despite normal functioning (e.g. Body Integrity Identity Disorder). Gains in one dimension may follow alterations in the other; prosthetic amputation can improve well-being in BIID (Blom, Hennekam, and Denys 2012). This implies that IFE supports plural approaches to health enhancement, rather than privileging the restoration of original biological function.