Science begins with questions, but understanding requires more than answers. Explore how evidence, reasoning, and uncertainty shape what we can know about the world.
By Niraj Kumar (Founder of NOESIS Lab) · August 20, 2026
Imagine opening your phone tomorrow morning and seeing this headline “New study finds that Supplement X reduces the risk of heart disease.” It sounds reassuring. Perhaps even convincing. But should you believe it? Before accepting the claim, a scientist would begin asking a series of questions: How many people were studied? Was there a comparison group? Were the participants randomly assigned? How large was the effect? Could something else explain the result? Has another research group found the same thing? Is there a plausible biological mechanism? And how certain are we that the observed effect is real? These questions point toward something deeper than the result of a single experiment. They lead to a more fundamental question: How do we know what we know? Science is sometimes presented as a neat sequence:
question → hypothesis → experiment → conclusion
That sequence is useful, but real science is rarely so tidy. It is better understood as a continuing process in which ideas are confronted with evidence, challenged by alternative explanations, revised when necessary, and tested again. As the UC Berkeley Understanding Science project emphasizes, treating science as a simple linear recipe oversimplifies how scientific investigation actually works. The important thing is not simply that scientists collect facts, but that they use evidence to distinguish between competing explanations. The important thing is not simply that scientists collect facts. It is that they use evidence to distinguish between competing explanations.
The process of reasoning
(In practice, the flow can loop: a conclusion often leads to new questions, refining or replacing the initial hypothesis)
Suppose you notice that plants growing in one part of a garden are taller than plants growing somewhere else. That is an observation. You might measure their heights and find that the difference is real. You now have data. But the data do not automatically tell you why the plants are different. Perhaps one area receives more sunlight. Perhaps the soil contains more nitrogen. Perhaps it retains more water. Perhaps the plants belong to different varieties. The observation is: Plants in location A are taller than plants in location B. An explanation might be: Greater sunlight causes increased plant growth. Those are not the same statement. This distinction is at the heart of scientific reasoning. An observation describes something we detect or measure. A hypothesis proposes an explanation for it. And a scientific hypothesis becomes especially useful when it makes a prediction that could turn out to be wrong. That is where science begins to become a test rather than a story.
Imagine a scientist proposes a hypothesis that plants grow faster when they receive more sunlight. The hypothesis itself is not yet evidence. But it leads to a prediction that if sunlight promotes growth, then plants receiving more light should show greater growth than comparable plants receiving less light. Now the idea has become testable. Scientists can design an experiment, measure growth, compare groups, and ask whether the results match the prediction.
This distinction between hypothesis and prediction is small but powerful. A hypothesis attempts to answer why might this happen? A prediction asks if that explanation is correct, what should we observe?
If the prediction fails, the scientist has learned something important. The result may indicate that the hypothesis needs to be modified or rejected, but it may also point to problems with the experimental design, the assumptions behind the prediction, or other factors that were not considered. If the prediction succeeds, the hypothesis gains support, but it has not been proven to be true. Further tests are needed to determine whether the explanation continues to hold under different conditions. This is one of the important characteristics of scientific knowledge.
Evidence can strengthen an explanation without turning it into absolute certainty.
Key Points:
Separate what we observe from what we think it means. Data tell us what happened; interpretation attempts to explain why it happened. Keeping the two separate helps prevent assumptions from being mistaken for evidence.
A hypothesis explains; a prediction puts the explanation at risk. A hypothesis proposes why something happens. A prediction states what we should observe if that explanation is correct. The prediction gives us something specific to test.
Correlation can reveal a pattern, but not necessarily its cause. When two things change together, look for alternative explanations, hidden variables, and other evidence before concluding that one caused the other.
Good comparisons make causal claims stronger. Controls, randomization, and careful experimental design help researchers distinguish the effect of an intervention from effects caused by other differences between groups.
One result is evidence, not the final answer. Confidence grows when findings survive independent studies, different methods, larger datasets, and repeated testing. A result that cannot be reproduced deserves closer examination.
Strong scientific conclusions often come from converging evidence. Statistical patterns can show that a relationship exists; experiments, observations, and mechanistic studies can help explain how or why it occurs. Different lines of evidence can reinforce one another.
Uncertainty is part of scientific knowledge. Scientists rarely deal in absolute certainty. The strength of a claim depends on the quality, quantity, consistency, and limitations of the available evidence.
Scientific knowledge remains open to revision. A conclusion is not made stronger by being declared certain; it becomes stronger when it continues to withstand attempts to test, challenge, and refine it.
One of the most famous examples began not with a perfectly planned experiment, but with something that looked like contamination. In 1928, Alexander Fleming noticed that a mold had appeared on a bacterial culture and that bacteria were not growing around it. The observation was intriguing. Where the mold was present, bacterial growth was inhibited. But Fleming still needed an explanation. Perhaps the mold was producing a substance that could kill or inhibit bacteria. That idea led to further investigation. Years later, Howard Florey, Ernst Chain, and their colleagues developed methods for working with purified penicillin and tested its antibacterial effects, including experiments in infected mice. Their work helped establish penicillin as a powerful antibacterial treatment. The scientific story therefore was not simply that Fleming discovered penicillin and antibiotics were born. It was a longer process that began with an unexpected observation, followed by a possible explanation, further testing, purification, animal experiments, clinical investigation, and eventually medical use.
A chance observation opened the door. Evidence had to keep walking through it.
Now consider another common scientific trap. Imagine that students who spend more time on their phones tend to get lower grades. At first glance, the two variables appear to be related, and statistically they may indeed be correlated. But it would be a mistake to conclude that spending more time on a phone directly causes lower academic performance. Other factors may be involved. Students who are struggling academically may use their phones more as a way to relax or avoid difficult work. They may also have less time available for studying because of differences in workload, family responsibilities, sleep, or access to educational resources. This is the difference between correlation and causation. Correlation tells us that two variables are associated, whereas causation means that a change in one variable actually produces a change in the other. The distinction is important because the world is full of factors that can influence what we observe. People who exercise more may also have different diets, people who take a particular supplement may already be more health conscious, and people who live near polluted roads may differ from those living elsewhere in terms of income, occupation, housing, or access to healthcare. A statistical association can therefore be genuine without representing a direct causal relationship. The scientific challenge is to determine whether the relationship remains when plausible alternative explanations are considered and, whenever possible, to find evidence that can distinguish cause from coincidence.
One of the most effective ways to investigate causation is to compare what happens when an intervention is introduced with what happens when it is not. This is the basic logic of a controlled experiment. Suppose researchers want to know whether a new medicine reduces blood pressure. They might divide participants into groups, with one group receiving the treatment and another receiving a placebo or an appropriate control. If the groups are sufficiently comparable and the study is properly designed, differences in their outcomes can provide evidence about the effect of the treatment. The basic principle is straightforward. Researchers change one important factor while keeping other relevant conditions as similar as possible, and then observe what happens. Randomization can make this comparison more reliable by reducing systematic differences between groups. This is why randomized controlled trials are an important tool for evaluating causal effects in many areas of medicine.
However, experiments are not always possible. We cannot randomly assign people to smoke cigarettes for twenty years simply to determine whether smoking causes lung cancer. For questions like this, scientists have to bring together different forms of evidence. This leads to one of the most important ideas in science.
A strong conclusion rarely rests on one kind of evidence alone.
The history of smoking and lung cancer provides a useful example. Researchers observed that smokers developed lung cancer at much higher rates than non-smokers, producing a striking pattern in epidemiological studies. But the statistical relationship alone left an important question. Could something else be responsible? Scientists therefore examined the relationship from several different directions. They looked at what happened in different populations, whether the risk changed with the amount of smoking, whether it changed with the duration of exposure, and whether people who stopped smoking experienced different risks. They also asked whether there were biological reasons to expect tobacco smoke to cause cancer. Over time, evidence accumulated from epidemiology, toxicology, pathology, and molecular biology. Today, the causal relationship is firmly established. The World Health Organization identifies tobacco as a major cause of cancer and many other diseases, while the International Agency for Research on Cancer reports that tobacco smoking is responsible for approximately 85% of lung cancer cases.
The lesson is not simply that smoking is associated with cancer. It is that multiple independent lines of evidence converge on the same explanation. That convergence is one of the foundations of scientific confidence.
Scientific evidence is not one-dimensional. Different kinds of evidence can answer different questions, and some become much more informative when considered together.
Statistical evidence can show whether an association is large, consistent, or unlikely to be explained by random variation alone. Researchers may report effect sizes, uncertainty intervals, probability values, or other statistical measures to describe what the data reveal. But a statistical result does not automatically tell us what is happening underneath the pattern. That is where mechanistic evidence becomes important. Mechanistic evidence asks a different question. How could this happen? In the case of smoking, biological research has identified carcinogenic components of tobacco smoke and mechanisms through which exposure can damage cells and contribute to cancer. Statistical evidence can show us that a relationship exists, while mechanistic evidence can help explain how that relationship could occur. Neither type of evidence has to stand alone. When different forms of evidence point toward the same explanation, they can provide a much stronger scientific case.
But even when several lines of evidence point in the same direction, scientists still have to ask whether the finding holds beyond the original study. A convincing result from one research team is important, but science becomes more confident when other researchers can examine the same question and arrive at similar findings. This is where replication and reproducibility become part of the story.
The experiment is not the end of the story
Imagine that a research team conducts an experiment and obtains an exciting result. The researchers publish their findings, but should we immediately consider the question settled? Not necessarily. A scientific result becomes more convincing when it can withstand attempts to check it, repeat it, or reproduce it under appropriate conditions. Science therefore asks another important question. Can someone else obtain a similar result?
This is where replication and reproducibility become important. The National Academies distinguishes between these two concepts. In its 2019 report, reproducibility refers, in the computational context discussed in the report, to obtaining consistent results using the same data, computational steps, methods, code, and analytical conditions. Replicability refers to obtaining consistent results across different studies that address the same scientific question using new data.
The distinction matters because an independent confirmation can increase confidence in a finding. If a result repeatedly fails to appear, scientists need to investigate why. Perhaps the original result was mistaken. Perhaps the effect occurs only under particular conditions. Perhaps the phenomenon is more variable than expected. It is also possible that the second study used a different method that produced a different outcome. A failed replication therefore does not automatically mean that the original researchers were wrong. It becomes another piece of evidence that must be considered alongside everything else.
The National Academies emphasizes that scientific claims should be evaluated in the context of the whole body of evidence, rather than being judged solely by one study or one replication attempt. And this brings us to another important point. Even when a finding has been tested repeatedly and supported by several lines of evidence, science does not suddenly remove every remaining uncertainty. More evidence can increase confidence, but it does not make uncertainty disappear.
Why science does not promise certainty
This may be one of the most misunderstood features of science. When people hear scientists say that there is uncertainty, they sometimes interpret it as meaning that scientists do not really know anything. In many cases, the opposite is true. Uncertainty is a way of describing the limits of what the available evidence can establish.
Every measurement has limits. Every instrument has finite precision. Every biological system contains variation. Every model simplifies some part of reality. Every experiment samples only a portion of the world. Good science does not hide these limitations. It tries to measure and communicate them as clearly as possible. A study might therefore report an estimated effect together with an uncertainty interval rather than presenting a single number as though it were exact. There is also an important point to remember about 95% confidence intervals. A 95% confidence interval does not mean that there is a 95% probability that the particular true value lies within that interval. In the frequentist framework, it means that the method used to construct the interval would capture the true parameter in about 95% of repeated samples, assuming the underlying statistical model and its assumptions are appropriate. That may sound less intuitive than saying that researchers are 95% certain, but it is scientifically more precise. The National Academies similarly emphasizes that scientific results carry uncertainty and that researchers should identify, characterize, and communicate relevant sources of uncertainty.
Uncertainty is not the enemy of knowledge. It is part of honest knowledge.
But uncertainty does not mean that every scientific claim is equally uncertain. When evidence accumulates, confidence can become very strong, especially when different kinds of evidence converge and when a prediction can be tested against an observation that was not simply expected, but specifically predicted in advance. This is perhaps one of the clearest ways to see how scientific knowledge becomes powerful. Sometimes scientists begin with something they can observe and work toward an explanation. At other times, they begin with an explanation and ask nature to provide the evidence. When a theory makes a precise prediction and that prediction is later observed, the relationship between theory and evidence becomes especially striking.
A prediction written in the language of the universe
Sometimes scientists do not begin with an observation. They begin with a theory that makes a prediction so specific that nature itself can provide the test. Consider gravitational waves. Albert Einstein's general theory of relativity predicted that accelerating massive objects could produce ripples in spacetime. For decades, these ripples remained undetected directly. Then, on September 14, 2015, the two LIGO detectors in the United States recorded a signal that matched the predicted pattern of gravitational waves produced by the merger of two black holes. The event was announced in 2016 as the first direct detection of gravitational waves and the first observation of a binary black hole merger.
The importance of the discovery was not simply that an instrument produced an unusual signal. The signal had a predicted form, it was observed by two separated detectors, and it was consistent with the theoretical waveform expected from a black hole merger. The observation also survived detailed analysis. Here, theory and observation met. A prediction made a century earlier became something the universe itself could answer.
The gravitational-wave discovery is an extraordinary example, but the underlying reasoning is not unique to astronomy. The same process can be seen in much more familiar scientific stories. Penicillin began with an unexpected observation, the connection between smoking and lung cancer emerged from patterns in human populations, and gravitational waves began as a prediction from a theory of gravity. The questions and methods were different, but in each case scientists had to connect evidence with explanations and determine whether the evidence was strong enough to support a particular conclusion. Looking at these stories side by side makes the common logic of scientific inquiry easier to see.
Three stories, one scientific logic
The examples may look completely different. Penicillin began with an unexpected observation. The connection between smoking and lung cancer emerged from epidemiological patterns followed by decades of investigation. Gravitational waves began as a theoretical prediction and were eventually detected using extraordinarily sensitive instruments. Yet the underlying logic is remarkably similar. Scientists ask what happened, consider possible explanations, and work out what they should expect to observe if those explanations are correct. They then look for ways to test those predictions, consider whether something else could explain the result, and ask whether the finding can be reproduced or replicated. They also examine whether there is a plausible mechanism connecting cause and effect and how much uncertainty surrounds the conclusion. As new evidence arrives, an explanation may survive, change, or fail. This is why science is much more than a collection of facts. It is a method for deciding how much confidence a claim deserves. But confidence in science rarely comes from a single experiment or a single moment of discovery. It develops over time as findings are examined, challenged, refined, and connected with new evidence. What begins as a result from one study can gradually become part of a much larger body of knowledge.
Knowledge is built, not announced
A scientific paper can be published in a single day, but scientific knowledge usually takes much longer to build. One study provides a result. Another group examines it. A third investigates the same question in a different population. Someone tests the proposed mechanism, while someone else may find an exception or a result that does not fit the original explanation. A meta-analysis can bring together findings from many studies, while new instruments can produce more precise measurements. As this evidence accumulates, theories can be refined and explanations that once seemed convincing may eventually become incomplete. This is why scientific knowledge is rarely the product of one experiment or one researcher. It is built gradually through a continuing process of testing and correction. Sometimes the evidence strengthens an existing explanation. Sometimes it reveals that an explanation needs to be changed. And sometimes it shows that an idea does not hold up at all.
The process is not perfectly linear, and it is not immune to mistakes. Scientists can be biased. Experiments can be poorly designed. Measurements can be wrong. Statistical analyses can mislead. Published findings can sometimes fail to replicate. But these weaknesses do not make science useless. They are precisely why science has developed practices such as controls, blinding, randomization, statistical analysis, replication, peer criticism, data sharing, and transparent reporting of uncertainty. The goal is not to make scientists incapable of being wrong. The goal is to create a system in which wrong ideas can be challenged, evidence can be examined, and explanations can be corrected.
That is what allows scientific knowledge to grow. It does not grow because every scientific claim is immediately certain. It grows because claims are continually tested against reality.
So, how do we know what we know?
We know because some explanations survive confrontation with reality better than their alternatives. We observe, we ask questions, we propose explanations, we make predictions, and we test them. We compare competing explanations, look for alternative causes, repeat important findings, measure uncertainty, and examine the mechanisms that could connect cause and effect. Most importantly, when new evidence arrives, we remain willing to change our minds.
That last part may be the most important. Science does not say that something is true simply because a scientist said so. Nor does it say that anything could be true. Instead, science asks a more demanding question. What evidence would make this explanation more credible, and what evidence would force us to reconsider it? That is the logic behind scientific evidence. Perhaps that is also the deeper lesson. Scientific knowledge is not a monument built from certainty. It is a structure continually tested against the world. The stronger the evidence, the more confidence we can place in it. When the evidence changes, the structure can change too. That is not a weakness of science. That is how science learns.
Science does not end with an answer. It continues with the next question.
Which part of scientific reasoning do you find most difficult to understand? Is there a scientific claim you have encountered that made you wonder how we actually know it to be true? Share your thoughts, questions, or examples below. Your question may become the starting point for another exploration.