Toward a Philosophy of Scientific Evaluation
Why the institutions responsible for selecting science often optimize for everything except scientific quality—and a proposal for a new system of scientific criticism.
One of the more surprising realizations I had during my PhD was that the scientific community has built extraordinary institutions for producing knowledge, but comparatively weak ones for evaluating its quality. Every day, journals decide which papers to publish, funding agencies allocate billions of dollars in research support, hiring committees shape the next generation of scientists, and investors determine which discoveries are worth commercializing. Few professions devote more time to evaluating ideas.
Yet despite the enormous influence of these decisions, much of scientific evaluation still relies on tradition, intuition, and imperfect proxies rather than an explicit philosophy of what constitutes good science.
Ask a room full of scientists what they value in a paper and the answers are reassuringly consistent: originality, rigor, careful interpretation, reproducibility, and a meaningful advance in knowledge. At the level of principle, there is broad consensus.
The difficulty begins when those principles have to be translated into decisions. Which grant deserves funding? Which paper belongs in Nature? Which applicant should receive the faculty position? Which company is worth investing in? These are not questions with objective answers; they are judgments made under uncertainty, using imperfect measures of quality.
Why, then, do our decisions so often diverge from the ideals we claim to value?
I suspect the answer lies in incentives.
Economists have long understood that incentives shape behaviour. Scientists are no exception. If we want to understand why research looks the way it does, we should spend less time asking whether scientists are making the “right” decisions and more time asking what, exactly, the system rewards.
We reward publication in prestigious journals, so scientists learn to write papers that appeal to prestigious journals. We reward citation counts, so researchers gravitate toward questions likely to attract attention. We reward novelty, so incremental but essential work often struggles to find a home. Funding agencies increasingly reward translational potential, so even basic science proposals promise clinical relevance long before such claims can reasonably be justified.
None of this requires bad actors. Rational people respond to rational incentives. If we want different scientific behaviour, we need to think more carefully about the incentive structures we create.
The publishing ecosystem illustrates this particularly well.
The most prestigious journals were never designed simply to publish the most rigorous experiments. Journals such as Nature, Science, and Cell built their reputations by identifying discoveries that were surprising, broadly relevant, and capable of changing conversations across disciplines. Their editors are not simply evaluating methodological quality; they are making editorial judgments about what readers will find compelling and what will attract attention. Those are entirely legitimate objectives, but they are not identical to identifying the highest-quality science.
The papers that transform science are often both exciting and rigorous. Yet history is equally full of influential discoveries that failed to replicate, alongside more humble studies that gradually became foundational. Editorial decisions inevitably balance scientific merit against novelty, breadth of appeal, and perceived importance. That is not a criticism of elite journals; it is simply an acknowledgment that they optimize for multiple objectives, only one of which is scientific quality.
The same tension exists throughout academia. Grant panels predict which projects are most likely to succeed before the experiments have even been performed. Hiring committees select scientists not only for the quality of their work, but for the value they are expected to create for the institution itself—financially, strategically, and socially. Venture capital firms evaluate biotechnology companies according to their expected financial upside. Every institution responsible for selecting science is optimizing for many things simultaneously, and scientific quality is only one of them.
If science is serious about improving the quality of its discoveries, it should devote as much creativity to building institutions that evaluate knowledge as it has to building institutions that produce it.
One of the most interesting recent attempts comes from QED Science. Their AI-based framework evaluates life science manuscripts along two independent dimensions: originality, defined as how far a study advances what the field already knows, and validity, defined as how well the evidence supports the conclusions being drawn.
I find this framework compelling—not because I believe artificial intelligence can completely replace scientific judgment; I still believe that meaningful evaluation of science will always require human judgment and scientific taste, much as criticism in art, literature, or film ultimately depends on human discernment rather than algorithms. Rather, I find it compelling because it provides a structured way of thinking about questions that scientists have been asking all along.
Did we learn something genuinely new?
Should we believe it?
Originality asks whether a study introduces an idea, observation, or mechanism that meaningfully extends what was previously known. Does it ask a genuinely new question, uncover an unexpected phenomenon, or offer a novel explanation? Originality is about the intellectual contribution a paper makes at the moment it enters the scientific record.
Validity asks a different question. Assuming the finding is novel, do the data actually support the conclusions? How well do those conclusions fit within our current understanding? If they challenge that understanding—as many important discoveries eventually do—is the evidence sufficiently strong to justify doing so? This is where scientific rigor lives. Experimental design, appropriate controls, statistical analysis, reproducibility, consideration of alternative explanations, and methodological transparency are not independent virtues—they are all different ways of evaluating the same underlying question: how much confidence should we place in these conclusions?
Viewed this way, many of the criteria scientists routinely discuss are not separate dimensions of scientific quality but different components of validity. They help us distinguish between a compelling idea and a compelling demonstration.
The final dimension is impact.
Unlike originality and validity, impact cannot be meaningfully assessed at the moment of publication. It is a judgment made in retrospect, emerging over years—or often decades—as a discovery is challenged, replicated, extended, and ultimately influences science or society.
Scientific impact asks whether a discovery fundamentally deepened our understanding of biology. Did it reshape how we think about the principles of life? Did it establish a new paradigm or become foundational for future discoveries?
Societal impact asks a different question. Did the discovery meaningfully improve human lives? Did it change clinical practice, enable transformative technologies, influence public policy, or solve an important real-world problem?
There is a reason the Nobel Prizes are typically awarded decades after the original discoveries. Their purpose is not simply to recognize outstanding papers, but to recognize discoveries whose influence has stood the test of time. That judgment requires hindsight.
I think of these three dimensions as operating on different timescales.
Originality is potential.
At the moment a paper is published, we can ask whether it contributes a genuinely new idea. Originality reflects the potential of a discovery to advance a field, not whether that potential will ultimately be realized.
Validity is credibility.
Validity asks whether the available evidence justifies believing the claims being made. Unlike originality, however, validity is not fixed. As new evidence accumulates—through replication, contradictory findings, improved methodology, or deeper mechanistic understanding—our confidence can increase or decrease. Scientific knowledge is provisional by design, and our assessment of validity should evolve alongside it.
Impact is legacy.
Impact asks whether the discovery ultimately mattered. Did it fundamentally change our understanding of biology? Did it meaningfully improve human lives? Unlike originality and validity, impact cannot be judged contemporaneously. It can only be assessed with the benefit of hindsight, once history has had an opportunity to reveal the true influence of a discovery.
Even if we agreed on how to evaluate science, we would still face another problem: creating conditions where scientific criticism can be expressed openly.
Modern research is an intensely social enterprise. We collaborate across institutions, review one another’s grants, write recommendation letters, serve on hiring committees, and depend on professional relationships throughout our careers. These connections strengthen science, but they also make criticism more complicated than we often acknowledge. The person whose paper you reject today may review your grant tomorrow. The investigator whose conclusions you publicly challenge may later evaluate your promotion package.
In such an environment, criticism acquires a cost. And whenever a behaviour becomes costly, it becomes less common. Diplomacy becomes the rational strategy.
Science advances through criticism, yet our incentive structures often reward collegiality more than candour. As a result, many papers receive less critical scrutiny than they probably should—not necessarily because scientists lack integrity, but because open disagreement carries professional consequences. One of science’s most important corrective mechanisms is therefore weakened by the very social structures that make collaboration possible.
I wonder whether science would benefit from developing its own culture of independent criticism.
In almost every other domain, criticism has evolved into an independent profession. Film critics do not work for movie studios. Political journalists are not employed by governments. Their value lies precisely in their independence. Science, by contrast, relies almost exclusively on active participants to evaluate one another. The people reviewing papers are also writing papers. The people deciding grants are also competing for grants.
There may be room for a complementary model: independent scientific critics, methodologists, retired investigators, or carefully moderated anonymous post-publication review. More interestingly, the internet has made it possible to create entirely new forums for scientific evaluation outside traditional academic structures. Thoughtful online communities, moderated discussion platforms, long-form newsletters, and reputation-based review systems could allow ideas to be challenged more openly than is often possible within the professional networks of academia itself.
The goal is to build intellectual spaces where scientific ideas can be evaluated as independently as possible from the incentives, relationships, and institutional pressures that inevitably shape modern research.
That is the experiment I want to run with Beyond the Abstract.
My hope is simply to create a space where scientific ideas are evaluated from first principles, discussed openly, and debated on their merits.
Science is often described as self-correcting, and over long enough timescales that is probably true. But self-correction can take years or even decades. The more immediate challenge is deciding what deserves our confidence today.
I don’t expect to settle that question, but I hope you’ll join me in exploring it.



