Science is still persuasion
On researcher degrees of freedom and the future of science, again
Having a tool that speeds up science gives us a unique opportunity to think about what science actually is. Yes, such tool is AI.
We can argue about what a unit of science really is or does or should be, but let’s start with a simple definition that it identifies something true about the world: elephants have rough skin, the average height of an American in 2025 is 66 inches, a 10% increase in capital investment in non-farm industries causes a 1% increase in output over 10 years.
AI will speed up a lot of things in science; right now it is making data work super quick, obvious to many empirical social scientists. “Speeding up” means that we go from questions to answers quicker and then can, in principle, collect more true facts in the same amount of time. It also means we can probe the same questions more times, which I will focus on here.
Of course those examples above come with some uncertainty. We did not touch every elephant, we did not measure the height of every American in 2025, and we did not check with every non-farm firm. The uncertainty resulting from sampling certain elephants, people, and firms from a broader pool is aptly called sampling uncertainty, and we measure this numerically with a model-based “standard error.” It is a thought experiment about what would happen if we resampled from a superpopulation a bunch of times—how different would those estimates be?
Some facts about the world are causal, like capital investment causing an increase in output. Causal claims are usually supported with experimental design-based evidence, which adds another layer of uncertainty. If we chose 100 similar non-farm firms and invested in capital for a random half of them and then measured their output, our “causal fact” would be the difference in outcomes between the treatment group and the control group. But there is uncertainty about the hidden characteristics of the specific firms, so we do another thought experiment: what would happen if we randomized the treatment assignment to different firms within our sample? Confusingly this is also called a standard error, usually design-based standard error.
The size of the standard errors matters for persuasion. Should anyone believe the fact? Let me be clear, it is not the only thing or main thing that matters for persuasion: are you in general trustworthy, do you have a good track record, do you present your results in clear enough ways that others can understand? But if you are estimating a quantity, reporting small standard errors is a nice reassurance.
Everyone knows that statistical significance doesn’t matter if it’s achieved nefariously: p-hacking, negligence, miscalculation, fake data, etc. all taint our credence. In fact, as readers, we are basically all good Bayesians such that we never take any single study as God-given truth, no matter how much we trust the authors and methods.
On his Substack Scott Cunningham has been leading an interesting discussion on how he uses Claude Code, and the latest has been on what he calls non-standard errors. A non-standard error would be the calculational equivalent thought experiment where you hold fixed the data—the statistical sample—and if applicable, also the treatment assignment, but vary the researcher doing the science. Or, more specifically, every time a researcher makes a choice about how to proceed in the process, we think about the possibility of going down the road not taken. Let’s call that a researcher degree of freedom. Combinatorially, this blows up fast:
From raw data cleaning through estimation through table construction, you might face ten major decision points. At each one, you might have two reasonable options. That’s 2^10 possible situations that could’ve occurred by the time the estimates were calculated or 1,024.
If there were 3 options for each of those ten tasks, then it’s 3^10 possible ensuing situations or 59,049. That’s 59k different hypothetical estimates.
Again, this matters for persuasion. Suppose you produce a unit of science. Imagine instead of 10 that there were 100 different decisions, 2 defensible options each, made in the process of producing the unit, including which title to use, which numbers to include in the abstract, maybe even which font to use. The number of different paths is absolutely enormous,1 but suppose that I agree wholesale with 90 of them such that those don’t sway my belief in the results.2 Then we are back to the 10-node case; then there are 10 decisions I might have made the opposite decision as you. If you could convince me that those 10 decisions don’t qualitatively change the results, and you can do that for every possible readers’ overlapping but different set of 10, that is equivalent to shrinking the non-standard error to zero.
As a note on the way we currently do empirical social science (or is it already in the past?), this is exactly why we do robustness checks, and it is exactly why different reviewers care about different robustness checks. “Persuade me!” they say, “show me that if I would have done the analysis slightly differently, the estimates wouldn’t change.”3
Speeding up science changes how we can calculate these bounds on uncertainty from non-standard errors. In fact this will likely scale faster than the bounds we can put on standard errors: there is only so many ways to increase sample sizes and replicated experiments. But in observational studies, AI makes it quite a bit easier to map out many different paths.
What’s next?
My expectation is that, at first, it will be the job of the authors to report the range of estimates produced, reporting something like a non-standard error, just like it is (was) their job to run the myriad of robustness checks. Basically a self-replication exercise baked into the original estimates.
But this will evolve to a demand-side system where “papers” are accompanied by a platform of all the tools used in the process, and the reader will ask their AI their “what if we did X instead of Y, does that change the estimate?” Like if I were reading an experimental chemistry paper, and it came with a pre-set lab with all the ingredients, a lab director and assistants, and I could ask them as I read the paper, “what if we tweaked the proportions by X?” and they did it right there in front of me and together we saw the outcome.
Journals are in a great place to foster this in this new age of science. Mandatory code and data sharing is the first step to this. But that’s more like watching a video of what happened in the lab—better than nothing for increasing my personal credence in the unit of science, but it can’t answer my personal “what ifs?”
More than half of the work done currently ends up in supplemental appendices anyway, so journals should just standardize how this gets processed such that it is easy to keep “living” papers. When I ask my “what if X” question to the paper, I have my AI pick up the tools off the work bench in the lab (sitting standardized with the journal) and produce a new little addition to the output and update the non-standard errors. If someone comes later with the same question, the answer is already there; if they have a new question, rinse and repeat.
I wouldn’t mind if journals somehow became obsolete in the age of AI-assisted science, but I do understand how interest groups steer technological change. Additionally, we will still want trustworthy clearinghouses, and journals are the natural torch bearers.
All that to say, Scott is on to something. The current process of “reading a literature” to get the sense of some fact, updating our priors a little with each unit—this will become much more formalized into estimate-able non-standard errors.
If we are measuring them, though, we can’t let them be called non-standard! How about analytic standard errors, based on analytic variance, due to procedural uncertainty?
Sounds more persuasive.
See also:
Scott Cunningham’s excellent series.
My related vision: The New Era of Science.
Gauti Eggertsson says: “I now find myself replicating papers and experimenting with frontier methods in an evening or a few days using Claude Code. That would have taken weeks before — which in practice meant I wouldn’t have done it at all.” Tyler Cowen says he’s too conservative.
Where is my vision too conservative?
1.268 nonillion or 1.268 x 10^30.
Paper idea: do catchier titles get cited more? Let me know if you want to write this with me, I have the data.
This is also why even though modern macroeconomics is a cool sophisticated form of the philosophy of counterfactual thinking, I have a hard time taking any number very seriously—there is just too many researcher degrees of freedom!

