Able to solve difficult math problems, but unable to reproduce research papers

Recent discussions surrounding PaperBenchX report that ChatGPT’s rate of reproducing results in scientific paper replication evaluations is only 13.98%. This figure reminds the industry that solving a math problem does not equate to being able to independently reproduce a research study; however, when the evaluation details remain insufficient, a single score should not be taken directly as an overall ranking of a model’s research capabilities.
Mathematical Problems Can Be Solved, Yet Research Results May Not Be Reproducible
PaperBenchX has recently brought an issue that can easily be obscured by model demonstrations into the spotlight: according to relevant reports, ChatGPT achieved a 13.98% reproduction rate in this scientific research reproducibility benchmark. This figure stands in sharp contrast to the mathematical research capabilities the model has recently demonstrated, and it makes a frequently conflated question concrete again: whether a model can produce an elegant answer and whether it can turn a study described in a paper into verifiable results are not the same capability.
This is not evidence that “models cannot do mathematics.” On the contrary, a report dated October 7 mentioned that OpenAI is publicly releasing a batch of mathematical research results generated by an internal frontier model, including Lean-formalized proofs. Formal proofs can be checked by computers, reducing the risk that proof texts which appear rigorous actually contain flaws. At the same time, the report also mentioned that each result consumed, on average, computing resources equivalent to approximately three hours of deep reasoning on ChatGPT Pro. The presentation of mathematical results and the reproduction rate of research papers measure different stages of the research process; one cannot be used to endorse the other.

13.98% Does Not Measure “Answer Accuracy”
Research reproduction is not a matter of reading a paper’s abstract and then generating a conclusion that looks reasonable. A complete reproduction task typically requires researchers to identify the methods and experimental setup from the paper, prepare the code and dependencies, process the data, run the experiments, and then compare the results with the metrics reported in the paper to determine whether the claims hold. A deviation at any stage can ultimately lead to the situation where “the code runs, but the paper’s results do not appear.”
This is also what distinguishes reproduction rates from common mathematics, knowledge, or coding benchmarks. Conventional questions usually have clear answers, and programming problems often have inputs and outputs that can be checked automatically. Reproducing a paper is more like an open-ended engineering task. A paper may not specify every implementation decision, a codebase may depend on a particular version or computing environment, and data preprocessing may affect the final metrics. Even if a model accurately summarizes a paper’s contributions and writes structurally complete code, that does not mean it has reproduced the paper’s experimental conclusions.
Therefore, 13.98% is a signal worth paying attention to, but considered in isolation, the percentage is still insufficient to determine exactly which stages the model handles well and where it fails. It may aggregate multiple tasks and multiple scoring dimensions, and it may also be affected by task difficulty, execution budgets, environment settings, and evaluation criteria. The materials currently available do not fully specify PaperBenchX’s task set, sample size, success criteria, model version, or test configuration. Without this information, the figure should not be interpreted as meaning that “ChatGPT can reproduce only 13.98% of scientific research,” nor can it be used to directly rank the research capabilities of different models.
An Entire Chain Separates Proof Checking from Paper Reproduction
Formal proofs and paper reproduction both pursue verifiability, but they verify different objects. Lean checks whether a formal proposition and its proof steps conform to logical rules. As long as the definitions and formalization process are correct, whether the proof passes verification has relatively clear criteria. This is highly important for mathematical research: it advances the standard from “peers find the reasoning credible” to “the proof steps can be checked one by one by a tool.”
Paper reproduction covers a much broader scope. The conclusions of machine learning papers typically depend on data, preprocessing, training code, hyperparameters, hardware, randomness, and other factors. Even when both the code and paper are publicly available, researchers may still need to determine which details must match exactly and which differences are acceptable. For experimental results, “obtaining a result” does not equal “successfully reproducing the study”: one must assess whether the metrics are sufficiently close, whether the variation falls within a reasonable range, whether key comparisons hold, and whether the main findings claimed by the paper have been reproduced.
To use an engineering analogy, Lean is more like an instrument that strictly checks circuit logic: the input format and rules are clear, and the output can be judged. Paper reproduction is more like rebuilding an entire production line from a set of design specifications. The drawings may not specify the requirements for every screw, and the material batch, equipment condition, and operating procedures may all affect the output. Successful verification of the former does not automatically prove that the latter can reliably produce the same product.
This also explains why “AI produced original mathematical results” and “AI can reliably reproduce scientific papers” can both be true. A system may propose an unpublished proof on certain constrained mathematical tasks while still being poor at managing a reproduction workflow spanning code, data, environments, and experimental analysis. The former demonstrates the model’s potential on specific reasoning tasks; the latter tests a more comprehensive ability to execute research.
OpenAI’s PaperBench Is Relevant Background, but It Is Not the Same Score as PaperBenchX
OpenAI released PaperBench in 2025 to evaluate AI agents’ ability to reproduce cutting-edge AI research. Its public description states that the benchmark selected 20 Spotlight and Oral papers from ICML 2024 and required agents to understand the papers’ contributions from scratch, develop codebases, and conduct experiments. The tasks were then broken down into smaller subtasks through hierarchical grading rubrics. This reflects an evaluation approach closer to actual research work: rather than looking only at the final answer, it examines the process and outputs involved in completing a research task.
However, PaperBench and the PaperBenchX mentioned in recent reports cannot be treated as the same benchmark merely because their names are similar. The former is an evaluation project publicly introduced by OpenAI; “13.98%” is a result reported in connection with PaperBenchX. The materials currently available do not provide sufficient information to prove that the two use the same task set, scoring method, or experimental configuration. Applying PaperBench’s publicly described design directly to PaperBenchX, or describing 13.98% as the score for OpenAI’s PaperBench, would create factual confusion.
This distinction matters. The results of scientific benchmarks depend heavily on the evaluation protocol: whether the model is allowed one run or multiple attempts, how much computing power it may use, whether it can call external tools, whether code and data are provided, and how human evaluators handle partial completion can all change the final score. A setup that allows the model only one attempt measures a different capability from one that allows repeated debugging, information searches, and access to a code execution environment. A benchmark name is not a methodological description, and a score cannot stand independently of its evaluation settings.
The Real Weakness May Lie in “Finishing the Research”
The value of reproduction tasks lies in bringing AI research capabilities back from demonstration-style output to the workflow itself. The model must turn an unstructured paper into clear implementation steps, inspect the code and environment, locate execution errors, determine whether the results match the paper, and identify which deviations could undermine the conclusions. This involves reasoning, software engineering, experimental design, and fault diagnosis.
For developers, the most practical distinction is this: a model’s ability to generate code does not mean it can complete an experiment; its ability to explain a paper does not mean it can reproduce the paper; and obtaining similar numbers in a single demonstration does not mean the result is reproducible. Research work often requires following the chain of evidence: Where did the data come from? Was the preprocessing consistent? Was the baseline fair? Were the parameters configured according to the paper? Were changes in the results caused by random seeds or implementation deviations? If a model cannot preserve and inspect this chain of evidence, it is more like an assistant that can accelerate local tasks than a researcher capable of independently taking responsibility for research conclusions.
This does not diminish AI’s value in research. On the contrary, a low reproduction rate can help research teams allocate responsibilities more clearly: models can assist with organizing papers, building code scaffolding, troubleshooting dependency conflicts, completing experiment records, and generating comparison tests; researchers can continue to review key results, while important conclusions can be confirmed through independent runs and peer verification. The benefits of models may initially appear in reduced search and debugging time, rather than in the complete elimination of researchers’ responsibility to verify results.
The Next Round of Evaluations Should Explain “Where the Failures Occur”
If PaperBenchX is to become an influential scientific benchmark, publishing a single overall score is only the starting point. More useful results should break down task completion: whether the model correctly understood the paper, whether it could produce a runnable implementation, whether it successfully configured the environment, whether the experiment ran to completion, and how far the final metrics were from those in the paper. The task samples, model versions, tool permissions, execution budgets, number of repetitions, and inter-rater consistency should also be disclosed so that other teams can verify where the numbers came from.
For research applications, failure types may be more useful for guiding improvements than the overall success rate. If the model mainly gets stuck installing dependencies and debugging code, improving tool use and software engineering capabilities may be more effective than continuing to scale up the language model; if the problem is that it cannot extract implicit assumptions from papers, improving long-context retrieval, structured experiment planning, and uncertainty communication may be more important; if the model can complete an experiment but cannot determine whether the results support the paper’s conclusions, the evaluation needs to examine statistical analysis and causal judgment rather than merely checking the code’s exit status.
Those deploying these systems should also treat process records from reproduction tasks as part of the output. Code versions, execution parameters, data sources, random seeds, and intermediate results should all be recorded. A model that submits only the final numbers is ill-equipped to support high-risk scientific decisions; a system that can explain deviations, flag uncertainty, and provide reproducible execution steps is more likely to enter the workflows of real research teams.
Conclusion: From “Able to Discover” to “Able to Complete Verifiably”
PaperBenchX’s 13.98% shifts the discussion from “Can models solve difficult problems?” to “Can models complete research tasks?” This shift is necessary. Formal proofs demonstrate AI’s progress in rigorous mathematical reasoning, while paper reproduction exposes the remaining distance between understanding a paper, implementing its code, and validating its experiments. The two do not contradict each other; they describe different boundaries of research automation.
The most cautious judgment for now is that AI can already produce results on some mathematical problems that deserve serious scrutiny, but it still has a significant gap to close before it can reliably and reproducibly take on complete research reproduction. The 13.98% figure is an evaluation signal that requires further verification and decomposition, not a final verdict on all models, all disciplines, or the research process as a whole. What matters next is not only whether models can produce another elegant result, but also whether the benchmark can publish a publicly verifiable protocol and whether models can make every step of the research process withstand scrutiny.
References
- PaperBench code and evaluation materials, the entry point to OpenAI’s publicly available related code repository; note that PaperBench and PaperBenchX should not be treated as directly equivalent evaluations.
- PaperBenchX discussion thread, an entry point to related discussions in the Chinese online community; the 13.98% figure in this article is relayed according to the provided report, and the evaluation details still need to be verified against the original paper or project materials.
- Search for OpenAI mathematical research repositories, used to locate the mathematical research results and Lean proof materials mentioned in reports from October 2026.



