| A smaller algorithm for complex matrix multiplication Public algorithmClaim source ↗ Evidence & sourcesA bilinear algorithm multiplies two 4 × 4 complex matrices with 48 scalar multiplications. Provisional assessment: A new construction beyond the documented 49-multiplication baseline. The numerical gain is small, but the construction improves a long-standing bound in this specific setting. This does not establish a new matrix-multiplication exponent or a universal speedup. Prior art: Strassen's recursive construction ↗ · 1969 Applying Strassen's 2 × 2 construction twice gives 49 multiplications for 4 × 4 matrices. AlphaTensor's 47-multiplication binary-field result is not the same setting. Review limits: Provisional assessment from the announcement and disclosed result. We have not independently run symbolic verification or exhaustively searched all bilinear constructions. | Google DeepMind AlphaEvolve / Gemini | Mathematics | N2New result Medium confidence · provisional | 14 May 2025 | ↗ |
| Larger cap-set constructions Peer-reviewed / constructionClaim source ↗ Evidence & sourcesFunSearch produces improved cap-set constructions, including a set of size 512 in dimension eight. Provisional assessment: The paper documents new constructions for a well-studied extremal problem. This is new mathematical output within an existing problem and search framework, rather than a new field or a solution of the general cap-set problem. Prior art: Edel and Bierbrauer; generalized product caps ↗ · 2004 Earlier finite cap constructions and product-cap methods provide the comparison framework in the FunSearch paper. A full reconstruction of the record history remains outstanding. Review limits: Prior-art coverage is partial. We rely on the paper's comparison and have not reproduced its programs or its asymptotic construction. | Google DeepMind FunSearch / Codey | Mathematics | N2New result Medium confidence · provisional | 14 Dec 2023 | ↗ |
| A 47-multiplication construction over the binary field Peer-reviewed / algorithmClaim source ↗ Evidence & sourcesAlphaTensor finds a 4 × 4 matrix-multiplication algorithm requiring 47 scalar multiplications in arithmetic modulo two. Provisional assessment: A new construction for the specified field. It advances the documented baseline within an established tensor-decomposition framework; its novelty should not be generalized to ordinary real or complex arithmetic. Prior art: Strassen's recursive construction ↗ · 1969 Recursive use of the seven-product algorithm requires 49 scalar multiplications. AlphaTensor's improvement here relies on binary-field arithmetic. Review limits: We have not rerun the released constructions. Subsequent algorithms do not change the novelty assessment at the 2022 cutoff. | Google DeepMind AlphaTensor | Computer science | N2New result Medium confidence · provisional | 05 Oct 2022 | ↗ |
| A mixed sodium–lithium solid electrolyte Preprint / experimentClaim source ↗ Evidence & sourcesAI-assisted screening leads to a mixed sodium–lithium solid electrolyte, followed by experimental validation with PNNL. Provisional assessment: Provisionally an incremental compositional extension in an established family of solid electrolytes. Screening scale and development speed measure the discovery process, not the material's novelty or superiority. Prior art: Asano et al., solid halide electrolytes ↗ · 14 Sept 2018 High-conductivity halide solid electrolytes for solid-state batteries were already demonstrated. This establishes family-level precedent, not an exact structural match to the new composition. Review limits: Low-confidence rating. Exact composition matching, patents, and a standardized performance comparison could change this classification. No percentage novelty or performance gain is assigned. | Microsoft Azure Quantum Elements / materials AI | Materials science | N1Incremental Low confidence · provisional | 08 Jan 2024 | ↗ |
| Generated inhibitors of the known DDR1 target Peer-reviewed / experimentClaim source ↗ Evidence & sourcesGENTRL generates candidate DDR1 inhibitors, with synthesized compounds tested experimentally. Provisional assessment: An incremental medicinal-chemistry result on an established target. Existing selective DDR1 inhibitors and related scaffolds predate the work. Similarity is evidence of close prior art, not proof of copying or an identical molecule. Prior art: Selective DDR1 inhibitors and prior kinase scaffolds ↗ · 10 Apr 2013 Gao and colleagues had reported selective, orally bioavailable DDR1 inhibitors in 2013. Bender's contemporaneous analysis found close chemical neighbors for a generated compound in ChEMBL. Review limits: We have not computed molecular fingerprints or run a patent search. Bender's 75% search threshold is not a measured universal novelty score and is not reported as one. Assay results from different studies are not directly compared. | Insilico Medicine GENTRL | Chemistry | N1Incremental Medium confidence · provisional | 02 Sept 2019 | ↗ |
| A formalization of Fermat's Last Theorem METHOD CONTEXTClaimant report / formalizationClaim source ↗ Evidence & sourcesAnthropic reports a complete Lean formalization of Fermat's Last Theorem, following existing mathematical proofs. Provisional assessment: The theorem is a known result. The potentially new contribution is its formal verification artifact and the automation process. N0 applies only to discovery of the theorem, not to the value or novelty of formalization. Prior art: Wiles; Darmon, Diamond, and Taylor ↗ · 1995 Wiles proved the theorem in work published in 1995. Anthropic explicitly says its formalization follows the later exposition by Darmon, Diamond, and Taylor. Review limits: Method-context record, excluded from discovery rankings. This is acknowledged reuse, not a plagiarism allegation. The formalization's validity and method novelty have not been independently audited here. | Anthropic Claude / Prove2Me | Mathematics | N0Known result High confidence · provisional | 04 Sept 2026 | ↗ |
| A proposed resolution of the Navier–Stokes problem Claimant report / proof artifactsClaim source ↗ Evidence & sourcesOpenAI announces a proposed solution of the Navier–Stokes Millennium Prize problem and releases a paper and a formal proof artifact. Provisional assessment: A potentially major result requiring independent specialist assessment. The closest proof strategy, theorem assumptions, and priority relative to concurrent work have not been resolved by this initial review. Prior art: Tao's averaged Navier–Stokes blowup construction ↗ · 03 Feb 2014 Tao established finite-time blowup for a modified, averaged equation. That result is relevant background but is not a solution for the original Navier–Stokes equations. Review limits: No proof verification or adjudication of priority was performed. Concurrent results for different equations must not be treated as identical. No claim of copying is made. | OpenAI Internal research model | Physics | —Unassessed Review incomplete | 08 Sept 2026 | ↗ |
| A claimed improvement in the zeta-zero proportion Claimant report / proof artifactsClaim source ↗ Evidence & sourcesAnthropic reports raising a lower bound for the proportion of Riemann-zeta zeros on the critical line from 41.6% to 67.2%. Provisional assessment: Potentially a new result combining existing mathematical tools. The cited numerical baseline, theorem assumptions, and full proof require specialist review before this leaderboard assigns novelty credit. Prior art: Unconditional pair-correlation results ↗ · 08 Jun 2023 Baluyot, Goldston, Suriajaya, and Turnage-Butterbaugh developed an unconditional Montgomery theorem. The announcement also credits Aryan and Bombieri. These are stated ingredients, not a verified latest numerical record. Review limits: We have not checked the Lean artifact, its assumptions, or the complete prior-art literature. This is not a claimed proof of the Riemann hypothesis itself. | Anthropic Unreleased research Claude | Mathematics | —Unassessed Review incomplete | 10 Aug 2026 | ↗ |
| A claimed disproof of the unit-distance conjecture Claimant report / external checks reportedClaim source ↗ Evidence & sourcesOpenAI reports a construction giving a polynomial improvement over classical grid examples for the planar unit-distance problem. Provisional assessment: A potentially major new result, but this initial review has not compared the proof with the closest constructions or independently established its quantitative improvement. It remains unassessed rather than receiving a score from the announcement alone. Prior art: Classical grid constructions and incidence bounds ↗ · 21 Jul 2025 The existing literature distinguishes lower-bound constructions from upper bounds on the maximum number of unit distances. A 2025 paper provides pre-claim context; an upper bound is not an interchangeable baseline for a new lower bound. Review limits: Independent specialists are reported as checking the result, but this leaderboard has not verified the proof or its exact prior-art distance. No numerical improvement is inferred from unmatched upper and lower bounds. | OpenAI Unreleased reasoning model | Mathematics | —Unassessed Review incomplete | 20 May 2026 | ↗ |
| A generated study of compositional regularization Workshop reviews / withdrawnClaim source ↗ Evidence & sourcesAn AI-generated manuscript reports negative results for compositional regularization and receives workshop review scores above the stated acceptance threshold. Provisional assessment: Review scores assess a manuscript, not its distance from prior art. The specific negative result needs comparison with earlier regularization and compositional-generalization experiments before a novelty rating can be assigned. Prior art: Earlier compositional-generalization and regularization research ↗ · Unresolved The closest experiment has not been identified in this initial review. The earlier AI Scientist framework is process context and does not establish priority for this scientific result. Review limits: The paper was withdrawn under the experiment protocol; the workshop did not perform a final meta-review. This record does not describe an accepted main-conference publication. | Sakana AI The AI Scientist-v2 | Computer science | —Unassessed Review incomplete | 12 Mar 2025 | ↗ |
| A large catalogue of predicted stable crystals Peer-reviewed / computationalClaim source ↗ Evidence & sourcesGNoME reports 381,000 new entries on an updated computational convex hull, within a larger set of predicted crystal structures. Provisional assessment: A substantial catalogue expansion is documented. A blanket novelty rating would conceal differences between new compositions, new structures, and variants of known prototypes. Entry-level matching against announcement-date databases is required. Prior art: Materials Project and OQMD catalogues ↗ · 2021-03 The study starts from March 2021 database snapshots. These are useful baselines, but leave a gap before the November 2023 public claim. Review limits: No structure-by-structure deduplication, patent search, or synthesis validation was performed here. Computational stability is distinct from experimental synthesizability. | Google DeepMind GNoME | Materials science | —Unassessed Review incomplete | 29 Nov 2023 | ↗ |
| Learning to control tokamak plasma shapes METHOD CONTEXTPeer-reviewed / experimentClaim source ↗ Evidence & sourcesA learned controller produces and maintains multiple plasma configurations on EPFL's TCV tokamak. Provisional assessment: An experimentally demonstrated control-method advance. It is tracked as enabling research, not counted as a new physical law or a separate discovery for each plasma configuration. Prior art: Conventional TCV magnetic feedback control ↗ · 2021 Model-based control and the study of shaped plasmas predate this work. The Nature paper describes the conventional controller architecture it replaces and cites the preceding control literature. Review limits: Research-method context record; excluded from discovery rankings. A separate controller-novelty review would require comparison with earlier control architectures. | Google DeepMind Deep reinforcement learning | Physics | —Unassessed Review incomplete | 16 Feb 2022 | ↗ |
| Predicted structures across the human proteome Peer-reviewed / predictionsClaim source ↗ Evidence & sourcesAlphaFold expands structural coverage across the human proteome, reporting confident predictions for 58% of residues. Provisional assessment: Expanded predictive coverage can enable discoveries, but does not establish that every prediction is a new biological finding. Each structure needs its own comparison to experimental structures, templates, and earlier predictions. Prior art: Experimental structures and earlier AlphaFold predictions ↗ · 2020 The study reports experimental coverage of 17% of human-protein residues. Earlier structure-prediction methods already existed, including the CASP13 AlphaFold system. Review limits: This record aggregates a catalogue and receives no blanket novelty score. Predictions and experimental observations are explicitly distinguished. | Google DeepMind AlphaFold 2 | Biology | —Unassessed Review incomplete | 22 Jul 2021 | ↗ |
| An AI-designed antifibrotic drug candidate Later peer-reviewed clinical dataClaim source ↗ Evidence & sourcesInsilico reports an AI-selected antifibrotic target and AI-designed candidate, later disclosed as the TNIK inhibitor rentosertib. Provisional assessment: TNIK inhibition was known before the claim. The possible novelty lies in the particular compound and its antifibrotic use. Neither target novelty nor chemical novelty can be resolved from the early undisclosed-target announcement alone. Prior art: NCB-0846 and earlier TNIK inhibition ↗ · 26 Aug 2016 Masuda and colleagues published a small-molecule TNIK inhibitor in cancer research in 2016. A 2018 study also examined TNIK inhibition in an epithelial–mesenchymal transition context. Review limits: A dated patent and chemical-structure review is outstanding. Later clinical studies strengthen evidence for the candidate but cannot retroactively establish novelty at the 2021 cutoff. Clinical efficacy is not rated here. | Insilico Medicine PandaOmics / Chemistry42 | Medicine | —Unassessed Review incomplete | 24 Feb 2021 | ↗ |