bucket foundation — inverse omegabucket.foundation
§ Research · paper

Machine learning for scientific discovery: the solvability frontier

Gianangelo DichioBucket FoundationBucket Foundation report · 1.0 (full report)2026-10-07CC-BY-4.0 text; figures and data per file below

Corpus: solvability atlas, 5,080 problem statements across seven branches, bge-small-en-v1.5 embeddings, 50 stored neighbours per row

Abstract

A first pass, the Solver Gap Engine of 2026-09-30, ranked 1,364 open problems from the formal-conjectures repository by their nearest solved problem in another file and found 31 at cosine similarity 0.8 or above. The frontier is the second pass. We embed 5,080 problem statements across seven canon branches with a small sentence encoder, call a problem's reach its highest similarity to a solved problem, and draw a frontier at the 10th percentile of solved problems' reach, 0.781 on this set. Open problems are sorted into a reach class: 251 close to known results, 764 borderline, 911 that need a new idea and 103 in a branch with too few solved rows to judge.

A backtest at cutoffs 2005 and 2021 scores the rule on the questions posed by those years under two outcome codings, and the backtest is inconclusive. Settled only: at 2005 the inside rows settle at 21% against 6% outside (112 and 71 rows, permutation p = 0.006, AUC 0.800) and at 2021 at 12% against 1% (238 and 126 rows, p < 0.001, AUC 0.883), but 11 and 21 of the scored rows are solved rows with no resolved year, and without them the gaps are 13% against 6% (p = 0.131) and 3% against 1% (p = 0.269). Settled or advanced: 46% inside against 61% outside at 2005 (183 rows, p = 0.051, AUC 0.479) and 43% against 39% at 2021 (364 rows, p = 0.503, AUC 0.523). The verdict depends on how partial is coded, and the embeddings and stored neighbours come from the 2026 corpus, so the test is a check against known outcomes under that leak and no forecast.

Key findings

Figures

The solvability frontier drawn as a disc: solved rows inside, open rows inside the circle at reach 0.781, outside rows beyond it, unsampled rows as grey crosses
Figure 2. The frontier on 5,080 rows. Solved rows fill the inner disc by reach, open rows inside fill the ring up to the circle at 0.781, outside rows sit beyond it by how far their reach falls short, and unsampled rows are grey crosses. Angle is the row's rank along the first two principal components of the embeddings.
Bar charts of the 5,080 rows by branch, status, form, source, licence, resolution evidence, posed and resolved years, zone, class, level, text kind and length
Figure 1. Make-up of the 5,080 rows under every label the pipeline reads: mathematics holds 4,112 rows, formal-conjectures supplies 3,567, Apache-2.0 covers 3,557, 588 rows carry a posed year, 236 a resolved year, and 2,194 are variants of another row.
Resolved rate by 2026 inside and outside the cutoff frontier at 2005 and 2021 under both codings and by length band
Figure 3. Resolved rate by 2026 inside and outside the cutoff frontier under both codings, over all scored rows, with the undated solved rows removed, and by length band.
Stacked bars of open problems per reach class, coloured by branch
Figure 4. Reach classes by branch: 251 close to known results, 764 borderline, 911 that need a new idea and 103 unsampled. Mathematics holds 247 of the 251 and 737 of the 764.

Data and code

Cite this paper

@techreport{dichio2026solvabilityfrontier,
  title        = {Machine learning for scientific discovery: the solvability frontier},
  author       = {Dichio, Gianangelo},
  institution  = {Bucket Foundation},
  year         = {2026},
  month        = {10},
  url          = {https://www.bucket.foundation/research/papers/solvability-frontier},
  note         = {Full report, version 1.0, 2026-10-07}
}