bucket foundation — inverse omegabucket.foundation
§ Research · paper

Human-AI-computer interaction: measuring whether AI strengthens unaided judgment

Gianangelo DichioBucket FoundationBucket Foundation report · 1.1 (protocol report with the learning system, no results on a person)2026-10-07CC-BY-4.0 text; figures MIT code output over the repository's own synthetic metrics; code MIT

Corpus: hai probe in packages/bkt: 998 frozen four-choice items across the seven canon branches, 68 held out by review, no retested probes yet; learning system in learning/research-os/learning-system: 57,060 enumerated planning cases, Lean 4 contract, synthetic prediction experiment; the rules joining the two are design

Abstract

A tool that helps a person answer today and leaves them no better at answering alone next week has made them dependent on it. We state a measurement protocol that separates the two outcomes. On one task set, three scores are taken: the person's unaided score Hp, the model's score A, and the person's joint score Jp with the model's answer in view. The multiplier is mp = Jp / max(Hp, A), the gain is D = Jp - max(Hp, A), and the unaided test is repeated after seven days to give retention Rp = Hp(t + 7 days) - Hp(t) and learning L = Jp(t + 7 days) - Hp(t + 7 days). A tool that raises Jp while L stays at or below zero across its interval marks dependence.

The learning side of the measurement is the minimum prerequisite learning problem: the catalog is a directed acyclic graph of concepts, a learner's state is a closed set of mastered concepts, and the least a learner must add to reach a target is the required set R(t, M), the target's prerequisite closure minus their mastery, a count proved minimal in Lean 4 under stated assumptions. Three rules will join the two sides: probe items drawn at the learner's frontier, Rp reported against the engine's predicted recall, and mastery updated from the solo half only. All three are design; what exists is the item bank, the two conditions, the seven-day retest and the scoring, and the probe code holds no mastery or FSRS logic today.

The instrument is the hai probe in Bucket's bkt package: a frozen bank of 998 four-choice items, sessions of 40 items in 20 pairs matched on tier and on whether the model got the pair's items right, a 20 second think window before the model's answer appears, no feedback until the retest, guess-corrected scores and a bootstrap over pairs for every interval. The instrument is built and no probe has been retested, so this report carries no result on a person; the learning system's own evaluation, 57,060 enumerated planning cases with zero errors and a synthetic prediction experiment, is reported with its source. The report states the planned analysis, paired comparisons stratified by education level, a divergence hypothesis on how reliance on one model narrows the range of answers across a group, and the rule by which the scores decide which tasks Research OS hands to the model and which it keeps with the person.

Key findings

Figures

Timeline of probe sessions: day 0 probe with solo and pair items, day 7 unaided retest, day 14 next probe
Figure 1. The session timeline. Each probe splits 40 unseen items into solo and pair conditions on day 0, is retested unaided in full on day 7, and unlocks its scores only then; the next probe draws 40 new items.
Diagram of the four scores Hp, A, Jp and Hp at day 7 and the statistics derived from them
Figure 2. The four scores and the statistics derived from them: the multiplier and the gain from Hp, A and Jp; retention on the solo items and learning on the pair items from the day 7 scores; dependence where the gain is positive and learning is not.
Prerequisite chart with a verified foundation, two ready concepts and a blocked target, and the shortest knowledge-state path through the diamond
Figure 3. The prerequisite chart for a target T with one verified foundation, two ready concepts and a blocked target, and the shortest knowledge-state path from empty mastery through the same diamond: four transitions, equal to the required set.
Path of confirmed coordinates on two axes for a three-concept fixture, with coverage G and radial extent r at each milestone
Figure 4. The knowledge region on the three-concept fixture: the path of confirmed coordinates as a, then b, then t are confirmed, and the catalog coverage G with the radial extent r at each milestone, from analysis/results/metrics.json.
FSRS-5 retrievability curves over 120 days for four stabilities, and the mastery surface M = P R with the 0.7 contour
Figure 5. FSRS-5 retrievability over 120 days for four stabilities with the day 7 retest and the 90-day horizon marked, and the mastery surface M = P R over proficiency and retention with the 0.7 mastered contour, as implemented in src/lib/academy.
A layered catalog with known, frontier and blocked concepts and four probe items at the frontier, beside a one-probe timeline with FSRS predicted recall
Figure 6. The protocol over a learner's knowledge region: probe items drawn at the frontier, the day 0 scores, the day 7 retest and the FSRS-5 predicted recall for a solo and a pair item. The points show the pattern and carry no measured value.
Brier scores of three models on synthetic data and two residual correlations with bootstrap intervals
Figure 7. The synthetic statistical experiment: Brier scores of the three models on 4,800 test responses, and the two prespecified residual correlations with their 97.5% bootstrap intervals and Holm-adjusted p-values, from analysis/results/metrics.json.

Data and code

Cite this paper

@techreport{dichio2026humanaimultiplier,
  title        = {Human-AI-computer interaction: measuring whether AI strengthens unaided judgment},
  author       = {Dichio, Gianangelo},
  institution  = {Bucket Foundation},
  year         = {2026},
  month        = {10},
  url          = {https://www.bucket.foundation/research/papers/human-ai-multiplier},
  note         = {Protocol report, version 1.1, 2026-10-07, no results on a person}
}