Human-AI-computer interaction: measuring whether AI strengthens unaided judgment
Corpus: hai probe in packages/bkt: 998 frozen four-choice items across the seven canon branches, 68 held out by review, no retested probes yet; learning system in learning/research-os/learning-system: 57,060 enumerated planning cases, Lean 4 contract, synthetic prediction experiment; the rules joining the two are design
Abstract
A tool that helps a person answer today and leaves them no better at answering alone next week has made them dependent on it. We state a measurement protocol that separates the two outcomes. On one task set, three scores are taken: the person's unaided score Hp, the model's score A, and the person's joint score Jp with the model's answer in view. The multiplier is mp = Jp / max(Hp, A), the gain is D = Jp - max(Hp, A), and the unaided test is repeated after seven days to give retention Rp = Hp(t + 7 days) - Hp(t) and learning L = Jp(t + 7 days) - Hp(t + 7 days). A tool that raises Jp while L stays at or below zero across its interval marks dependence.
The learning side of the measurement is the minimum prerequisite learning problem: the catalog is a directed acyclic graph of concepts, a learner's state is a closed set of mastered concepts, and the least a learner must add to reach a target is the required set R(t, M), the target's prerequisite closure minus their mastery, a count proved minimal in Lean 4 under stated assumptions. Three rules will join the two sides: probe items drawn at the learner's frontier, Rp reported against the engine's predicted recall, and mastery updated from the solo half only. All three are design; what exists is the item bank, the two conditions, the seven-day retest and the scoring, and the probe code holds no mastery or FSRS logic today.
The instrument is the hai probe in Bucket's bkt package: a frozen bank of 998 four-choice items, sessions of 40 items in 20 pairs matched on tier and on whether the model got the pair's items right, a 20 second think window before the model's answer appears, no feedback until the retest, guess-corrected scores and a bootstrap over pairs for every interval. The instrument is built and no probe has been retested, so this report carries no result on a person; the learning system's own evaluation, 57,060 enumerated planning cases with zero errors and a synthetic prediction experiment, is reported with its source. The report states the planned analysis, paired comparisons stratified by education level, a divergence hypothesis on how reliance on one model narrows the range of answers across a group, and the rule by which the scores decide which tasks Research OS hands to the model and which it keeps with the person.
Key findings
- Four scores on the same items: Hp, A, Jp at day 0 and Hp again at day 7; mp = Jp / max(Hp, A), Rp = Hp(t + 7) - Hp(t), L = Jp(t + 7) - Hp(t + 7).
- Dependence flag: the 95% interval of Jp - Hp above zero and the interval of L at or below zero.
- Status 2026-10-07: bank frozen at 998 items, 68 flagged by review, no AI scores collected, no probe run, no retest. No numbers are reported.
- A half-width of 0.10 on Jp - Hp needs about 342 items per condition, 17 probes; no trend before 5 retested probes.
- Learning side: the minimum number of concepts to reach a target is |R(t, M)|, proved in Lean 4 (minimum_unit_distance, minimum_weighted_effort) under five stated assumptions; 57,060 enumerated planning cases on every DAG up to five nodes with zero errors.
- Mastery as shipped: M = P^alpha R^beta with alpha = beta = 1, P = sigmoid(theta + 0.2), R = FSRS-5 retrievability at 90 days, mastered at M >= 0.7. The frontier draw, the predicted-recall report and the solo-only mastery update are design; the probe code holds neither today.
Figures







Data and code
- probe source, packages/bkt/src/hai ↗
- bkt package README, hai commands ↗
- learning system brief, Lean contract and metrics.json; CC-BY-4.0 text, MIT code ↗
- mastery and FSRS-5 as implemented, src/lib/academy; MIT ↗
- paper figures, seven webp; MIT code output over the repository's synthetic metrics, no person data ↗
Cite this paper
@techreport{dichio2026humanaimultiplier,
title = {Human-AI-computer interaction: measuring whether AI strengthens unaided judgment},
author = {Dichio, Gianangelo},
institution = {Bucket Foundation},
year = {2026},
month = {10},
url = {https://www.bucket.foundation/research/papers/human-ai-multiplier},
note = {Protocol report, version 1.1, 2026-10-07, no results on a person}
}