AlienCode
A calculation language whose familiar-looking operators follow hidden semantics. Programs are run on private inputs, so passing takes rule-aware behaviour rather than memorised examples.
Two executable worlds whose rules contradict everything a model has read. An interpreter and a proof checker hold the ground truth, so every answer can be checked.
The two standard ways of testing it each fail, for a different reason.
A correct answer may simply have been memorised, and contamination cannot be ruled out for a black-box model.
Without a reliable verifier, a real discovery and a confident mistake look identical.
SHATTER looks destructive and multiplies. EMIT(100) prints 127. Nothing in the language adds. The rules were written for this benchmark, so no model has read them, and they are executable, so the evaluator can check every held-out answer against them.
Both are deterministic and executable. Every active rule has at least one held-out task that depends on it.
A calculation language whose familiar-looking operators follow hidden semantics. Programs are run on private inputs, so passing takes rule-aware behaviour rather than memorised examples.
A Fitch-style natural-deduction system with patched inference rules and opaque identifiers. A checker verifies every proof, and unprovable goals earn credit only for a correct refusal.
The manual describes standard semantics, and the world was perturbed underneath it. Pick a probe.
Weights never change. Everything learned lives in one context window, and every held-out answer is produced with the sandbox closed.
M0 after the seed, M4 after four rounds, gain G = M4 − M0. Each cell is the mean of three runs.
Reasoning labels name each provider's highest setting, not a common compute scale. The last column is diagnostic and never enters G. Rank is by M4.
Fifteen pairs climb to the endpoint, three saturate early, one peaks and falls back, one never moves.
Does exploration work, where does it stall, and what does a unit of budget buy?
Each bar is a median difference across twenty pairs; the dots count how many came out positive.
Against an arm that is simply handed the rules, AlienLogic runs sit 45.5 pp below a near-perfect endpoint. AlienCode runs nearly reach one that is only 69.8%.
Stating a rule raises the pass rate on the tasks it governs by 39.1 pp — and still leaves 14.8% failing.
The protocol rations what a system may ask of the world, never what it may think, so cost splits into two currencies that disagree on size.
Throughput buys chances to learn a rule without delivering one.
The paper's own plots, re-rendered from the same scripts. Numbering follows the PDF; click any figure to enlarge.
The construction is not specific to exploration: any ability confounded by pre-training, but verifiable in an unfamiliar world, can be measured this way.