Schema harness self-reports ~99% on ARC-AGI-3, far above frontier models
A system called Schema self-reported 98.98% Relative Human Action Efficiency on the public ARC-AGI-3 set using Opus 4.8 and Fable 5, and 95.35% with GPT-5.6, on a benchmark where agents must infer a game's rules from raw 64x64 grids with no instructions. Verified frontier-model performance had reached only 13.33% on the public set with GPT-5.6 Sol at max reasoning.
The results are self-reported and not yet ARC Prize-verified. ARC-AGI-3 has proven exceptionally hard because it offers no object list, rule sheet, stated goal, or shaped reward, forcing agents to form and revise hypotheses as they act.
View full digest for July 17, 2026