179 questions
No questions match those filters.
O3 spends roughly $200 per ARC-AGI task at 75% accuracy...
This is one of the questions in the full AI/ML interview bank. Pro unlocks all 1789 questions; Premium includes the same bank plus the highest daily Practice limit.
See plansThe O3 result indicates that the necessary representations for high-level reasoning are already latent within the model weights, but eliciting them requires extensive search and verification. If a smaller model achieved 75% accuracy via single-pass inference, it would suggest that the reasoning capability has been distilled more efficiently or that the model has discovered a structural shortcut specific to the ARC-AGI task distribution. These possibilities can be tested by evaluating the model on a held-out set of novel, ARC-AGI-style tasks.