ModelRefs / ARC-AGI Methodology — Methodology
ARC-AGI Methodology — Methodology
ARC-AGI tests abstract pattern-recognition on grid puzzles designed to resist memorization.
Overview
What it measures: Fluid intelligence — solving novel visual reasoning puzzles from a few examples.
How it works
- Each task: 3–5 input/output grid pairs demonstrating a rule, then a test input.
- Model must produce the correct output grid.
- Tasks are private and rotated to prevent training contamination.
Strengths
- Strong contamination resistance
- Tests generalization, not recall
Limitations
- Small eval set
- Grid format is unusual for LLMs
Best use cases
- AGI-progress measurement
- Reasoning-model stress test
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to ARC-AGI Methodology — Methodology.
Frequently asked questions
What does ARC-AGI measure?
Fluid intelligence — solving novel visual reasoning puzzles from a few examples.
What are its main limitations?
Small eval set Grid format is unusual for LLMs
When should I use this benchmark?
AGI-progress measurement Reasoning-model stress test