07/2026 - ARC-AGI 3 Review and What Monty Would Need to Solve it

@vclay presents on ARC-AGI 3, an interactive benchmark designed to test whether AI agents can efficiently adapt to unfamiliar tasks without relying on language, external knowledge, or memorized/pre-trained solutions. The team discusses how Monty’s current capabilities compare with the skills required to pass the test. The team also explores whether ARC-AGI-3 could serve as a practical test environment for prototyping new features in Monty, like causality, object segmentation, compositionality and goal inference.

Main Video

Summary Video

0:00 Introduction
0:13 ARC-AGI
0:59 What is ARC-AGI-3?
2:17 What’s an Unsaturated Benchmark?
3:22 ARC Prize Foundation & Prize
5:04 Francois Chollet Tweet on Benchmarking Intelligence
6:04 Past Challenges & Performance with ARC-AGI-1 and ARC-AGI-2
9:07 ARC-AGI 3 Targets Agentic Intelligence
10:10 ARC-AGI-3 Tests These Four Core Components: Exploration, Modeling, Goal-Setting, Planning and Execution
12:22 More Details about ARC-AGI-3
13:34 How Does ARC-AGI-3 Measure Efficiency?
14:21 How are ARC-AGI-3 Environments Designed?
16:44 Discussion: Hidden Priors in Games
17:59 How is Performance Measured?
22:16 Live Demo: Example ARC-AGI-3 Games
34:40 Current Approaches to Solve ARC-AGI-3 & Scores
37:04 Why This Benchmark Matters for Monty
39:06 How Does Monty Do on the Skills Required to Solve ARC-AGI-3?
40:21 Discussion on Causality
43:54 Using Learned Models for Planning & Goal Inference
46:11 Recognize/Segment Objects, Understand Symmetry, Rotation & Elementary Topology & Abstract Models/Generalization
49:48 Demo: Collaborative Agent Game
52:34 Discussion: Analyzing the Demo Game
58:11 Discussion: 2D vs 3D Modeling
1:04:35 What’s Missing in Monty to Solve ARC-AGI-3
1:10:32 Debate Over What’s Not Important to Solve ARC-AGI-3
1:17:59 Focus on Behaviors, Causality & Achieving Goals
1:18:50 Is ARC-AGI 3 a Good Benchmark for Monty?
1:28:58 Goals, Rewards & Curiosity
1:34:59 Implementation Approach & Prototyping Plans
1:38:44 Dynamic Compositionality & Forgetting Mechanisms
1:40:58 Finite vs Infinite Games

4 Likes

Indeed, the supremacy of TBT may not be evident until the project is in its final stages of development (when performance on ARC won’t matter anymore). If you are trying to get to the Moon, then a person climbing up a tree will report steady progress, unlike a person building a rocket. ARC looks more like an arms race between benchmark developers and programmers trying to make LLMs resemble AGI more closely (which, unfortunately, encourages investors’ belief in the unlimited potential of LLMs and will make the AI winter colder).

The worst thing is that the capabilities required from the system by ARC are tightly interconnected (as any practical industrial application demands), so there are no meaningful landmarks. Meanwhile, I completely agree: there are common computational problems that hinder any approach to AGI (starting from LLMs and ending with logic-based systems), and TBT certainly provides insights into how to solve these problems. If the community could state them a little more clearly (for newcomers who are limited in expertise and experience, like myself), it would make the learning process much faster.

A person may specialize in a certain problem for a long time, and they will point out right away what’s wrong with the algorithm (computational problem – TBT’s unique insights on how to solve it – flowchart of the algorithm currently adopted – implementation details). On the whole, a modular structure of the program, where each module is independent of the others (and focuses on a specific computational problem), is the best solution.

By the way, the discussion reminded me of some topics that I had forgotten about. Here are some of the questions I hope you may be interested in considering:

  1. How can a program be represented via reference frames? For example, could cooking breakfast be considered a program? How could functions (reusable program blocks – open a cupboard, take a plate, put some food on it) and variables (degrees of freedom, so to say – take a spoon/fork/knife/plate, put some cereal/porridge/bread) be represented?

  2. The question that bothers me most, perhaps: birds may be encoded by the length of their necks and legs, for example. But a leg is a complex object of irregular shape, as is a neck. How is this transition performed by the brain – from a complex object to just a point in space?

  3. Suppose I take a cup of tea in my hand. All five of my fingers are working together; they do not merely agree on the identity of the object – they are actively interacting with it and coordinating their actions. This process is not completely automatic – I can move each one of my fingers separately. How do they manage to perform this feat? How does the right hand know what the left is doing?

  4. For the sake of curiosity: they say “cross all boundaries.” Could there actually be a space where crossing a boundary triggers an action?

  5. Young children supposedly add “what” (the appearance of an object) only if “where” fails (its location). But how do they individuate an object? Movement alone may not be enough (there are objects we perceive as separate even though they may be glued together for their entire existence, and there are objects with parts moving relative to each other, like a running animal, that we consider to be one thing). What is an object, on the whole? Perhaps the concept of an object itself is derivative? Or do we learn so many things about the world later in life that this variability obscures the underlying principles of individuation?

  6. What is a role? How could it be represented via reference frames? For example, how do we know whether someone or something qualifies for a job? How are the capabilities of someone or something encoded?

  7. Does our mind care more about what properties an object should have to be something, or about what properties it shouldn’t have to be something? Perhaps both? If so, to what degree?

  8. Charles Darwin constructed a tree of species, and we can construct a genealogical tree. But we may know nothing at all about how these species evolved exactly (or we may know little about our ancestors). How do we still manage to infer links between them?

  9. Is the relationship between two positions of one object (the object is moving over time) the same as the relationship between two objects?

  10. To what degree can each person be considered a single cortical column? For example, people give orders to each other. Is there a similar format of communication between columns?

  11. Does our mind build models of the people around us (tiny identities)? Does it build models of objects in a similar way? Can we switch between these identities? What defines our sense of self, then? Memories, perhaps?

  12. How are operations represented in the brain (like the addition of two numbers)? What about division/multiplication?

  13. How are constraints represented in the human brain? How is searching performed in the human brain? Say I want to find some fruit. How do I impose constraints on the environment in which it may be growing (not a desert, nor a mountain, nor a dense forest)?

  14. Suppose I have a ball: half of it is white, and half is black. Another ball is colored with black and white chaotically and in the same proportion (like an old TV screen). How does the brain represent each ball’s coloring? What things does our mind consider accidental (the second ball is neither white nor black – but the first one is clearly not accidentally colored)? How does our mind measure structure? Perhaps it has something to do with the painter’s intentions?

  15. How are the relationships between two objects represented? For example, “thing A is on top of thing B” – how is “on top” represented?

2 Likes

I’d like to share three personal opinions, for discussion:

  1. I don’t believe the LLM path leads to AGI. LLMs will keep solving more and more problems, but in my view what they do is always replicating and recombining a known knowledge distribution — interpolation inside the distribution keeps getting more polished, while extrapolation beyond it has not appeared. The ARC series is a yardstick designed precisely for this weak spot: it measures not “how many skills you have” but “how efficiently you acquire new skills on tasks you have never seen”. Hence the pattern that keeps repeating: each time a new ARC generation is released, human scores stay stable while machine scores reset to zero; and when the scores do eventually get pushed up, it is usually by massive test-time search and compute, not by the “figure out the rules within a few interactions” that ARC-3 is trying to test. The scores will be pushed up sooner or later — but how they get pushed up will itself expose the nature of the route.

  2. To me, TBT is the most complete first-principles framework for AGI we currently have — but the list of problems it still has to solve is also very long. The capability list in this video speaks for itself about how large the gap is — the good news is that every item on it lives inside the framework’s own vocabulary: what is missing is implementation, not theoretical patches. My own experience running compositional object-recognition experiments says the same: even at the purely concrete level, part-whole compositional understanding is still the frontier. And the real key milestone, I believe, is the leap from understanding concrete objects to understanding abstract concepts — the theory says the same cortical algorithm should work over more abstract reference frames, but so far this step remains at the level of theory. What is interesting is that ARC-3 sits right across this boundary: the grid is a concrete space, while the rules — what a key does, what counts as winning — are abstract objects. As a forcing function for that leap, it is very well chosen.

  3. What is most unfavorable to TBT, I think, is neither theory nor engineering — it is resources. Over the past three years, global capital has bet almost entirely on the transformer lineage and its derivatives (LLMs, multimodal models, VLAs, physical AI, all kinds of world models). Compute, talent, tooling, benchmarks and narrative are all self-reinforcing in the same direction — the Matthew effect fulfills itself. For research like TBT — genuinely promising but requiring long-horizon investment — the biggest risk is being marginalized by this enormous inertia, left without adequate resources and attention for a long time. Seen from another angle, though: as the mainstream saturates benchmark after benchmark through scale, the marginal information of an orthogonal route is actually rising; neural networks themselves were once the under-resourced minority, from 2006 to 2012. What has to be fought is the technical inertia of an entire era — but isn’t that exactly the interesting part?

4 Likes

It’s just that, from my point of view, the project is currently facing two major difficulties:

1) For the project as a whole: not being able to distinguish TBT’s potential from the amount of attention it receives.

Programs cannot necessarily be designed in the same way that the brain is studied. The history of programming languages suggests that different levels of abstraction and different programming paradigms can be useful for expressing different kinds of computation. Assembly language, functional programming, object-oriented programming, and other approaches each provide different ways of structuring a computational system. Similarly, the unstructured use of commands such as goto is generally discouraged because certain forms of complexity are easier to manage when the structure of a program is made explicit.

The human brain illustrates the importance of interactions among components: certain functions cannot simply be isolated or switched off without affecting the wider system. This suggests that, for some questions, a complex cognitive system may need to be analyzed as an integrated whole rather than as a collection of independent parts.

However, when someone comes to the project structured in this way, they can easily feel inadequate because they have to spend huge amounts of time analyzing it without necessarily receiving any immediate reward. After several months, they may start consulting more experienced researchers while still feeling uncomfortable: Maybe I should have read more before taking up the developers’ time? How sensible are my questions?

I personally became fond of TBT because I had been reading all kinds of articles about AGI for about a year. As a result, I had developed many constraints on what I thought a coherent AGI framework should look like, and “the mind as reference frames” was indeed a good fit.

But it will be difficult to convince people who are not already interested in AGI of TBT’s potential. There is no way to completely avoid uncertainty or guarantee Monty’s future. In a sense, the community’s greatest achievement is also its greatest difficulty: it has sustained interest in the project despite not having a working prototype for a long period of time.

2) For individual researchers: not being able to separate their self-worth from the social acceptance they receive.

Some people imagine that their duty is to be able to give the correct answer under any circumstances. However, this can make them shy away from areas they understand poorly, which ultimately makes their knowledge more fragile.

After a certain point, passive reading starts to produce diminishing returns, demanding more and more time for the same amount of additional expertise. The greatest achievement of an individual may instead be the willingness to put oneself in a disadvantaged position for long periods of time by acknowledging the gaps and flaws in one’s understanding.

The most valuable skill is learning how to obtain the necessary facts robustly and relatively quickly.

As Einstein supposedly put it: “Education is what remains after one has forgotten what one has learned in school.”

1 Like

Hi, another very intresting discussion, thanks. Just with reference to the Francois Chollet comment ‘you don’t need intelligence to solve a problem via brute force’, some might say 86 billion neurons and trillions of synapses is brute force :slight_smile: . As @vclay noted, no animal or young child could solve these tasks.

I think a huge amount of prior knowldege is brough to the problem by the human brain. If you were to distill the task down to its fundamental level, you have a list of 4096 numbers and if you change a small group if those numbers other groups of numbers change in some complex pattern. Your goal is to get the 4096 numbers into a specific sequence. Impossibe for even a human brain. Now convert the numbers into real world concepts of rooms and walls and moving agents and it becomes far more solveable. But you must already have all of those concepts in your brain or system.

@tslominski made a very good point about the perspective. The god’s eye view of the world is completely unnatural and probably can only be conceived of in a brain of human level complexity. We can switch between first-person perspective games and third-person perpective games where we are represented by a single token or a whole army, as in chess.

As I think Tristan also suggested, the TBP could create its own, simpler ARC-AGI type challenges as a means of leanring the concepts required to play the game. For example start with a 64 x 64 room to learn the concept of a room, then add interior walls, then other moveable objects etc. But you would have to decide whether these concepts are learned from first person or third person perspective. A room concept from third person person perpective is very different from a room concept from first person perspective, humans effortlessly switch between the two.

Alex

1 Like

Well, looks like OpenAI came in like a wrecking ball. :laughing:

On the Semi-Private set, it obtained a high-score of 99.9% with ARC’s long-memory harness (“Provider Adapter” as they call it), and 62.7% with the standard harness. Still burned a small sedan’s worth of tokens, and understandably so because Astra scales computation at the moment of reasoning, but progress is progress.

“Astra matched and surpassed human parity.”

However, it’s far from over… François Chollet on X: "@Yossi_Dahan_ @polynoamial ARC-4 is in the works, to be released early 2027. ARC-5 is also planned. The final ARC will probably be 6-7. The point is to keep making benchmarks until it is no longer possible to propose something that humans can do and AI can't. AGI ~2030." / X

It’s hard to say in which direction the next challenges will go. If there’s 3D, Monty has a slight headstart, but it’s clear it will need more than shape memorization and recall. Things like full-scene segmentation based on dynamic saliency (e.g. singling out a moving car in parking lot full of identical cars), and temporal behavior memorization/recall (e.g. melodies/patterns, or X happens after Y). Just speculating here, but those might require some hefty heterarchical scaffolding, e.g. long chains of series-connected LMs.

So, it’s a good thing that Hawkins is pushing toward benchmarks not necessarily to win, but to identify dead angles.

It is mindboggling how much is being poured into transformers and GPUs. Putting all their eggs in the same trillion-dollar basket… I don’t really see inertia as a problem here, but rather the low societal ROI on that massive build-up, which seems to be progressively souring public opinion about AI in general.

Even among those interested in AGI, the definition of what AGI is remains a constant debate. I always spend a bit of time every week in pro-AI communities to stay updated, and there are some people who proclaim that Mythos is already AGI. (My go-to reply is “Can it make a sandwich?”)

However, whenever I bring up TBP, the people’s reactions are always positive. Some reminisce about reading Hawkins’ books in the past and are pleased to see progress. Those who never heard of it are curious. I even encountered a few TBP forum users in the wild. :wink:

The sentiment I get is that even some pro-AI people are starting to grow tired of LLM glorification. A server rack solving an Erdos problem doesn’t mean that much in comparison to a physical robot proficiently performing a basic chore.

Einstein spent his last 3 decades working on a unified field theory without success, but he showed us that true pioneers do not fear being at a disadvantage, no matter how long! :face_with_monocle:

2 Likes

The approach GPT-6 used is really interesting:

“It represented game mechanics as logical rules and developed its own domain-specific language shorthand to track state and plan actions”

The human brain will tend to convert the problem into our familiar mind set of rooms and walls and pushing things. The AI converts the problem into the way it works, code, rules and states. Because the problem is set in the virtual world this works fine. Transfer the problem into the physical world with all its unpredictability and imprecision and it’s a very different challenge.

1 Like