I'm curious about the Provider Adapter Harness because the original score with the Standard Harness was 62.7% at max reasoning, whereas the Provider Adapter Harness reached 99.95% at high reasoning.
ARC-AGI-3 consists of various games specifically made for the competition, and the goal is to clear the game with as few moves as possible. So there's no answer key per se, but Astra was probably trained on game-like reinforcement-learning environments to the point where most games look very similar to a game it was already trained on.
In any case, the competition was designed to highlight an issue with existing agents at the time where they couldn't maintain coherence over long sequences of actions, and it seems that problem has been remedied.
I'm looking forward to finding out which weakness ARC-AGI-4 will be targeting.
I'm curious about the Provider Adapter Harness because the original score with the Standard Harness was 62.7% at max reasoning, whereas the Provider Adapter Harness reached 99.95% at high reasoning.
I wonder if the agent(s) somehow got their hands on the answer key beforehand?
ARC-AGI-3 consists of various games specifically made for the competition, and the goal is to clear the game with as few moves as possible. So there's no answer key per se, but Astra was probably trained on game-like reinforcement-learning environments to the point where most games look very similar to a game it was already trained on.
In any case, the competition was designed to highlight an issue with existing agents at the time where they couldn't maintain coherence over long sequences of actions, and it seems that problem has been remedied.
I'm looking forward to finding out which weakness ARC-AGI-4 will be targeting.