A video of about a minute and a half, without sound, in three parts, each with the agent's trajectory under it as an axis of numbered steps. First, standard inference, where fixed model weights guide every step. An agent works through a task one step at a time. The model's prior experience is distilled into fixed weights, and temperature sampling can produce many different paths through token space. Zooming in on the third step, its paths represent possible token sequences. Temperature sampling selects one token at a time at random from temperature-adjusted sampling, tracing one of many possible paths; run again, the model can take a different path and reach a different outcome. Second, Test Time Cognition (System 2 thinking), deliberate search where it matters: an agent's run of three steps under the model's learned weights and a bank of episodic memory. At the first step, episodes from memory fill in the path the agent takes, and the weights make a small, narrow funnel along it. At the second, nothing in memory applies, so the weights' funnel opens wide and explicit search finds the answer inside it. At the third, episodic memory recalls more than half the way and search finds the rest. Each step is saved back to episodic memory, and the run reaches its goal. Third, explicit search solves hard problems. On the same task, standard inference takes the favored path at the third step too. It is wrong, and nothing checks it, so the run drifts off course. Test Time Cognition searches at steps 3 and 5 and reaches the goal.
The Search Hypothesis
Intelligence is search.
Intelligence is the process of searching among possible actions, explanations, and solutions in pursuit of a goal. It involves generating and evaluating candidates, validating them against evidence, and pursuing promising paths.
Memory is what makes search more efficient.
Memory retains experience as retrievable episodes or knowledge distilled into model weights, allowing previous computation to be reused. It can guide search, narrow the search space dramatically, resume search from a useful checkpoint, or supply a known solution that bypasses explicit search.
Cognition is the coordination of search over time.
Cognition coordinates learning, knowledge, memory, and explicit exploration as an agent or organism interacts with its environment to achieve a goal. It guides when, where, and how to search, how to evaluate possibilities, and when to act, deliberate or stop.
Learning is getting faster, better, and cheaper at search.
Learning is the process by which experience improves a system’s ability to search. It can improve the quality and reliability of solutions, expand the range of problems the system can solve, and reduce the time and computation required.
Higher intelligence at lower cost, keeping the same model weights and agent harness.
The TTC inference layer is what the agent interacts with. The model weights sit inside it, and TTC directs the inference without changing the model weights or the agent harness. It is the site of inference-time computation: the cognitive architecture that governs how reasoning unfolds between input and output.