Unlocking System 2 AI1
through neuroscience.

Voaige is a research lab translating neuroscience & AI research into computational principles for cognition, reasoning, and intelligence.

We then engineer the resulting principles into state-of-the-art LLM systems that maximize intelligence per FLOP.

Explore our research

A video of about a minute and a half, without sound, in three parts, each with the agent's trajectory under it as an axis of numbered steps. First, standard inference, where fixed model weights guide every step. An agent works through a task one step at a time. The model's prior experience is distilled into fixed weights, and temperature sampling can produce many different paths through token space. Zooming in on the third step, its paths represent possible token sequences. Temperature sampling selects one token at a time at random from temperature-adjusted sampling, tracing one of many possible paths; run again, the model can take a different path and reach a different outcome. Second, Test Time Cognition (System 2 thinking), deliberate search where it matters: an agent's run of three steps under the model's learned weights and a bank of episodic memory. At the first step, episodes from memory fill in the path the agent takes, and the weights make a small, narrow funnel along it. At the second, nothing in memory applies, so the weights' funnel opens wide and explicit search finds the answer inside it. At the third, episodic memory recalls more than half the way and search finds the rest. Each step is saved back to episodic memory, and the run reaches its goal. Third, explicit search solves hard problems. On the same task, standard inference takes the favored path at the third step too. It is wrong, and nothing checks it, so the run drifts off course. Test Time Cognition searches at steps 3 and 5 and reaches the goal.

The Search Hypothesis

Search among possible pathsA search tree branches out into many candidates. A sage funnel highlights the promising region while the surrounding tree remains visible. Candidates within the funnel are evaluated and one path is selected in green.

Intelligence is search.

Intelligence is the process of searching among possible actions, explanations, and solutions in pursuit of a goal. It involves generating and evaluating candidates, validating them against evidence, and pursuing promising paths.

Recall a discovery to guide the next searchA first search spreads out in every direction before it finds a solution, and that route is kept as a memory. The next search reuses the remembered route to get most of the way, and only searches the part that remains unresolved.First searchNext search

Memory is what makes search more efficient.

Memory retains experience as retrievable episodes or knowledge distilled into model weights, allowing previous computation to be reused. It can guide search, narrow the search space dramatically, resume search from a useful checkpoint, or supply a known solution that bypasses explicit search.

Coordinate search along an agent’s trajectory, with learned weights and episodic memoryAn agent’s trajectory runs from left to right through three steps, Step 1 to Step 3, to its goal. Above it hang the model’s learned weights, a small network whose learned pathway is sage, and a bank of episodic memory that holds past episodes, each a short coral route. Whenever the learned weights kick in, a pulse runs through the network, which makes a small sage funnel that travels down to cognition’s dot and opens there. At the first step a few episodes are recalled into the dot and memory fills in the path the agent takes, and the weights’ funnel opens small and narrow along it; the step is saved to the bank as a new episode. At the second step nothing in the bank applies, so the weights’ funnel opens wide over the whole step, and a full explicit search looks inside it; what it finds is saved. At the third step several episodes are recalled and memory recalls more than half the way; the weights’ funnel then opens around the rest, and search finds and checks the answer, which is saved too. The run then reaches its goal.Step 1Step 2Step 3Agent’s trajectoryLearned weightsEpisodic memory

Cognition is the coordination of search over time.

Cognition coordinates learning, knowledge, memory, and explicit exploration as an agent or organism interacts with its environment to achieve a goal. It guides when, where, and how to search, how to evaluate possibilities, and when to act, deliberate or stop.

Get faster, better, and cheaper at searchFour new problems are solved one after another. The first search explores widely before it reaches its goal. With experience, a learned prior narrows where each later search looks, so it explores far less and reaches its goal sooner, by a more direct route.Experience

Learning is getting faster, better, and cheaper at search.

Learning is the process by which experience improves a system’s ability to search. It can improve the quality and reliability of solutions, expand the range of problems the system can solve, and reduce the time and computation required.

Read Intelligence Is Search

From hypothesis to system

Test Time
Cognition.

Higher intelligence at lower cost,
keeping the same model weights and agent harness.

The TTC inference layer is what the agent interacts with. The model weights sit inside it, and TTC directs the inference without changing the model weights or the agent harness. It is the site of inference-time computation: the cognitive architecture that governs how reasoning unfolds between input and output.

Explore Test Time Cognition
The agent gives Test Time Cognition a task and gets an answer backThe agent sits above a box labelled Test Time Cognition and sends it a task. Inside the box, cognition, a single dot, coordinates search: the model's learned weights and episodic memory of past episodes both feed into it, and it sets a search going that explores new paths. One path checks out, and the answer goes back up to the agent.AgentPLANNING · TOOLS · CONTEXTTEST TIME COGNITIONModelLEARNED WEIGHTSEpisodicmemoryPAST EPISODESSearchEXPLORES NEW PATHSCognitionCOORDINATES SEARCHTASKANSWERThe agent gives Test Time Cognition a task and gets an answer backThe agent sits above a box labelled Test Time Cognition and sends it a task. Inside the box, cognition, a single dot, coordinates search: the model's learned weights and episodic memory of past episodes both feed into it, and it sets a search going that explores new paths. One path checks out, and the answer goes back up to the agent.AgentPLANNING · TOOLS · CONTEXTTEST TIME COGNITIONModelLEARNED WEIGHTSEpisodicmemoryPAST EPISODESSearchEXPLORES NEW PATHSCognitionCOORDINATES SEARCHTASKANSWER

Our Research Applied

Our research delivers gains while keeping the model weights and agent harness fixed.

3 distinct use cases, 6 agents evaluated, over 307B tokens.

Customer Service

Test Time Cognition achieves SOTA performance of 55.3%.

Customer service: Test Time Cognition achieves 55.3% performance. Voaige baselines: Grok, Qwen, GPT-5.5. External reference: Opus 5 from Sierra. Logarithmic cost axis. Arrows point from baseline configurations to TTC. Qwen 3.8 Max (xHigh reasoning): 52.41% performance, $0.7594; Grok 4.6 (High reasoning): 45.83% performance, $0.3949; Test Time Cognition — Qwen 3.8 Max (xHigh reasoning) + Grok 4.6 (High reasoning): 55.33% performance, $1.8053; GPT-5.5 (xHigh reasoning): 39.18% performance, $1.8166; Opus 5 (Max reasoning): 48.71% performance, $13.3215 Performance (%) ↑ 40 50 60 $0.25 $1 $5 $20 Qwen 3.8 Max (xHigh reasoning) Performance: 52.41% Mean cost per task: $0.7594 Qwen 3.8 Grok 4.6 (High reasoning) Performance: 45.83% Mean cost per task: $0.3949 Grok 4.6 Test Time Cognition — Qwen 3.8 Max (xHigh reasoning) + Grok 4.6 (High reasoning) Performance: 55.33% Mean cost per task: $1.8053 TTC GPT-5.5 (xHigh reasoning) Performance: 39.18% Mean cost per task: $1.8166 GPT-5.5 Opus 5 (Max reasoning) Performance: 48.71% Mean cost per task: $13.3215 Opus 5 Mean cost per task (USD, log scale)
Tau Knowledge

Coding

Test Time Cognition achieves higher performance at lower cost.

Coding: GPT-5, MiniMax, and Test Time Cognition with verified mean costs across 89 tasks. Lines connect measured reasoning levels. GPT-5 (minimal reasoning): 26.97% performance, $0.1163; GPT-5 (low reasoning): 37.45% performance, $0.2815; GPT-5 (high reasoning): 51.69% performance, $0.7722; MiniMax M2.5 Thinking: 46.82% performance, $0.0881; Test Time Cognition — GPT-5 (low reasoning) + MiniMax M2.5 (Thinking): 54.31% performance, $0.3424; Test Time Cognition — GPT-5 (medium reasoning) + MiniMax M2.5 (Thinking): 56.93% performance, $0.4747 Performance (%) ↑ 20 30 40 50 60 70 $0 $0.25 $0.5 $0.75 $1 GPT-5 (minimal reasoning) Performance: 26.97% Mean cost per task: $0.1163 Minimal GPT-5 (low reasoning) Performance: 37.45% Mean cost per task: $0.2815 Low GPT-5 (high reasoning) Performance: 51.69% Mean cost per task: $0.7722 GPT-5 High MiniMax M2.5 Thinking Performance: 46.82% Mean cost per task: $0.0881 MiniMax Test Time Cognition — GPT-5 (low reasoning) + MiniMax M2.5 (Thinking) Performance: 54.31% Mean cost per task: $0.3424 TTC Test Time Cognition — GPT-5 (medium reasoning) + MiniMax M2.5 (Thinking) Performance: 56.93% Mean cost per task: $0.4747 TTC Mean cost per task (USD)
Terminal-Bench 2.0

Financial Research

Test Time Cognition achieves higher performance at lower cost.

Financial research: Kimi and DeepSeek with Test Time Cognition. Logarithmic cost axis. Arrows point from baseline configurations to TTC. Kimi K3 (Max reasoning): 64.20% performance, $0.9930; Test Time Cognition — Kimi K3 (Max reasoning): 72.84% performance, $0.3169; DeepSeek V4 Flash Thinking: 53.01% performance, $0.0197; Test Time Cognition — DeepSeek V4 Flash Thinking: 70.37% performance, $0.0100 Performance (%) ↑ 50 60 70 80 $0.01 $0.1 $1 Kimi K3 (Max reasoning) Performance: 64.20% Mean cost per task: $0.9930 Kimi K3 Test Time Cognition — Kimi K3 (Max reasoning) Performance: 72.84% Mean cost per task: $0.3169 TTC DeepSeek V4 Flash Thinking Performance: 53.01% Mean cost per task: $0.0197 DeepSeek V4 Flash Test Time Cognition — DeepSeek V4 Flash Thinking Performance: 70.37% Mean cost per task: $0.0100 TTC Mean cost per task (USD, log scale)
Finance Agent V2

Stay in the loop

Follow our research.