Building System 2 AI
through neuroscience.

Voaige is an AI research lab translating insights from cognitive and systems neuroscience into computational principles for reasoning, memory, and search.

We are engineering these principles into LLMs under the constraint of building state-of-the-art,
commercially deployable systems.

Explore our research

A video of under a minute, without sound, in two parts, with the agent's trajectory under it throughout as an axis of numbered steps. First, an agent works through a task one step at a time. Every step opens onto several possible future states, and the model's learned priors guide which one to pursue. At the first two steps the priors clearly favor one path and the agent takes it. At the third the paths waver and none stands out, so cognition, which coordinates search and decides when and where to intervene, steps in: the learned priors narrow the directions, episodic memory recalls a similar past episode, and explicit search explores the remaining paths and checks the answer. Second, explicit search solves hard problems. On the same task, standard inference, where fixed model priors guide every step, takes the favored path at the third step too. It is wrong, and nothing checks it, so the run drifts off course. Test Time Cognition (System 2 thinking), deliberate search where it matters, searches at steps 3 and 5 and reaches the goal.

Our research

From neuroscience to
high-performing AI systems.

We pursue two connected lines of research to build state-of-the-art,
commercially deployable systems.

From existing literature to a working systemPapers from existing literature are distilled into principles, drawn as simple shapes. The principles are engineered into a system, drawn as a brain whose folds are rows of the same shapes.EXISTING LITERATUREPRINCIPLESSYSTEM From existing literature to a working systemPapers from existing literature are distilled into principles, drawn as simple shapes. The principles are engineered into a system, drawn as a brain whose folds are rows of the same shapes.EXISTING LITERATUREPRINCIPLESSYSTEM

Distill the principles.

Draw on existing neuroscience research to identify computational mechanisms for learning, memory, and the coordination of search.

Engineer the systems.

Investigate the mechanics of large language models to turn those principles into algorithms and working systems.

Our Guiding Scientific Hypotheses

Search among possible pathsA search tree branches out into many candidates. A sage funnel highlights the promising region while the surrounding tree remains visible. Candidates within the funnel are evaluated and one path is selected in green.

Intelligence is search.

Intelligence is the process of searching among possible actions, explanations, and solutions in pursuit of a goal. It involves generating and evaluating candidates, validating them against evidence, and pursuing promising paths.

Recall a discovery to guide the next searchA first search spreads out in every direction before it finds a solution, and that route is kept as a memory. The next search reuses the remembered route to get most of the way, and only searches the part that remains unresolved.First searchNext search

Memory is what makes search more efficient.

Memory is the retention and retrieval of experience that allows previous computation to be reused. It can guide and narrow search, resume it from a useful earlier point, or supply a known solution that bypasses explicit search.

Coordinate learning, memory, and searchCognition makes three deliberate looks. Memory in an upper region provides a coral clue pointing right. Search in a lower region finds a green clue pointing back left. Both signals meet at the same junction, guiding the funnel to that region. Memory and search then work together to verify the answer.

Cognition is the coordination of search over time.

Cognition brings learned knowledge, memory, and exploration together to decide where to search, how to evaluate possibilities, and when to continue, change direction, or stop.

Get faster, better, and cheaper at searchFour new problems are solved one after another. The first search explores widely before it reaches its goal. With experience, a learned prior narrows where each later search looks, so it explores far less and reaches its goal sooner, by a more direct route.Experience

Learning is getting faster, better, and cheaper at search.

Learning is the process by which experience improves a system’s ability to search. It can improve the quality and reliability of solutions, expand the range of problems the system can solve, and reduce the time and computation required.

Read Intelligence Is Search

From hypothesis to system

Test Time
Cognition.

Higher intelligence at lower cost,
keeping the same model weights and agent harness.

The TTC inference layer is what the agent interacts with. The model weights sit inside it, and TTC directs the inference without changing the model weights or the agent harness. It is the site of inference-time computation: the cognitive architecture that governs how reasoning unfolds between input and output.

Explore Test Time Cognition
Test Time Cognition coordinates search inside the inference layerThe agent sends a task down into the Test Time Cognition layer, which contains the model. Inside it, cognition coordinates three ways of searching: the model's learned weights narrow where to look, memory recalls a route from a past episode, and a short new search resolves the rest. The answer goes back up to the agent, and the outcome is kept in memory as a new episode.AGENTPlanning · Tools · ContextTEST TIME COGNITIONMODELLEARNED WEIGHTSMEMORYPAST EPISODESTASKANSWERKEEP AS A NEW EPISODE

Our Research Applied

Our research delivers gains across models, domains, and agent harnesses.

The results below are drawn from our broader evaluation program.

Tokens used for evals
307B
Use cases evaluated
3
Agents tested
6

Customer Service

Test Time Cognition achieves SOTA performance.

55.3%Performance
Customer service: Test Time Cognition achieves 55.3% performance. Voaige baselines: Grok, Qwen, GPT-5.5. External reference: Opus 5 from Sierra. Logarithmic cost axis. Arrows point from baseline configurations to TTC. Qwen 3.8 Max (xHigh reasoning): 52.41% performance, $0.7594; Grok 4.6 (High reasoning): 45.83% performance, $0.3949; Test Time Cognition — Qwen 3.8 Max (xHigh reasoning) + Grok 4.6 (High reasoning): 55.33% performance, $1.8053; GPT-5.5 (xHigh reasoning): 39.18% performance, $1.8166; Opus 5 (Max reasoning): 48.71% performance, $13.3215 Performance (%) ↑ 40 50 60 $0.25 $1 $5 $20 Qwen 3.8 Max (xHigh reasoning) Performance: 52.41% Mean cost per task: $0.7594 Qwen 3.8 Grok 4.6 (High reasoning) Performance: 45.83% Mean cost per task: $0.3949 Grok 4.6 Test Time Cognition — Qwen 3.8 Max (xHigh reasoning) + Grok 4.6 (High reasoning) Performance: 55.33% Mean cost per task: $1.8053 TTC GPT-5.5 (xHigh reasoning) Performance: 39.18% Mean cost per task: $1.8166 GPT-5.5 Opus 5 (Max reasoning) Performance: 48.71% Mean cost per task: $13.3215 Opus 5 Mean cost per task (USD, log scale)
  • Baseline · Opus 5 from Sierra
  • TTC: Qwen + Grok

TTC = Test Time Cognition

Tau Knowledge

Coding

Test Time Cognition achieves higher performance at lower cost.

+5.2 ppPerformance
38.5%Lower cost
Coding: GPT-5, MiniMax, and Test Time Cognition with verified mean costs across 89 tasks. Lines connect measured reasoning levels. GPT-5 (minimal reasoning): 26.97% performance, $0.1163; GPT-5 (low reasoning): 37.45% performance, $0.2815; GPT-5 (high reasoning): 51.69% performance, $0.7722; MiniMax M2.5 Thinking: 46.82% performance, $0.0881; Test Time Cognition — GPT-5 (low reasoning) + MiniMax M2.5 (Thinking): 54.31% performance, $0.3424; Test Time Cognition — GPT-5 (medium reasoning) + MiniMax M2.5 (Thinking): 56.93% performance, $0.4747 Performance (%) ↑ 20 30 40 50 60 70 $0 $0.25 $0.5 $0.75 $1 GPT-5 (minimal reasoning) Agent: mini-SWE-agent 2.2.1 Performance: 26.97% Mean cost per task: $0.1163 Minimal GPT-5 (low reasoning) Agent: mini-SWE-agent 2.2.1 Performance: 37.45% Mean cost per task: $0.2815 Low GPT-5 (high reasoning) Agent: mini-SWE-agent 2.2.1 Performance: 51.69% Mean cost per task: $0.7722 GPT-5 High MiniMax M2.5 Thinking Agent: mini-SWE-agent 2.2.1 Performance: 46.82% Mean cost per task: $0.0881 MiniMax Test Time Cognition — GPT-5 (low reasoning) + MiniMax M2.5 (Thinking) Agent: mini-SWE-agent 2.2.1 Performance: 54.31% Mean cost per task: $0.3424 TTC Test Time Cognition — GPT-5 (medium reasoning) + MiniMax M2.5 (Thinking) Agent: mini-SWE-agent 2.2.1 Performance: 56.93% Mean cost per task: $0.4747 TTC Mean cost per task (USD)
  • Baseline: GPT-5 · MiniMax
  • TTC: GPT-5 + MiniMax

TTC = Test Time Cognition

Terminal-Bench 2.0 · mini-SWE-agent 2.2.1

Financial Research

Test Time Cognition achieves higher performance at lower cost.

+13.0 ppPerformance
58.6%Avg. cost reduction
Financial research: Kimi and DeepSeek with Test Time Cognition. Logarithmic cost axis. Arrows point from baseline configurations to TTC. Kimi K3 (Max reasoning): 64.20% performance, $0.9930; Test Time Cognition — Kimi K3 (Max reasoning): 72.84% performance, $0.3169; DeepSeek V4 Flash Thinking: 53.01% performance, $0.0197; Test Time Cognition — DeepSeek V4 Flash Thinking: 70.37% performance, $0.0100 Performance (%) ↑ 50 60 70 80 $0.01 $0.1 $1 Kimi K3 (Max reasoning) Performance: 64.20% Mean cost per task: $0.9930 Kimi K3 Test Time Cognition — Kimi K3 (Max reasoning) Performance: 72.84% Mean cost per task: $0.3169 TTC DeepSeek V4 Flash Thinking Performance: 53.01% Mean cost per task: $0.0197 DeepSeek V4 Flash Test Time Cognition — DeepSeek V4 Flash Thinking Performance: 70.37% Mean cost per task: $0.0100 TTC Mean cost per task (USD, log scale)
  • Baseline: Kimi K3 · DeepSeek V4 Flash
  • Same models + TTC

TTC = Test Time Cognition

Finance Agent V2

Stay in the loop

Follow our research.