Advancing the state
of the art in AI
through neuroscience.

Voaige is an AI research lab translating insights from cognitive and systems neuroscience into computational principles for reasoning, memory, and search.

We are engineering these principles into LLMs under the constraint of building state-of-the-art,
commercially deployable systems.

Explore our research

Begin with Fixed Learning.

Fixed learning and relevant experiences guide search for novel solutions. Useful outcomes return as evidence and become memories that guide future search.

Fixed learning, recalled experiences, and live search form one coordinated system. An evaluated outcome is retained as a specific episode that can guide the next search; broader patterns can be learned over time.

Our research

From neuroscience to
high-performing AI systems.

We pursue two connected lines of research to build state-of-the-art,
commercially deployable systems.

I.

Distill the principles.

Draw on existing neuroscience research to identify computational mechanisms for learning, memory, and the coordination of search.

II.

Engineer the systems.

Investigate the mechanics of large language models to turn those principles into algorithms and working systems.

Our guiding hypothesis

Search among possible paths The branching search network from the hero circuit, with one evaluated route selected from many alternatives.

Intelligence is search.

Find useful actions, explanations, and solutions among many possibilities. Enumerate candidates, evaluate them, and select a promising path.

Coordinate learning, memory, and search Learned patterns and remembered experiences connect to search, while outcomes feed back into memory.

Cognition coordinates search.

Decide what to explore, what to recall, and when further computation is worth its cost. Bring learning, memory, and search into one coordinated system.

Recall a discovery to guide the next search A retained experience in the memory bank connects to the selected search path. Evidence from its outcome returns to memory.

Memory makes search efficient.

Reuse prior discoveries to guide the next search. Retain useful experiences, recall them when relevant, and deliberate on what remains unresolved.

Read Intelligence Is Search

From hypothesis to system

Test Time
Cognition.

Our research, engineered into
the way AI reasons.

The TTC layer sits between the agent and the model, operating without modifying either. It is the site of inference-time computation: the cognitive architecture that governs how reasoning unfolds between input and output.

Explore Test Time Cognition
Agent

Planning · Tools · Context

Voaige

Test Time Cognition

Search · Evaluate · Allocate effort

Model

Weights · Architecture · Training

Model weights and agent harness code remain unchanged.

Our Research Applied

Higher intelligence at lower cost for the same model, same agent.

Test Time Cognition delivers gains across models, domains, and agent harnesses.

Without changing model weights or agent harness code.

Tokens
307B
Domains
3
Agents
6

Customer Service

Test Time Cognition achieves SOTA performance.

55.3%Performance
Customer service: Test Time Cognition achieves 55.3% performance. Voaige baselines: Grok, Qwen, GPT-5.5. External reference: Opus 5 from Sierra. Logarithmic cost axis. Dashed arrows compare selected configurations. Qwen 3.8: 52.41% performance, $0.7594; Grok 4.6 High: 45.83% performance, $0.3949; Test Time Cognition: 55.33% performance, $1.8053; GPT-5.5 xHigh (Voaige baseline): 39.18% performance, $1.8166; Opus 5 Max (Sierra leaderboard): 48.71% performance, $13.3215 Performance (%) ↑ 40 50 60 $0.25 $1 $5 $20 Qwen 3.8: 52.41%, $0.7594 Qwen 3.8 Grok 4.6 High: 45.83%, $0.3949 Grok 4.6 Test Time Cognition: 55.33%, $1.8053 TTC GPT-5.5 xHigh (Voaige baseline): 39.18%, $1.8166 GPT-5.5 Opus 5 Max (Sierra leaderboard): 48.71%, $13.3215 Opus 5 Mean cost per task (USD, log scale)
  • Baseline · Opus 5 from Sierra
  • TTC: Qwen + Grok
Tau Knowledge

Coding

Test Time Cognition achieves higher performance at lower cost.

+5.2 ppPerformance
38.5%Lower cost
Coding: GPT-5, MiniMax, and Test Time Cognition with verified mean costs across 89 tasks. Lines connect measured reasoning levels. Minimal: 26.97% performance, $0.1163; Low: 37.45% performance, $0.2815; GPT-5 High: 51.69% performance, $0.7722; MiniMax M2.5 Thinking: 46.82% performance, $0.0881; Test Time Cognition — GPT-5 low + MiniMax M2.5 Thinking: 54.31% performance, $0.3424; Test Time Cognition — GPT-5 medium + MiniMax M2.5 Thinking: 56.93% performance, $0.4747 Performance (%) ↑ 20 30 40 50 60 70 $0 $0.25 $0.5 $0.75 $1 Minimal: 26.97%, $0.1163 Minimal Low: 37.45%, $0.2815 Low GPT-5 High: 51.69%, $0.7722 GPT-5 High MiniMax M2.5 Thinking: 46.82%, $0.0881 MiniMax Test Time Cognition — GPT-5 low + MiniMax M2.5 Thinking: 54.31%, $0.3424 TTC Test Time Cognition — GPT-5 medium + MiniMax M2.5 Thinking: 56.93%, $0.4747 TTC Mean cost per task (USD)
  • Baseline: GPT-5 · MiniMax
  • TTC: GPT-5 + MiniMax
Terminal-Bench 2.0 · mini-SWE-agent 2.2.1

Financial Research

Test Time Cognition achieves higher performance at lower cost.

+13.0 ppPerformance
58.6%Avg. cost reduction
Financial research: Kimi and DeepSeek with Test Time Cognition. Logarithmic cost axis. Dashed arrows compare selected configurations. Kimi K3: 64.20% performance, $0.9930; Test Time Cognition — Kimi K3 Max: 72.84% performance, $0.3169; DeepSeek V4 Flash Thinking: 53.01% performance, $0.0197; Test Time Cognition — DeepSeek V4 Flash Thinking: 70.37% performance, $0.0100 Performance (%) ↑ 50 60 70 80 $0.01 $0.1 $1 Kimi K3: 64.20%, $0.9930 Kimi K3 Test Time Cognition — Kimi K3 Max: 72.84%, $0.3169 TTC DeepSeek V4 Flash Thinking: 53.01%, $0.0197 DeepSeek V4 Flash Test Time Cognition — DeepSeek V4 Flash Thinking: 70.37%, $0.0100 TTC Mean cost per task (USD, log scale)
  • Baseline: Kimi K3 · DeepSeek V4 Flash
  • Same models + TTC
Finance Agent V2

Join us

A unique place to work on the mission.

We bring together cognitive and systems neuroscientists, AI researchers, mathematicians, and engineers around a single shared challenge: to develop a coherent computational framework for cognition and memory, and test it in AI systems that push beyond the current frontier.

Join Us

Stay in the loop

Follow our research.