Advancing the state
of the art in AI
through neuroscience.

Voaige is an AI research lab translating insights from cognitive and systems neuroscience into computational principles for reasoning, memory, and search.

We are engineering these principles into LLMs under the constraint of building state-of-the-art, commercially deployable systems.

Explore our research

Begin with Fixed Learning.

Fixed learning and relevant experiences guide search for novel solutions. Useful outcomes return as evidence and become memories that guide future search.

Fixed learning, recalled experiences, and live search form one coordinated system. An evaluated outcome is retained as a specific episode that can guide the next search; broader patterns can be learned over time.

Our research

From neuroscience to
high-performing AI systems.

We pursue two connected lines of research to build state-of-the-art, commercially deployable systems.

I.

Distill the principles.

Draw on existing neuroscience research to identify computational mechanisms for learning, memory, and the coordination of search.

II.

Engineer the systems.

Investigate the mechanics of large language models to turn those principles into algorithms and working systems.

Every design element must serve a computational purpose. We evaluate its contribution against strong baselines, measuring both what the system can solve and the computation it requires.

Our guiding hypothesis

Search among possible paths The branching search network from the hero circuit, with one evaluated route selected from many alternatives.

Intelligence is search.

Find useful actions, explanations, and solutions among many possibilities. Enumerate candidates, evaluate them, and select a promising path.

Coordinate learning, memory, and search Learned patterns and remembered experiences connect to search, while outcomes feed back into memory.

Cognition coordinates it.

Decide what to explore, what to recall, and when further computation is worth its cost. Bring learning, memory, and search into one coordinated system.

Recall a discovery to guide the next search A retained experience in the memory bank connects to the selected search path. Evidence from its outcome returns to memory.

Memory makes it efficient.

Reuse prior discoveries to guide the next search. Retain useful experiences, recall them when relevant, and deliberate on what remains unresolved.

Read Intelligence Is Search

From principles to measured gains

Higher performance.
Lower cost.

Trajectories evaluated
120k+
Domains
3
Agents
6

Customer Service

Cognition achieves SOTA performance.

55.3%Performance
Customer service: cognition achieves 55.3% performance. Voaige baselines: Grok, Qwen, GPT-5.5. External reference: Claude Opus 5 from Sierra. Logarithmic cost axis. Dashed arrows compare selected configurations. Qwen 3.8: 52.41% performance, $0.7594; Grok 4.6 High: 45.83% performance, $0.3949; Cognition: 55.33% performance, $1.8053; GPT-5.5 xHigh (Voaige baseline): 39.18% performance, $1.8166; Claude Opus 5 Max (Sierra leaderboard): 48.71% performance, $13.3215 Performance (%) ↑ 40 50 60 $0.25 $1 $5 $20 Qwen 3.8: 52.41%, $0.7594 52.4% Qwen 3.8 Grok 4.6 High: 45.83%, $0.3949 45.8% Grok 4.6 Cognition: 55.33%, $1.8053 55.3% Cognition GPT-5.5 xHigh (Voaige baseline): 39.18%, $1.8166 39.2% GPT-5.5 Claude Opus 5 Max (Sierra leaderboard): 48.71%, $13.3215 48.7% Claude Opus 5 Mean cost per task (USD, log scale)
  • Qwen 3.8 Max · Grok 4.6 High · GPT-5.5 xHigh
  • Sierra: Claude Opus 5 Max
  • Voaige cognition: Qwen + Grok
Tau Knowledge

Coding

Cognition achieves higher performance at lower cost.

+5.2 ppPerformance
38.5%Lower cost
Coding: GPT-5, MiniMax, and cognition with verified mean costs across 89 tasks. Lines connect measured reasoning levels. Minimal: 26.97% performance, $0.1163; Low: 37.45% performance, $0.2815; GPT-5 High: 51.69% performance, $0.7722; MiniMax M2.5 Thinking: 46.82% performance, $0.0881; Cognition — GPT-5 low + MiniMax M2.5 Thinking: 54.31% performance, $0.3424; Cognition — GPT-5 medium + MiniMax M2.5 Thinking: 56.93% performance, $0.4747 Performance (%) ↑ 20 30 40 50 60 70 $0 $0.25 $0.5 $0.75 $1 Minimal: 26.97%, $0.1163 Minimal Low: 37.45%, $0.2815 Low GPT-5 High: 51.69%, $0.7722 51.7% GPT-5 High MiniMax M2.5 Thinking: 46.82%, $0.0881 MiniMax Cognition — GPT-5 low + MiniMax M2.5 Thinking: 54.31%, $0.3424 Cognition Cognition — GPT-5 medium + MiniMax M2.5 Thinking: 56.93%, $0.4747 56.9% Cognition Mean cost per task (USD)
  • GPT-5
  • MiniMax M2.5 Thinking
  • Voaige cognition: GPT-5 + MiniMax
Terminal-Bench 2.0 · mini-SWE-agent 2.2.1

Financial Research

Memory achieves higher performance at lower cost.

+13.0 ppPerformance
58.6%Avg. cost reduction
Financial research: separate Kimi and DeepSeek memory comparisons. Logarithmic cost axis. Dashed arrows compare selected configurations. Kimi K3: 64.20% performance, $0.9930; Memory — Kimi K3 Max: 72.84% performance, $0.3169; DeepSeek V4 Flash Thinking: 53.01% performance, $0.0197; Memory — DeepSeek V4 Flash Thinking: 70.37% performance, $0.0100 Performance (%) ↑ 50 60 70 80 $0.01 $0.1 $1 Kimi K3: 64.20%, $0.9930 64.2% Kimi K3 Memory — Kimi K3 Max: 72.84%, $0.3169 72.8% Memory DeepSeek V4 Flash Thinking: 53.01%, $0.0197 53.0% DeepSeek V4 Flash Memory — DeepSeek V4 Flash Thinking: 70.37%, $0.0100 70.4% Memory Mean cost per task (USD, log scale)
  • Kimi K3 Max · DeepSeek V4 Flash Thinking
  • Same models + Voaige memory
Finance Agent V2

Join us

A unique place to work on the mission.

We bring together cognitive and systems neuroscientists, AI researchers, mathematicians, and engineers around a single shared challenge: to develop a coherent computational framework for cognition and memory, and test it in AI systems that push beyond the current frontier.

Join Us

Stay in the loop

Follow our research.