Advancing the state
of the art in AI
through neuroscience.

Voaige is an AI research lab translating insights from cognitive and systems neuroscience into computational principles for reasoning, memory, and search.

We are engineering these principles into LLMs under the constraint of building state-of-the-art,
commercially deployable systems.

Explore our research
Learned priors, memory and search, coordinated by cognition. Each outcome is kept as a memory that makes the next search smaller.

A radial search forest with cognition at its root. For three related problems in turn, learned priors narrow the search to one region, memory recalls the closest past route, and explicit search tries the branches that remain and selects a path. Each result is retained as a memory, so the second search tries far fewer branches than the first, and the third fewer still. The memories are then consolidated into learned priors.

Our research

From neuroscience to
high-performing AI systems.

We pursue two connected lines of research to build state-of-the-art,
commercially deployable systems.

Research in two phases: distill the principles, then engineer the systemPapers from six fields of neuroscience research fall in and fold into Tetris blocks as they cross a line, where each is judged. Most are rejected and pile up in the gutters; a few are kept as principles. Then the principles are fitted together in a well, the last completing four rows at once, and they fuse into one working system.NEUROSCIENCE RESEARCHDISTILL THE PRINCIPLESIDEAS →︎ PRINCIPLESENGINEER THE SYSTEMPRINCIPLES →︎ SYSTEM Research in two phases: distill the principles, then engineer the systemPapers from four fields of neuroscience research fall in and fold into Tetris blocks as they cross a line, where each is judged. Most are rejected and pile up in the gutters; a few are kept as principles. Then the principles are fitted together in a well, the last completing four rows at once, and they fuse into one working system.NEUROSCIENCE RESEARCHDISTILL THE PRINCIPLESIDEAS →︎ PRINCIPLESENGINEER THE SYSTEMPRINCIPLES →︎ SYSTEM

Distill the principles.

Draw on existing neuroscience research to identify computational mechanisms for learning, memory, and the coordination of search.

Engineer the systems.

Investigate the mechanics of large language models to turn those principles into algorithms and working systems.

Our Guiding Scientific Hypotheses

Search among possible pathsA search tree branches out into many candidates. A sage funnel highlights the promising region while the surrounding tree remains visible. Candidates within the funnel are evaluated and one path is selected in green.

Intelligence is search.

Intelligence is the process of searching among possible actions, explanations, and solutions in pursuit of a goal. It involves generating and evaluating candidates, validating them against evidence, and pursuing promising paths.

Coordinate learning, memory, and searchCognition makes three deliberate looks. Memory in an upper region provides a coral clue pointing right. Search in a lower region finds a green clue pointing back left. Both signals meet at the same junction, guiding the funnel to that region. Memory and search then work together to verify the answer.

Cognition coordinates search.

Cognition is the coordination of search over time. As a problem unfolds, it brings learned knowledge, memory, and exploration together to decide where to search, how to evaluate possibilities, and when to continue, change direction, or stop.

Recall a discovery to guide the next searchA first search spreads out in every direction before it finds a solution, and that route is kept as a memory. The next search reuses the remembered route to get most of the way, and only searches the part that remains unresolved.First searchNext search

Memory makes search more efficient.

Memory is the retention and retrieval of experience that allows previous computation to be reused. It can guide and narrow search, resume it from a useful earlier point, or supply a known solution that bypasses explicit search.

Get faster, better, and cheaper at searchFour new problems are solved one after another. The first search explores widely before it reaches its goal. With experience, a learned prior narrows where each later search looks, so it explores far less and reaches its goal sooner, by a more direct route.Experience

Learning is getting faster, better, and cheaper at search.

Learning is the process by which experience improves a system’s ability to search. It can improve the quality and reliability of solutions, expand the range of problems the system can solve, and reduce the time and computation required.

Read Intelligence Is Search

From hypothesis to system

Test Time
Cognition.

Our research, engineered into
the way AI reasons.

The TTC inference layer is what the agent interacts with. The model weights sit inside it, and TTC directs the inference. It is the site of inference-time computation: the cognitive architecture that governs how reasoning unfolds between input and output.

Explore Test Time Cognition
Test Time Cognition coordinates search inside the inference layerThe agent sends a task down into the Test Time Cognition layer, which contains the model. Inside it, cognition coordinates three ways of searching: the model's learned weights narrow where to look, memory recalls a route from a past episode, and a short new search resolves the rest. The answer goes back up to the agent, and the outcome is kept in memory as a new episode.AGENTPlanning · Tools · ContextTEST TIME COGNITIONMODELLEARNED WEIGHTSMEMORYPAST EPISODESTASKANSWERKEEP AS A NEW EPISODE

Our Research Applied

Higher intelligence at lower cost for the same model, same agent.

Test Time Cognition delivers gains across models, domains, and agent harnesses.

Without changing model weights or agent harness code.

tokens used for evals
307B
use cases evaluated on
3
agents tested
6

Customer Service

Test Time Cognition achieves SOTA performance.

55.3%Performance
Customer service: Test Time Cognition achieves 55.3% performance. Voaige baselines: Grok, Qwen, GPT-5.5. External reference: Opus 5 from Sierra. Logarithmic cost axis. Dashed arrows compare selected configurations. Qwen 3.8: 52.41% performance, $0.7594; Grok 4.6 High: 45.83% performance, $0.3949; Test Time Cognition: 55.33% performance, $1.8053; GPT-5.5 xHigh (Voaige baseline): 39.18% performance, $1.8166; Opus 5 Max (Sierra leaderboard): 48.71% performance, $13.3215 Performance (%) ↑ 40 50 60 $0.25 $1 $5 $20 Qwen 3.8: 52.41%, $0.7594 Qwen 3.8 Grok 4.6 High: 45.83%, $0.3949 Grok 4.6 Test Time Cognition: 55.33%, $1.8053 TTC GPT-5.5 xHigh (Voaige baseline): 39.18%, $1.8166 GPT-5.5 Opus 5 Max (Sierra leaderboard): 48.71%, $13.3215 Opus 5 Mean cost per task (USD, log scale)
  • Baseline · Opus 5 from Sierra
  • TTC: Qwen + Grok

TTC = Test Time Cognition

Tau Knowledge

Coding

Test Time Cognition achieves higher performance at lower cost.

+5.2 ppPerformance
38.5%Lower cost
Coding: GPT-5, MiniMax, and Test Time Cognition with verified mean costs across 89 tasks. Lines connect measured reasoning levels. Minimal: 26.97% performance, $0.1163; Low: 37.45% performance, $0.2815; GPT-5 High: 51.69% performance, $0.7722; MiniMax M2.5 Thinking: 46.82% performance, $0.0881; Test Time Cognition — GPT-5 low + MiniMax M2.5 Thinking: 54.31% performance, $0.3424; Test Time Cognition — GPT-5 medium + MiniMax M2.5 Thinking: 56.93% performance, $0.4747 Performance (%) ↑ 20 30 40 50 60 70 $0 $0.25 $0.5 $0.75 $1 Minimal: 26.97%, $0.1163 Minimal Low: 37.45%, $0.2815 Low GPT-5 High: 51.69%, $0.7722 GPT-5 High MiniMax M2.5 Thinking: 46.82%, $0.0881 MiniMax Test Time Cognition — GPT-5 low + MiniMax M2.5 Thinking: 54.31%, $0.3424 TTC Test Time Cognition — GPT-5 medium + MiniMax M2.5 Thinking: 56.93%, $0.4747 TTC Mean cost per task (USD)
  • Baseline: GPT-5 · MiniMax
  • TTC: GPT-5 + MiniMax

TTC = Test Time Cognition

Terminal-Bench 2.0 · mini-SWE-agent 2.2.1

Financial Research

Test Time Cognition achieves higher performance at lower cost.

+13.0 ppPerformance
58.6%Avg. cost reduction
Financial research: Kimi and DeepSeek with Test Time Cognition. Logarithmic cost axis. Dashed arrows compare selected configurations. Kimi K3: 64.20% performance, $0.9930; Test Time Cognition — Kimi K3 Max: 72.84% performance, $0.3169; DeepSeek V4 Flash Thinking: 53.01% performance, $0.0197; Test Time Cognition — DeepSeek V4 Flash Thinking: 70.37% performance, $0.0100 Performance (%) ↑ 50 60 70 80 $0.01 $0.1 $1 Kimi K3: 64.20%, $0.9930 Kimi K3 Test Time Cognition — Kimi K3 Max: 72.84%, $0.3169 TTC DeepSeek V4 Flash Thinking: 53.01%, $0.0197 DeepSeek V4 Flash Test Time Cognition — DeepSeek V4 Flash Thinking: 70.37%, $0.0100 TTC Mean cost per task (USD, log scale)
  • Baseline: Kimi K3 · DeepSeek V4 Flash
  • Same models + TTC

TTC = Test Time Cognition

Finance Agent V2

Join us

A unique place to work on the mission.

We bring together cognitive and systems neuroscientists, AI researchers, mathematicians, and engineers around a single shared challenge: to develop a coherent computational framework for cognition and memory, and test it in AI systems that push beyond the current frontier.

Join Us
Five fields, one problem, Next Gen AI Cognitive neuroscience, systems neuroscience, AI research, mathematics and engineering each draw a line in their own notation: a brain rhythm, spikes, network nodes, a wave and a clock signal. The lines merge and fall into a black hole, the O of the Voaige logo, and one line leaves it for Next Gen AI. COGNITIVE NEUROSCIENCE SYSTEMS NEUROSCIENCE AI RESEARCH MATHEMATICS ENGINEERING Next Gen AI

Stay in the loop

Follow our research.