Higher intelligence for less.Same model, same agent.

Cognition and memory deliver gains across open-source and closed-source models, domains, and agent harnesses.

Without changing model weights or agent harnesses.

307BTokens across evaluations
3Domains
6Different agents

Customer Service

Cognition achieves SOTA performance.

55.3% SOTAPerformance
Customer service: cognition achieves 55.3% performance. Voaige baselines: Grok, Qwen, GPT-5.5. External reference: Claude Opus 5 from Sierra. Logarithmic cost axis. Dashed arrows compare selected configurations. Qwen 3.8: 52.41% performance, $0.7594; Grok 4.6 High: 45.83% performance, $0.3949; Cognition: 55.33% performance, $1.8053; GPT-5.5 xHigh (Voaige baseline): 39.18% performance, $1.8166; Claude Opus 5 Max (Sierra leaderboard): 48.71% performance, $13.3215 Performance (%) ↑ 40 50 60 $0.25 $1 $5 $20 Qwen 3.8: 52.41%, $0.7594 52.4% Qwen 3.8 Grok 4.6 High: 45.83%, $0.3949 45.8% Grok 4.6 Cognition · SOTA: 55.33%, $1.8053 55.3% SOTA Cognition GPT-5.5 xHigh (Voaige baseline): 39.18%, $1.8166 39.2% GPT-5.5 Claude Opus 5 Max (Sierra leaderboard): 48.71%, $13.3215 48.7% Claude Opus 5 Mean cost per task (USD, log scale)
  • Qwen 3.8 Max · Grok 4.6 High · GPT-5.5 xHigh
  • Sierra: Claude Opus 5 Max
  • Voaige cognition: Qwen + Grok
Tau Knowledge

Coding

Cognition achieves higher performance at lower cost.

+5.2 ppPerformance
38.5%Lower cost
Coding: GPT-5, MiniMax, and cognition with verified mean costs across 89 tasks. Lines connect measured reasoning levels. Minimal: 26.97% performance, $0.1163; Low: 37.45% performance, $0.2815; GPT-5 High: 51.69% performance, $0.7722; MiniMax M2.5 Thinking: 46.82% performance, $0.0881; Cognition — GPT-5 low + MiniMax M2.5 Thinking: 54.31% performance, $0.3424; Cognition — GPT-5 medium + MiniMax M2.5 Thinking: 56.93% performance, $0.4747 Performance (%) ↑ 20 30 40 50 60 70 $0 $0.25 $0.5 $0.75 $1 Minimal: 26.97%, $0.1163 Minimal Low: 37.45%, $0.2815 Low GPT-5 High: 51.69%, $0.7722 51.7% GPT-5 High MiniMax M2.5 Thinking: 46.82%, $0.0881 MiniMax Cognition — GPT-5 low + MiniMax M2.5 Thinking: 54.31%, $0.3424 Cognition Cognition — GPT-5 medium + MiniMax M2.5 Thinking: 56.93%, $0.4747 56.9% Cognition Mean cost per task (USD)
  • GPT-5
  • MiniMax M2.5 Thinking
  • Voaige cognition: GPT-5 + MiniMax
Terminal-Bench 2.0 · mini-SWE-agent 2.2.1

Financial Research

Memory achieves higher performance at lower cost.

+13.0 ppPerformance
58.6%Avg. cost reduction
Financial research: separate Kimi and DeepSeek memory comparisons. Logarithmic cost axis. Dashed arrows compare selected configurations. Kimi K3: 64.20% performance, $0.9930; Memory — Kimi K3 Max: 72.84% performance, $0.3169; DeepSeek V4 Flash Thinking: 53.01% performance, $0.0197; Memory — DeepSeek V4 Flash Thinking: 70.37% performance, $0.0100 Performance (%) ↑ 50 60 70 80 $0.01 $0.1 $1 Kimi K3: 64.20%, $0.9930 64.2% Kimi K3 Memory — Kimi K3 Max: 72.84%, $0.3169 72.8% Memory DeepSeek V4 Flash Thinking: 53.01%, $0.0197 53.0% DeepSeek V4 Flash Memory — DeepSeek V4 Flash Thinking: 70.37%, $0.0100 70.4% Memory Mean cost per task (USD, log scale)
  • Kimi K3 Max · DeepSeek V4 Flash Thinking
  • Same models + Voaige memory
Finance Agent V2

Appendix

How we evaluate

We evaluate autonomous agents with room to work through a task, using the same evaluation protocol to compare baseline performance with cognition and memory.

Room to complete the task

We use generous time and step limits so agents can pursue longer solutions. The goal is to measure task-solving performance with enough room for the agent to finish, rather than have tight execution limits determine the outcome.

Repeated trials

Our standard protocol uses three samples per problem and configuration. Repeated attempts help account for variation between runs.

Framework and sandboxes

We use the Harbor framework to run and evaluate tasks, with execution environments hosted in Daytona sandboxes.