Higher intelligence for less.Same model, same agent.
Cognition and memory deliver gains acrossopen-source and closed-source models, domains, and agent harnesses.
Without changing model weights or agent harnesses.
307BTokens across evaluations
3Domains
6Different agents
Customer Service
Cognition achieves SOTA performance.
55.3% SOTAPerformance
Qwen 3.8 Max · Grok 4.6 High · GPT-5.5 xHigh
Sierra: Claude Opus 5 Max
Voaige cognition: Qwen + Grok
Tau Knowledge
Coding
Cognition achieves higher performance at lower cost.
+5.2 ppPerformance
38.5%Lower cost
GPT-5
MiniMax M2.5 Thinking
Voaige cognition: GPT-5 + MiniMax
Terminal-Bench 2.0 · mini-SWE-agent 2.2.1
Financial Research
Memory achieves higher performance at lower cost.
+13.0 ppPerformance
58.6%Avg. cost reduction
Kimi K3 Max · DeepSeek V4 Flash Thinking
Same models + Voaige memory
Finance Agent V2
Appendix
How we evaluate
We evaluate autonomous agents with room to work through a task, using the same evaluation protocol to compare baseline performance with cognition and memory.
Room to complete the task
We use generous time and step limits so agents can pursue longer solutions. The goal is to measure task-solving performance with enough room for the agent to finish, rather than have tight execution limits determine the outcome.
Repeated trials
Our standard protocol uses three samples per problem and configuration. Repeated attempts help account for variation between runs.
Framework and sandboxes
We use the Harbor framework to run and evaluate tasks, with execution environments hosted in Daytona sandboxes.