NautilusBench v1

Opus 5.596%
GPT-6 Sol96%
GPT-6 Astra96%
GPT-5.6 Sol95%
GPT-5.6 Terra95%

Research · Announcements

NautilusBench: which AI model does trading research best

We graded 12 AI models on 77 closed trading research tasks in Chartnaut, re-run on market data they never saw. Three tied at 96%. GPT-6 Astra is the strongest overall, GPT-6 Sol is the best buy, and the most expensive model on the board is not the most accurate.

CJCJ11 min read

We gave 12 AI models the same 77 research tasks in Chartnaut and checked every number they reported against an answer we’d certified from the raw bars. Three models solved 74 of 77: Opus 5.5, GPT-6 Sol and GPT-6 Astra. If price doesn’t matter, GPT-6 Astra is the one to use. If it does, GPT-6 Sol gets the same score for 15 cents a solved task, about a third of what the other two cost.

Top solve rate

96.1%

74 of 77 tasks, three models tied

Best overall

GPT-6 Astra

top on hard tasks and code speed

Best value at the top

$0.15

per solved task, GPT-6 Sol

Hardest task

8%

one model in twelve solved it

What NautilusBench measures

Chartnaut’s agents write the indicators, definitions and studies traders run on their charts. General benchmarks can’t tell us which model does that well, because they don’t test the things that go wrong in trading code: an opening range defined in New York time on bars that arrive in UTC, a weekly level that mustn’t look at a week still in progress, a study whose sample quietly shrinks across a data gap.

So we built a benchmark out of those jobs. It has 77 tasks in three tiers. The 26 Foundation tasks are one script and one number. The 20 Hard tasks each have at least two traps that give a plausible wrong number. The 31 Expert tasks are full research chains: a definition, a run over history, a study over the events and statistics on top, sometimes across several runs that have to be pooled.

Every model runs in the same harness with the same tools, 20 minutes and the same token budget per task. A task counts as solved only if three things hold. The number the model reports matches the run it cites. Its saved script, re-run on instruments and dates it never saw, matches an answer we computed independently from raw market data. And the script passes Chartnaut’s validator. The tasks are closed and unpublished, so no model has seen them.

The leaderboard

Solve rate with its 95% intervalNautilusBench v1
30%40%50%60%70%80%90%100%2.0 min4.0 min6.0 min8.0 min10.0 minminutes per task →↑ tasks solvedOpus 5.5GPT-6 SolGPT-6 AstraGPT-5.6 SolGPT-5.6 TerraFable 5.1Sonnet 5GPT-5.6 LunaDeepSeek V4.1 FlashGLM 5.3 FlashGPT-6 LunaHaiku 4.5

Each vertical line is a 95% bootstrap interval over the 77 tasks. Where lines overlap, the benchmark cannot separate the models on solve rate alone. Plotted against time per task.

The top five sit within one task of each other and their intervals overlap almost completely. With 77 tasks, a single task is worth 1.3 points, so a gap of one or two tasks near the top is noise. The gaps further down are not. Sonnet 5 at 89.6% is clearly behind the leaders, and Haiku 4.5 at 48.1% is in a different class.

The best model, with price set aside

LeaderboardNautilusBench v1
#ModelSolvedTier-weightedPer solve
1Opus 5.5Anthropic96.1%94.3$0.45
1GPT-6 SolOpenAI96.1%95.0$0.15
1GPT-6 AstraOpenAI96.1%95.6$0.53
2GPT-5.6 SolOpenAI94.8%94.3$0.32
2GPT-5.6 TerraOpenAI94.8%93.7$0.17
3Fable 5.1Anthropic92.4%90.1$1.27
4Sonnet 5Anthropic89.6%86.8$0.53
5GPT-5.6 LunaOpenAI84.4%83.6$0.021
6DeepSeek V4.1 FlashDeepSeek80.5%74.2$0.050
7GLM 5.3 FlashZhipu AI72.7%64.8$0.12
8GPT-6 LunaOpenAI64.9%57.2$0.008
9Haiku 4.5Anthropic48.1%40.3$0.38

Tier-weighted counts Hard tasks twice and Expert three times, so easy tasks can’t carry a score. Per solve is the whole run at list price divided by tasks solved.

Three models tie on the headline number, so we break the tie on the things that matter when a trader relies on the output. Weight the harder tiers and GPT-6 Astra leads on 95.6, ahead of GPT-6 Sol on 95.0 and Opus 5.5 on 94.3. Astra and Opus both solved every Hard task. On Expert work Astra and Sol solved 94% and Opus 90%.

Then there’s the code. We time every script a model saves against our own reference. GPT-6 Astra’s ran at 1.04 times the reference, which is effectively as fast as ours. GPT-6 Sol’s took 1.51 times as long and Opus 5.5’s 1.62. Astra also finished tasks fastest, at 2.2 minutes each. Right answers, the hardest tasks handled and the quickest code make GPT-6 Astra the strongest model on NautilusBench.

The best model for the money

Solve rate against cost per solved taskNautilusBench v1
40%50%60%70%80%90%100%$0.005$0.010$0.020$0.050$0.10$0.20$0.50$1.00$2.00cost per solved task, log scale →↑ tasks solvedOpus 5.5GPT-6 SolGPT-6 AstraGPT-5.6 SolGPT-5.6 TerraFable 5.1Sonnet 5GPT-5.6 LunaDeepSeek V4.1 FlashGLM 5.3 FlashGPT-6 LunaHaiku 4.5

Cost is every token the model used across the 77 tasks at its lab’s list price, divided by the tasks it solved, so failed attempts count. The dashed line joins the models nothing beats on both axes.

Only three models sit on the efficient frontier, and between them they cover every budget. GPT-6 Sol is the one to take at the top: the same 96.1% as Astra and Opus at 15 cents a solve, against 53 and 45 cents. No model scores higher at any price, so paying more than GPT-6 Sol buys code speed, not accuracy.

Below it, GPT-5.6 Luna solved 84.4% for about 2 cents a task, the cheapest way to get a model that’s right more than four times in five. GPT-6 Luna is cheaper still at under a cent, but at 64.9% you’d be checking a third of its work. GPT-5.6 Terra is worth a mention off the frontier: 94.8% for 17 cents, with code as fast as Astra’s.

  • Best regardless of price: GPT-6 Astra, 96.1% and the strongest on hard tasks and code speed.
  • Best value at the top: GPT-6 Sol, 96.1% at $0.15 a solve.
  • Fast and nearly as accurate: GPT-5.6 Terra, 94.8% at $0.17, 2.2 minutes a task.
  • On a tight budget: GPT-5.6 Luna, 84.4% at about 2 cents.

The most expensive model is not the best one

Fable 5.1 is Anthropic’s top-priced model, at $10 per million input tokens and $50 per million output. That’s two and a half times Opus 5.5’s list price. On NautilusBench it solved 92.4% of the tasks graded so far, below Opus 5.5’s 96.1%, and every solve cost $1.27, nearly three times Opus’s 45 cents.

It isn’t a weak model. Fable solved every Hard task, the same as Opus and Astra. It loses ground on Expert work, 84% against Opus’s 90%, and it takes more tokens to get there: about 733 thousand a task, against Opus’s 494 thousand. On this kind of work the premium buys nothing we could measure.

Where the field comes apart

Solve rate by tierNautilusBench v1
20%30%40%50%60%70%80%90%100%FoundationHardExpertOpus 5.5 90%GPT-6 Astra 94%Sonnet 5 84%DeepSeek V4.1 Flash 61%GLM 5.3 Flash 48%Haiku 4.5 32%

Share of each tier solved: 26 Foundation, 20 Hard and 31 Expert tasks.

Foundation tasks barely separate anyone. Most models solve nine in ten of them, which is why a single easy benchmark would call them all roughly equal. The separation happens on Expert work. DeepSeek V4.1 Flash solved 96% of Foundation tasks and 61% of Expert ones. GLM 5.3 Flash went from 92% to 48%, Haiku 4.5 from 77% to 32%. The leaders barely move: GPT-6 Astra goes from 96% to 94%.

That shape says what the smaller models are good for. They write a clean indicator or a single definition reliably. What they lose is the thread across a long chain, where one wrong assumption early on makes every number after it wrong.

What still beats the best models

One task was solved by one model in twelve. It asks for a year of 1-minute breakouts. No single run holds a year of minute bars, so the model has to split the year into runs that meet exactly, then pool the event-level results itself rather than average the per-run figures. GPT-6 Astra managed it. The others ran out of time or reported an average of averages, which gets the median wrong.

The next hardest measures relative volume against the same minute on the previous six days, solved by 3 of 12. After that is a trade model with a partial exit and a stop that moves to breakeven, solved by 5. All three need state held correctly across thousands of bars, where a rule that’s nearly right gives a number that’s plainly wrong.

Multi-timeframe work splits even strong models. Sonnet 5 solved one of the four tasks that read a higher timeframe from a lower chart without peeking at an unfinished bar. The three leaders solved all four.

Right and fast are different skills

Solve rate against code speedNautilusBench v1
40%50%60%70%80%90%100%1×1.5×2×2.5×code run time against our reference →↑ tasks solvedour referenceOpus 5.5GPT-6 SolGPT-6 AstraGPT-5.6 SolGPT-5.6 TerraFable 5.1Sonnet 5GPT-5.6 LunaDeepSeek V4.1 FlashGLM 5.3 FlashGPT-6 LunaHaiku 4.5

Code speed is the run time of the model’s own saved scripts on the hidden data, against our reference scripts for the same tasks. The dashed vertical line is our reference.

Speed and accuracy barely correlate. Haiku 4.5 writes code as fast as our reference and solves half the tasks. Opus 5.5 solves 96% and its code takes 1.62 times as long. Only two models are both right and fast: GPT-6 Astra and GPT-5.6 Terra. On a live chart that’s the difference between an indicator that keeps pace with the market and one that lags it.

What we take from it

At the top the benchmark is close to saturated. Three models solve the same 74 tasks, so the next version needs more of the work only one model can do: long horizons, results pooled across runs, and state that has to survive gaps in the data.

The more useful finding is how little accuracy costs once you know where to look. GPT-6 Sol matches the most accurate models for a third of their price, and a 2-cent model solves more than four tasks in five. A higher price tag didn’t buy a better answer anywhere on this board.

A note from me

That’s the first full run, and there’ll be more as new models come out.

We built NautilusBench because our agents are only as good as the model behind them, and we wanted to stop guessing which one that should be. Now we measure it the same way every time, on tasks we know are hard.

For a manual trader the useful part is simple. A model that gets your study right on data it has never seen is one you can build on, and some of the best ones cost pennies.

Everything in this post, and the model picker, is on evals.chartnaut.com. Have a look and tell me what you’d want us to test next.

Thanks for reading.

Best,

CJCJFounder, Chartnaut

More from the blog

Keep reading