← All research
Research

Raven Flash Performance Report: Sub-Second Velocity and #1 on AutomationBench

Raven Flash takes #1 on AutomationBench (49.2), leads the velocity tier on Terminal Bench 4.0 (34.1), and ties #1 on the Intelligence Index, at $0.50 per 1M input tokens.

Raven Flash is our high-velocity frontier model: engineered for low latency, responsive autonomous agents, and high-frequency code execution without compromising accuracy. This report presents its empirical evaluation, run with maximum reasoning effort enabled, against the current frontier peers.

#1 on AutomationBench

On AutomationBench, which measures end-to-end autonomous task completion, Raven Flash scores 49.2, outperforming Claude Opus 4.8 by 8.2 points:

Raven Flash
49.2
Claude Opus 4.8
41
DeepSeek-V4-Vision
38.8
Gemini 3.7 Flash
38.6
GPT-5.6 Terra
37.2

#1 in the velocity tier on Terminal Bench 4.0

On Terminal Bench 4.0, Raven Flash scores 34.1 in the velocity tier, ahead of GPT-5.6 Terra at 24.6 and Claude Opus 4.8 at 21.7:

Raven Flash
34.1
GPT-5.6 Terra
24.6
Claude Opus 4.8
21.7
DeepSeek-V4-Vision
18.2
Gemini 3.7 Flash
17.5

Headline results

BenchmarkRaven FlashBest peerMargin
AutomationBench49.2 (#1)Claude Opus 4.8 · 41.0+8.2
Terminal Bench 4.0 (velocity tier)34.1 (#1)GPT-5.6 Terra · 24.6+9.5
Intelligence Index43 (tied #1)Claude Opus 4.8 · 43tied

Why Flash wins on velocity

Raven Flash shares the Raven family's native tool execution and inspectable reasoning streams, tuned for sub-second first tokens. Its advantage shows in agentic loops: plan, call a tool, read the result, revise, and continue, without the model losing the thread between steps. On high-frequency workloads, that pacing compounds into the margins above.

Production economics

Velocity is only useful if it is affordable. Raven Flash bills $0.50 per 1M input tokens, $0.05 per 1M cached input, and $2.50 per 1M output tokens, with automatic prompt caching discounts applied at the gateway. Full pricing is on the pricing page.

All evaluations run with maximum reasoning effort enabled. Full methodology and additional benchmarks appear on the models page.

Get started

Try what we just shipped.

The real models, free in your browser, or download DCode and put them to work.