Raven Model Family · High-Velocity Frontier

Raven Flash

Frontier intelligence at production speed

Raven Flash is the high-velocity frontier model in the Raven family, engineered for low latency, responsive autonomous agents, and high-frequency code execution without compromising accuracy. It streams inspectable reasoning and completions in real time.

Specifications

API model IDraven-flash
ModalitiesText + Vision
Context window1M tokens
Output limit128K tokens
Tool useNative function calling
Input price$0.50 / 1M tokens ($0.05 cached)
Output priceundefined
Release dateOctober 1, 2026
Reasoning controlsreasoning_effort low or high, default when omitted, with an inspectable reasoning stream

Prices are per 1M tokens in USD. See the pricing page for plans and credit packs.

Key workloads

Autonomous Coding in DCode · Interactive Chat · CI/CD Pipelines · High-Volume Classification

Part of the family

Raven Max adds deep, long-horizon reasoning for decisions where precision is critical.

Meet Raven Max

Empirical evaluation

Benchmark results, as published.

Evaluated with maximum reasoning effort enabled across standard industry benchmarks. Source: Dipole ML Technical Report (October 2026).

BenchmarkRaven FlashRankBest peer
Terminal Bench 4.034.1#1GPT-5.6 Terra · 24.6
Agent's Last Exam26.6#2Claude Opus 4.8 · 27
AutomationBench49.2#1Claude Opus 4.8 · 41
GDPVal-AA1651#2Claude Opus 4.8 · 1731
DeepSWE v1.165.4#2GPT-5.6 Terra · 69.6
Intelligence Index43tied #1Claude Opus 4.8 · 43
HLE w/ Tools55.8#2Claude Opus 4.8 · 57.9
AA-Briefcase v1.11462#2Claude Opus 4.8 · 1480

Read the full evaluation in the Raven Flash performance report, or see See the whole family compared.

Put Raven Flashto work.

Available in the Raven app, DCode, and through the OpenAI-compatible DipoleML API.