Raven Model Family · High-Velocity Frontier
Raven Flash
Frontier intelligence at production speed
Raven Flash is the high-velocity frontier model in the Raven family, engineered for low latency, responsive autonomous agents, and high-frequency code execution without compromising accuracy. It streams inspectable reasoning and completions in real time.
Specifications
| API model ID | raven-flash |
| Modalities | Text + Vision |
| Context window | 1M tokens |
| Output limit | 128K tokens |
| Tool use | Native function calling |
| Input price | $0.50 / 1M tokens ($0.05 cached) |
| Output price | undefined |
| Release date | October 1, 2026 |
| Reasoning controls | reasoning_effort low or high, default when omitted, with an inspectable reasoning stream |
Prices are per 1M tokens in USD. See the pricing page for plans and credit packs.
Key workloads
Autonomous Coding in DCode · Interactive Chat · CI/CD Pipelines · High-Volume Classification
Part of the family
Raven Max adds deep, long-horizon reasoning for decisions where precision is critical.
Meet Raven MaxEmpirical evaluation
Benchmark results, as published.
Evaluated with maximum reasoning effort enabled across standard industry benchmarks. Source: Dipole ML Technical Report (October 2026).
| Benchmark | Raven Flash | Rank | Best peer |
|---|---|---|---|
| Terminal Bench 4.0 | 34.1 | #1 | GPT-5.6 Terra · 24.6 |
| Agent's Last Exam | 26.6 | #2 | Claude Opus 4.8 · 27 |
| AutomationBench | 49.2 | #1 | Claude Opus 4.8 · 41 |
| GDPVal-AA | 1651 | #2 | Claude Opus 4.8 · 1731 |
| DeepSWE v1.1 | 65.4 | #2 | GPT-5.6 Terra · 69.6 |
| Intelligence Index | 43 | tied #1 | Claude Opus 4.8 · 43 |
| HLE w/ Tools | 55.8 | #2 | Claude Opus 4.8 · 57.9 |
| AA-Briefcase v1.1 | 1462 | #2 | Claude Opus 4.8 · 1480 |
Read the full evaluation in the Raven Flash performance report, or see See the whole family compared.
Put Raven Flashto work.
Available in the Raven app, DCode, and through the OpenAI-compatible DipoleML API.