Skip to main content
POST

Overview

DeepSeek V4 Flash is DeepSeek’s speed-first Mixture-of-Experts model, engineered to deliver fast inference without sacrificing the quality you’d expect from a large-scale architecture. With 284 billion total parameters and only 13 billion activated per forward pass, V4 Flash achieves an exceptional throughput-to-quality ratio—making it ideal for latency-sensitive applications where you still need capable, coherent outputs. Its massive 1 million token context window sets it apart from most models in its class, enabling long-document analysis, extended multi-turn conversations, and large codebase comprehension in a single request. Routing DeepSeek V4 Flash through Vibetool gives you a clean, OpenAI-compatible interface with no vendor-specific SDK to learn. One API key, one endpoint, and a simple model slug are all you need to start sending requests. This makes it straightforward to benchmark V4 Flash against other models in your stack, swap it in for cost-sensitive workloads, or build pipelines that need to process very long inputs—like entire repositories, lengthy legal documents, or multi-session chat histories—without chunking or summarization hacks.
Model Slug: deepseek/deepseek-v4-flashUse this exact slug when making API requests to Vibetool.

Model Specifications

  • Architecture: Mixture-of-Experts (MoE)
  • Total Parameters: 284B
  • Activated Parameters: 13B
  • Context Window: 1,048,576 tokens (1M)
  • Input Modalities: Text
  • Output Modalities: Text

Pricing

Pricing: See vibetool.ai/pricing for current rates.

Use Cases

  • Long-Document Processing: Analyze entire books, legal contracts, or large codebases in a single request thanks to the 1M-token context window.
  • High-Throughput Pipelines: Power batch processing, classification, and summarization workflows where speed and cost efficiency are critical.
  • Extended Multi-Turn Assistants: Maintain coherent, context-rich conversations across very long sessions without losing earlier context.
  • Rapid Prototyping: Iterate quickly on prompts and features with fast response times during development and experimentation.

Authorizations

Authorization
string
header
required

Bearer authentication header of the form Bearer <token>, where <token> is your auth token.

Body

application/json
model
enum<string>
required
Available options:
deepseek-v4-flash
messages
object[]
required
stream
boolean
default:false
temperature
number
Required range: 0 <= x <= 2
max_tokens
integer

Response

Chat completion

id
string
object
string
Example:

"chat.completion"

model
string
choices
object[]
usage
object