Compact · Sep 2026

DeepSeek V4.1 Flash

Multimodal 552B-parameter MoE with 1M-token context, compressed KV cache, and efficient agentic reasoning.

By DeepSeek · Text + Vision · MIT

Context

1M tokens

Modality

Text + Vision

Architecture

CED MoE + CSA2

License

MIT

Per 1M tokens $0.30 in · $1.20 out
Model details

Specs & substance

Developed by
DeepSeek
Model family
DeepSeek
Use case
Compact
Modality
Text + Vision
Context window
1M tokens
Architecture
CED MoE + CSA2
Version
V4.1 Flash
License
MIT
Pricing
$0.30 in · $1.20 out · $0.006 cache read
Released
Sep 2026
Endpoint
parasail-deepseek-v41-flash

Multimodal 552B-parameter MoE with 1M-token context, compressed KV cache, and efficient agentic reasoning.

DeepSeek V4.1 Flash is part of the DeepSeek family by DeepSeek, and sits in the compact category. It supports a 1M tokens context window and is built on a CED MoE + CSA2 architecture.

Key strengths: efficient, agentic, multimodal, long-context. Parasail serves DeepSeek V4.1 Flash on a global fleet of current-gen GPUs behind a single OpenAI-compatible endpoint — with per-token pricing, no minimums, and dedicated capacity options when you need guaranteed throughput.

Integrate

Drop-in via the OpenAI SDK

Point any OpenAI-compatible client at Parasail and change the model name. That's it.

python parasail · DeepSeek V4.1 Flash
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.parasail.io/v1"
)

response = client.chat.completions.create(
    model="parasail-deepseek-v41-flash",
    messages=[
        {"role": "user", "content": "Hello, what can you do?"}
    ],
    stream=True,
    max_tokens=1000
)

for chunk in response:
    if chunk.choices and chunk.choices[0].delta.content is not None:
        print(chunk.choices[0].delta.content, end="", flush=True)
More models

Explore the library

All models

Start building today

Instantly run any open model — popular or specialized.