Specs & substance
- Developed by
- Alibaba
- Model family
- Qwen
- Use case
- Vision
- Modality
- Text + Vision
- Context window
- 128K tokens
- Architecture
- Dense VL
- Version
- VL 72B
- License
- Apache 2.0
- Pricing
- $0.80 in · $1.00 out · $0.40 cache read
- Released
- Jan 2025
- Endpoint
- parasail-qwen25-vl-72b-instruct
Proven vision-language model for OCR, document QA, and visual grounding.
Qwen2.5-VL 72B is part of the Qwen family by Alibaba, and sits in the vision category. It supports a 128K tokens context window and is built on a Dense VL architecture.
Key strengths: vision, document. Parasail serves Qwen2.5-VL 72B on a global fleet of current-gen GPUs behind a single OpenAI-compatible endpoint — with per-token pricing, no minimums, and dedicated capacity options when you need guaranteed throughput.
Drop-in via the OpenAI SDK
Point any OpenAI-compatible client at Parasail and change the model name. That's it.
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://api.parasail.io/v1"
)
response = client.chat.completions.create(
model="parasail-qwen25-vl-72b-instruct",
messages=[
{"role": "user", "content": "Hello, what can you do?"}
],
stream=True,
max_tokens=1000
)
for chunk in response:
if chunk.choices and chunk.choices[0].delta.content is not None:
print(chunk.choices[0].delta.content, end="", flush=True) curl https://api.parasail.io/v1/chat/completions \
-H "Authorization: Bearer $PARASAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "parasail-qwen25-vl-72b-instruct",
"messages": [{"role": "user", "content": "Hello, what can you do?"}],
"stream": true,
"max_tokens": 1000
}' Explore the library
Qwen3-VL 235B-A22B
Large vision-language model with strong document understanding and visual reasoning.
Qwen3-VL 8B
Compact vision-language model — fast and affordable for image understanding at scale.
Qwen3.5 397B-A17B
Large-scale MoE model with frontier reasoning and multilingual capabilities.
Qwen3.6 35B-A3B
Efficient MoE model with 3B active parameters — great price-to-performance.
Start building today
Instantly run any open model — popular or specialized.