Vision · Jan 2025

Qwen2.5-VL 72B

Proven vision-language model for OCR, document QA, and visual grounding.

By Alibaba · Text + Vision · Apache 2.0

Context

128K tokens

Modality

Text + Vision

Architecture

Dense VL

License

Apache 2.0

Per 1M tokens $0.80 in · $1.00 out
Model details

Specs & substance

Developed by
Alibaba
Model family
Qwen
Use case
Vision
Modality
Text + Vision
Context window
128K tokens
Architecture
Dense VL
Version
VL 72B
License
Apache 2.0
Pricing
$0.80 in · $1.00 out · $0.40 cache read
Released
Jan 2025
Endpoint
parasail-qwen25-vl-72b-instruct

Proven vision-language model for OCR, document QA, and visual grounding.

Qwen2.5-VL 72B is part of the Qwen family by Alibaba, and sits in the vision category. It supports a 128K tokens context window and is built on a Dense VL architecture.

Key strengths: vision, document. Parasail serves Qwen2.5-VL 72B on a global fleet of current-gen GPUs behind a single OpenAI-compatible endpoint — with per-token pricing, no minimums, and dedicated capacity options when you need guaranteed throughput.

Integrate

Drop-in via the OpenAI SDK

Point any OpenAI-compatible client at Parasail and change the model name. That's it.

python parasail · Qwen2.5-VL 72B
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.parasail.io/v1"
)

response = client.chat.completions.create(
    model="parasail-qwen25-vl-72b-instruct",
    messages=[
        {"role": "user", "content": "Hello, what can you do?"}
    ],
    stream=True,
    max_tokens=1000
)

for chunk in response:
    if chunk.choices and chunk.choices[0].delta.content is not None:
        print(chunk.choices[0].delta.content, end="", flush=True)
More models

Explore the library

All models

Start building today

Instantly run any open model — popular or specialized.