Qwen3-VL is Alibaba's vision-language model. It accepts an image and a text prompt and returns a text response. We support Qwen3-VL through our Serverless Cloud API, Dedicated Deployments, and self-hosted Inference.
Qwen3-VL API
Get your API Key
Create a Roboflow account, find your key on the Roboflow API settings page and make it available to your shell:
export ROBOFLOW_API_KEY="your-key-here"Install the dependencies
Install the Inference SDK:
pip install -U inference-sdk supervisionRun the model
The sample prompts the qwen3vl-2b-instruct checkpoint to describe an image and prints the response.
import os
import supervision as sv
from inference_sdk import InferenceHTTPClient
image = sv.load_image_from_url("https://media.roboflow.com/quickstart/dog.jpeg")
client = InferenceHTTPClient(
api_url="https://serverless.roboflow.com",
api_key=os.environ["ROBOFLOW_API_KEY"],
)
result = client.infer_lmm(
image,
model_id="qwen3vl-2b-instruct",
prompt="Describe this image briefly.",
max_new_tokens=128,
)
print(result["response"])The code above prints the model response to the terminal:
A man in a white t-shirt and red shorts is carrying a beagle dog on his shoulders. The dog is wearing a black harness and is looking forward. The man is walking on a paved path in a residential area with apartment buildings in the background. There is a small garden with green grass and white flowers to the left.
Qwen3-VL inference speed
Latency measured with Roboflow Inference on 1x NVIDIA L4, batch size 1, generating exactly 128 tokens with greedy decoding from a fixed prompt. Latency scales with output length, so use tokens/sec to estimate other lengths.
| Alias | Latency, 128 tokens (ms) | Tokens/sec |
|---|---|---|
qwen3vl-2b-instruct | 4057 | 32 |
qwen25-vl-7b | 5603 | 23 |
qwen25-vl-7b is the earlier Qwen2.5-VL checkpoint. It is listed here because it shares this alias namespace and runs through the same block.
Set api_url to match your deployment target:
https://serverless.roboflow.comfor the Serverless Cloud API.http://localhost:9001for a local Inference server.- Your Dedicated Deployment URL for a private endpoint.
You can train your own Qwen3-VL checkpoint on Roboflow and call it by its per-model {workspace}/{model-slug} ID (see Versions, Trainings, and Models).