Qwen3-VL

Use Alibaba's Qwen3-VL vision-language model through our Serverless Cloud API

Qwen3-VL is Alibaba's vision-language model. It accepts an image and a text prompt and returns a text response. We support Qwen3-VL through our Serverless Cloud API, Dedicated Deployments, and self-hosted Inference.

Qwen3-VL API

1

Get your API Key

Create a Roboflow account, find your key on the Roboflow API settings page and make it available to your shell:

export ROBOFLOW_API_KEY="your-key-here"
2

Install the dependencies

Install the Inference SDK:

pip install -U inference-sdk supervision
3

Run the model

The sample prompts the qwen3vl-2b-instruct checkpoint to describe an image and prints the response.

import os
import supervision as sv
from inference_sdk import InferenceHTTPClient

image = sv.load_image_from_url("https://media.roboflow.com/quickstart/dog.jpeg")

client = InferenceHTTPClient(
    api_url="https://serverless.roboflow.com",
    api_key=os.environ["ROBOFLOW_API_KEY"],
)
result = client.infer_lmm(
    image,
    model_id="qwen3vl-2b-instruct",
    prompt="Describe this image briefly.",
    max_new_tokens=128,
)
print(result["response"])

The code above prints the model response to the terminal:

A man in a white t-shirt and red shorts is carrying a beagle dog on his shoulders. The dog is wearing a black harness and is looking forward. The man is walking on a paved path in a residential area with apartment buildings in the background. There is a small garden with green grass and white flowers to the left.

Qwen3-VL inference speed

Latency measured with Roboflow Inference on 1x NVIDIA L4, batch size 1, generating exactly 128 tokens with greedy decoding from a fixed prompt. Latency scales with output length, so use tokens/sec to estimate other lengths.

AliasLatency, 128 tokens (ms)Tokens/sec
qwen3vl-2b-instruct405732
qwen25-vl-7b560323

qwen25-vl-7b is the earlier Qwen2.5-VL checkpoint. It is listed here because it shares this alias namespace and runs through the same block.

Set api_url to match your deployment target:

  • https://serverless.roboflow.com for the Serverless Cloud API.
  • http://localhost:9001 for a local Inference server.
  • Your Dedicated Deployment URL for a private endpoint.

You can train your own Qwen3-VL checkpoint on Roboflow and call it by its per-model {workspace}/{model-slug} ID (see Versions, Trainings, and Models).