Qwen3 VL 8B Instruct
Alibaba · 1 configuration
Intelligence—
Input / 1M$0.12
Output / 1M$0.45
Context262.1K
Tariff: OpenRouter · Source: openrouter · USD per million tokens
Thinking configuration
What does more thinking buy you?
Compare this model’s measured thinking configurations.
◌
No pair of measurements available
Change the filters or pick another benchmark. Missing values are not estimated.
Selected: thinking default · Click a point to switch configuration
Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon...
This model is in the catalog but has no imported benchmarks. Capabilities are not estimated.
Specifications from the catalog
Inputimage, text
Outputtext
Tool callingYes
Structured outputYes
SpeedNot available
First tokenNot available
Maximum output32.8K
Catalog listing date2025-10-14