qwen

Qwen: Qwen2.5 VL 32B Instruct

Qwen2.5-VL-32B is a multimodal vision-language model fine-tuned through reinforcement learning for enhanced mathematical reasoning, structured outputs, and visual problem-solving capabilities. It excels at visual analysis tasks, including object recognition, textual interpretation within images, and precise event localization in extended videos. Qwen2.5-VL-32B demonstrates state-of-the-art performance across multimodal benchmarks such as MMMU, MathVista, and VideoMME, while maintaining strong reasoning and clarity in text-based tasks like MMLU, mathematical problem-solving, and code generation.

Input Cost
$0.05
per 1M tokens
Output Cost
$0.22
per 1M tokens
Context Window
16,384
tokens
Compare vs GPT-4o
Developer ID: qwen/qwen2.5-vl-32b-instruct

Related Models

qwen
$0.40/1M

Qwen: Qwen Plus 0728 (thinking)

Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoni...

📝 1,000,000 ctx Compare →
qwen
$0.20/1M

Qwen: Qwen2.5-VL 7B Instruct

Qwen2.5 VL 7B is a multimodal LLM from the Qwen Team with the following key enhancements: ...

📝 32,768 ctx Compare →
qwen
$1.20/1M

Qwen: Qwen3 Max

Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in ...

📝 256,000 ctx Compare →