- Alibaba launched Qwen-Image-3.0 on July 21, featuring 4.5k token input and one-shot complex layouts.
- Early tests indicate mixed performance in Japanese text rendering, with some awkward translations noted.
- API trials are now open on Alibaba Cloud Bailian and the Qwen AI Platform.
- The previous version, Qwen-Image-2.0, supported approximately 1,000 tokens, making the new version's input capacity roughly 4.5 times larger.
- Qwen-Image-3.0 is designed to enhance the understanding of complex prompts and improve the precision of text-and-image layouts.
Alibaba's Qwen-Image-3.0, launched on July 21, represents a significant leap in AI image generation, supporting ultra-long prompts of up to 4,500 tokens. This enhancement allows for complex layouts and multilingual outputs, including 12 languages and over 20 fonts.15
Despite its advanced capabilities, early evaluations indicate that the model struggles with Japanese text rendering. For instance, while generating a fictional blog post about ramen, the output contained awkward translations, highlighting ongoing challenges in accurately depicting Japanese characters.2
The model excels in generating realistic images and can produce intricate designs such as multi-layered UI interfaces and knowledge infographics. It is currently ranked as the top model in China for text-to-image evaluations, just below GPT Image 2.

However, the launch has been met with scrutiny due to the absence of benchmark scores and technical documentation, which raises concerns about the verifiability of its claimed improvements. Tech media outlets have noted that this lack of transparency could hinder trust in the model's performance.
API trials are now open, potentially reducing production costs in advertising and creative design, but the model's effectiveness in real-world applications remains to be fully assessed.
As Alibaba continues to innovate, the balance between advanced features and practical usability will be crucial for the success of Qwen-Image-3.0.
“Early tests show Qwen-Image-3.0 struggles with Japanese text despite supporting 12 languages. The model is currently top-ranked in China below GPT Image 2, but Alibaba did not release benchmark scores or a technical report, and API trials are now open.”
