Hello @vit0c72708961,
This is a great question and worth testing directly. The short answer is that it appears to be model dependent, and a simple test can help you determine how any given model handles non-English prompts.
A useful diagnostic is what I call the “elephant test”: ask the model to generate an image of an object with a label on it that says the word “elephant” — but write the prompt entirely in Arabic (صندوق عليه ملصق مكتوب عليه: فيل). If the label in the generated image renders the word in English, the model translated your entire Arabic prompt to English before processing. If the label renders in Arabic script, the model seems to have processed the Arabic natively.
Testing this across a few models in Firefly produced these results:
- Firefly native models appear to translate non-English prompts to English before processing, so the label renders “elephant” in English
- HUMAIN Image 1 appears to behave similarly, also rendering the word in English despite being designed for Arabic content
- Gemini Flash (Nano Banana 2) processed the Arabic prompt natively and rendered فيل correctly in Arabic script on the label (well, I do not read Arabic, but Claude told me that was elephant in Arabic)
For your specific use case — creating Saudi cultural stock images with accurate cultural detail — this suggests that if Arabic text rendering within images is important to you, Gemini Flash currently appears to handle this better than HUMAIN Image 1. For cultural subject matter accuracy more broadly, HUMAIN Image 1 may still produce more accurate results for Saudi-specific subjects such as traditional clothing, architecture, and food compared to standard Firefly models — but this would be due to its Saudi-specific training data rather than any advantage in prompt language processing, since the prompt appears to be translated to English regardless.
The practical recommendation would be to run your own elephant test with each model you are considering, as behaviour may change as these models are updated.
droopy