How to choose & budget for generation models?
Unlike text models billed per token, generation models bill three ways: images per image, video per second, speech per character. Within a category, price also shifts with resolution, duration and voice tier — so pick the category first, then compare inside it.
Chinese models are strong here: Wan, Doubao Seedream, Kling, Dreamina for image/video are usually a tier cheaper than global options and handle Chinese well; global Veo, Sora, Imagen still lead on quality and duration. For TTS, iFlytek, Doubao, CosyVoice sound natural in Chinese, while ElevenLabs leads on multilingual voices.
Per-call prices look small, but at volume (thousands of images / long video / long voiceover) they add up fast. To estimate total spend, reuse the approach in the cost calculator.














