Picking the right model for image work
When to reach for text-to-image vs image-to-image, and what each of the curated models is actually good at.
Two image modes are live on VideoGen today, and the question most beginners get wrong is which to pick.
Text-to-Image (T2I)generates a fresh image from a prompt alone. Use it when you have an idea but no reference image — “cinematic portrait of an architect in a sunlit studio”. Two curated models cover the bulk of T2I work: ideogram-v3-balanced for crisp typography and bold compositions, and nano-banana-pro/text-to-imagewhen you want Google’s flagship quality and prompt adherence. Knobs worth knowing: aspect ratio, resolution, and seed when you want to iterate on a specific result.
Image-to-Image (I2I) takes a reference image and a text instruction, and produces a new image that follows the instruction while keeping the subject. Use it for restyles, background swaps, and color grading. The curated model is nano-banana-pro/edit — Gemini 3.0 Pro Image under the hood, up to 4K, and strong identity preservation. The most useful knob is the prompt itself: describe what should change, not what should stay.
Both render from the same studio surface. If you find yourself describing “a photo I already have, but changed”, that’s i2i — paste it in instead of trying to describe it. Live catalogue on /models.