Images: picking a model and writing a prompt
How to pick a model for the job, describe a frame so the model gets it on the first try, and edit your own photo. With examples of weak and strong prompts and ready-made wordings.
Updated:
How to pick a model
The studio has dozens of image models, each with its own character. The full list with prices is in the catalogue, and every model has its own page with features and tips. Below is where each model is strongest.
- Nano Banana 2 — the workhorse: fast, close to Pro level, readable text in the image, 1K to 4K, handles complex prompts and keeps one character consistent across a series.
- Nano Banana Pro — for final assets: logos, banners, layouts with accurate small text, complex scenes that keep a style and characters from references.
- GPT Image — understands prompts well and writes text in images, including Cyrillic, and edits uploaded photos neatly.
- Seedream — long prompts with many conditions, photo edits that keep the face and key objects.
- Flux — follows detailed instructions precisely, photorealism and product shots.
- Grok Imagine — a recognisable aesthetic of its own, handy for trying out ideas.
- Qwen Image — simple generation and edits: remove an object, swap the background, fix light and colour.
- Ideogram — covers and promo images with lettering; Ideogram Character keeps your appearance from a photo.
- Recraft and Topaz — utility models: remove the background and upscale a finished image.
If you are not sure where to start, take Nano Banana 2: it is fast, inexpensive and handles most jobs. For Cyrillic lettering use GPT Image 2, Nano Banana 2 or Nano Banana Pro — the studio warns you when the selected model may distort Russian text.
From text and from a photo
There are two ways to get an image. From text — you describe the frame in words and the model draws it from scratch. From a photo — you add one or more images to the References block and the model builds on them: it takes a face, an object, a style or a composition.
- How many references you can add depends on the model: one image for some, more than ten for others. The limit is on the model page, and the studio warns you and skips extra files.
- Some models require a reference — these are editing models. Others work from text only and do not accept references; then the References block is unavailable.
- Send a finished result straight to references with the To references button and keep working on it with the next prompt.
- Write in any language; there is no need to translate your prompt.
Editing your own photos
An edit is the same prompt, just with your photo in the references. Do not write “make it better” — say exactly what to change and what to keep.
- Background: “replace the background with a light studio and a soft shadow”.
- Clutter: “remove the people in the background and the cables on the desk”.
- Light and colour: “make the light warm, like at sunset, and the colours richer”.
- Marketplace product photo: “put the product on a white background, soft studio light, a shadow below, no lettering”.
- What to keep: add “do not change the face”, “keep the shape and the packaging text as in the photo”.
Models don't create images with a transparent background. If you need one, generate the image on a plain background, then remove the background with the Recraft background removal model.
How to write a prompt
A universal structure that works for most tasks: subject → action → environment → style → technical parameters. Each block answers a specific question:
- Subject — who or what is in the frame
- Action — what is happening, and in what pose
- Environment — where it takes place and what surrounds it
- Style — photorealism, illustration, 3D, a specific technique
- Technical parameters — lighting, camera angle, lens, color palette
A young woman in a white doctor's coat sits at a desk, holding a tablet and reviewing patient data. A modern medical office with large windows, soft natural light from the left. Realistic photograph, 50mm lens, f/2.8 depth of field, neutral color palette.
Weak prompt
A bicycle
Strong prompt
A red vintage bicycle with a wicker basket of lemons leaning against a turquoise wall in a narrow sunny alley, hard morning light and long shadows, bougainvillea above the wall, 35mm street photography, eye level
The second prompt names the subject, the place, the light, the details and the kind of shot, so the model has almost nothing to guess. With the first one it picks everything itself, and the result is random.
Practical tips
Quick answers for common situations — what belongs in the prompt and what is better set in the request parameters.
If you are editing a photo
Don't write “make it better” — say exactly what should change: “remove the people in the background”, “replace the background with an office”, “keep the face unchanged”.
If you uploaded several references
Specify what to take from each one: “take the face from photo 1, the outfit from photo 2, the background from photo 3”.
If you need text in the image
Write the exact phrase in quotes and specify the language. For text in images, GPT Image 2, Nano Banana 2, and Nano Banana Pro work best.
If you need a transparent background
Models cannot generate transparency directly. Ask for a solid-color background — for example, plain black — and then remove it with Recraft Remove Background. You will get a clean PNG with transparency.
If you want to set the resolution or format
Don't write “make it 4K” or “3:2 aspect ratio” in the prompt — the model will ignore it. Resolution and aspect ratio are chosen in the request parameters before generation.
If you don't like the result
Change one parameter at a time: first the style, then the lighting, then the camera angle. Otherwise it will be hard to tell what actually improved the image.
Composition and framing
Controlling composition is the most underrated lever in prompting. Specify:
- Shot size: close-up, medium, wide
- Angle: eye level, from above, from below, isometric
- Focal length: 24mm for wide scenes, 50mm for natural perspective, 85mm+ for portraits
- Depth of field: f/1.4 for strong bokeh, f/8 for a sharp wide shot
- Composition techniques: rule of thirds, centered composition, leading lines
Style and aesthetics
It is better to describe style through technique and characteristics than through the names of artists or brands. Technical terms are more universal and predictable:
- “Watercolor illustration with soft color transitions”
- “Isometric 3D graphics, flat shadows, limited palette”
- “Cinematic photography, warm palette, subtle grain”
- “Technical drawing, thin lines, monochrome”
If you need to convey a specific mood, describe it with adjectives: “cozy”, “austere”, “dynamic”, “contemplative”. Models understand these words just as well as technical terms.
Working with references
Reference images are a shortcut to the result you want when it is hard to describe in words. A few practical rules:
- 1 reference — to hold on to a single aspect: composition, color, or a character.
- 2–4 references — for a series of related frames with a unified style.
- 5–10 references — for fine-tuning the aesthetics, when averaging the character matters.
- Don't mix fundamentally different styles. If one reference is minimalist and another is packed with baroque detail, the model will average them into a mediocre result.
Negatives and exclusions
Not all models handle “don't do this” instructions equally well, so phrase your requirements as affirmative statements:
- Instead of “no people” — “empty space”
- Instead of “not blurry” — “sharp image, high detail”
- Instead of “not bright” — “muted color palette”
Aspect ratios, resolution and settings
Besides the prompt itself, every model has a set of parameters configured before you run it: resolution, aspect ratio, output file type, and a number of model-specific options. There is no need to put them in the prompt — the model will ignore them.
Resolution: 1K, 2K, 4K
Resolution controls the level of detail in the image.
- 1K — faster and cheaper. Good for tests, ideas, social media, and quick drafts.
- 2K — the sweet spot. Noticeably more detail, a solid choice for most tasks.
- 4K — maximum quality, but usually more expensive and slower. Needed for print, banners, commercial layouts, and complex scenes.
If in doubt, start with 1K or 2K and raise the resolution once the content itself is right.
Image format
Format here means the aspect ratio, not the file type.
- 1:1 — square. For avatars, cards, and posts.
- 4:5 — vertical social media post.
- 16:9 — wide frame: banners, presentations, covers, video thumbnails.
- 9:16 — Stories, Reels, Shorts.
- 3:4 and 4:3 — more classic photo proportions.
- 21:9 — ultrawide frame for cinematic compositions and headers.
Pick the format for the task up front, so you don't have to crop important parts of the image later.
Output file type
- PNG — usually heavier, but cleaner; better for graphics, text, interfaces, and cases where precision matters.
- JPG / JPEG — lighter files, a good fit for regular photos and quick publishing.
- MP4 — used by video models when the output is a clip rather than an image.
Models don't create images with a transparent background. If you need one, generate the image on a plain background, then remove the background with the Recraft background removal model.
Additional settings
Not every model has these parameters — the exact set depends on the model you choose.
Quality (medium / high). Available in GPT Image, for example. Medium is for quick iterations and drafts, high is for final material. Affects the generation cost.
Google Search. Available in Nano Banana 2 and Seedance. Pulls up-to-date information from the web straight into the request context — useful for generations about recent topics, events, and real-world objects.
NSFW Checker. An extra check of the result for sensitive content. Turn it on when the output goes to public channels or materials for a broad audience.
Common mistakes
A prompt that is too short. A request like “a beautiful office” leaves the model complete freedom of interpretation, so the result will be random — different every time. A minimum viable prompt contains at least a subject, an environment, and a style.
An overloaded prompt. If a request contains dozens of requirements, the model will start ignoring some of them. A useful rule: one request — one scene. If you need several related frames, run a series of requests with the same structure.
Contradictory instructions. “A minimalist composition with lots of detail”, “a photograph in watercolor style” — combinations like these produce unpredictable results. It is better to explicitly pick one approach and describe it in detail.
Templates for common tasks
[Product] on a [neutral/contrasting] background, [studio/natural] lighting [from a specific side]. Clean composition, minimal distracting elements, focus on material and texture. Realistic photograph, 50mm lens, f/8, 4K resolution.
Conceptual illustration on the topic of [article topic]. [Style: flat graphics / watercolor / engraving]. Limited palette of [3–4 colors]. Centered composition, simple background. 16:9 aspect ratio.
[Scene/subject]. Modern digital photography, distinct character, contrasty lighting. 4:5 aspect ratio for the feed. Composition with headroom at the top for text.
[Metaphor/image for the topic], minimalist visual language, neutral palette, a consistent look across the slide series. 16:9 aspect ratio, free space on the right side of the frame.
Ready-made prompts with sample results are in templates: open a template in the studio and the model, the aspect ratio and the prompt are filled in for you.
Iterative workflow
Getting a perfect result on the first try is the exception rather than the rule. Work in iterations:
- Run 3–4 draft requests with different wordings
- Pick the one closest to what you need
- Refine the details in the next request and add a reference
- Keep the prompt that worked so you don't have to rebuild it from scratch
The service keeps your full prompt history, so you can return to wordings that worked and use them as a starting point for new tasks.
How to make an image in the studio
- Open the studio. Log in to the web studio with your email. A new user gets 2 free tokens. You can use them to try creating an image or chatting with AI.
- Pick a model. Choose an image model in the Model field. It decides the aspect ratios, resolutions, number of references and the price.
- Describe the frame. In the Prompt field, write who or what is in the frame, what is happening, where, in what style and in what light. Need lettering? Put it in quotes and name the language.
- Add references. To build on your own photo or a style, add images to the References block and say what to take from each.
- Choose the aspect ratio and resolution. Aspect ratio and resolution are set in the settings, not in the prompt. You see the price in tokens before you run.
- Run and save. Press Run. Download the finished image, send it to references for the next step, or find it later in History.