Обзор Ideogram 4.0: открытые веса, ComfyUI, подсказки, API и ограничения

Обзор Ideogram 4.0: открытые веса, ComfyUI, подсказки, API и ограничения
Tested August 14, 2026

Быстрый ответ: Ideogram 4.0 is a 9.3B open-weight image model built for typography, structured layouts, and native 2K generation. In our hosted tests it handled bilingual poster text and materials well, but missed exact-count details and ignored the requested orientation once. Public weights are non-commercial; production requires the hosted API or separate commercial rights.

Ideogram 4.0 arrives with a rare combination: downloadable weights, a design-first prompt format, and a hosted API that can be used commercially. That makes it more interesting than a routine image-model update—but also easier to misunderstand.

We checked the official repository, prompting and inference guides, API reference, license, and pricing page. We also ran bilingual typography and product-photography tasks through Broly Anywhere, then compared the same two prompts with GPT Image 2 and Nano Banana 2 baselines. This is a practical review, not a claim that one image wins every category.

If you are surveying the wider market first, our guide to the best AI image generators in 2026 explains how design control, editing, realism, and workflow convenience pull in different directions.

Главная страница GlobalGPT

ИИ-платформа "все в одном" для написания текстов, создания изображений и видео с помощью GPT-5, Nano Banana и др.

What Is Ideogram 4.0, and What Changed?

Ideogram released Ideogram 4.0 on June 3, 2026 as its first open-weight text-to-image model trained from scratch. The official repository describes a 9.3B-parameter model with gated NF4 and FP8 variants.

The core idea is structured design. Ideogram 4 uses a 34-layer single-stream diffusion transformer and a frozen Qwen3-VL-8B-Instruct text encoder. That architecture matters less to most creators than the workflow it enables: long structured captions, explicit in-image text, normalized placement boxes, and named color palettes.

9.3Bparameters
2Knative maximum side
NF4 / FP8gated weight variants
JSONnative caption format

The inference guide documents image dimensions from 256 to 2048 pixels per side, in multiples of 16, and aspect ratios as wide as 6:1 or as tall as 1:6. It lists 48 steps for Quality, 20 for Default, and 12 for Turbo.

Official model zoo
Ideogram 4 official repository model table showing NF4 and FP8 variants
The official model table separates NF4 and FP8 support. Both weight routes link to Ideogram’s model agreement.

Is Ideogram 4 Open Source? The License Explained

The precise label is open weight. You can inspect the repository and request the downloadable model files, but access to weights is not the same as an unrestricted open-source license.

The public NF4 and FP8 weights link to the Ideogram Non-Commercial Model Agreement. It covers non-commercial purposes and internal non-production evaluation or research under the agreement. It should not be read as automatic permission to ship client work or a commercial image service.

Public weights

Research, evaluation, and uses allowed by the non-commercial agreement.

Commercial self-hosting

A separate Ideogram commercial-license conversation.

Hosted API

Paid per-image generation presented for commercial use.

If commercial rights are the deciding issue, use our broader commercial-use guide for AI images as a checklist, then apply the current Ideogram agreement and your own legal review to the actual project.

Official license boundary
Ideogram 4 Non-Commercial Model Agreement showing definitions and granted rights
The public model weights are governed by Ideogram’s Non-Commercial Model Agreement. The hosted API and separate commercial licensing are different routes.

Ideogram 4 Quality: Text, Layout, and Photorealism

Our hands-on venue was Broly Anywhere’s asynchronous image task API. It reported Ideogram as the upstream provider and exposed Default, Turbo, and Quality routes. GPT Image 2 and Nano Banana 2 baselines ran through GlobalGPT. Because the venues, seeds, and output sizes differ, the comparison is directional.

6 / 6

Required bilingual text groups correct in Ideogram T1.

+ LUMA

T2A added bottle text despite “no extra text.”

4 controls

T5 requested exactly two black knobs.

16–36 s

Server task times across six successful runs.

Test 1: bilingual typography and poster design

The poster required six exact English and Chinese text groups, a sculptural cobalt chair, an ivory/blue/red/black palette, a Swiss grid, and no extra logo or watermark. Ideogram 4 rendered all six groups correctly in this run, including 北港设计周.

All three baseline models rendered the required text correctly. Ideogram used the most explicit visible grid and the most negative space; GPT Image 2 used the densest editorial hierarchy; Nano Banana 2 used the largest centered copy. Our separate Nano Banana 2. Тест на отображение текста provides more model-specific context.

Hands-on test 1 / Typography
Ideogram 4, GPT Image 2 and Nano Banana 2 bilingual poster outputs side by side
All three models rendered the six required text groups correctly in one run. Design direction differed, and outputs came from different routes and sizes.

Test 2: product detail and photorealism

Ideogram 4 Quality produced convincing brushed steel, walnut, limestone, linen, reflections, and espresso crema. Morning light came from the left, and the still-life composition felt spacious and premium.

Instruction precision was weaker. The machine showed four black control elements instead of exactly two, the pressure-gauge markings did not make a near-9-bar reading reliably verifiable, and the route returned a 1792×2240 portrait image after a 1536×1024 landscape request.

For practical realism work, material quality is only half the job; inventory and geometry still need checking. Our workflow for how to make AI-generated images look real focuses on those production checks.

Hands-on test 2 / Product detail
Ideogram 4 Quality, GPT Image 2 and Nano Banana 2 espresso machine outputs side by side
Ideogram delivered strong materials but missed the exact control count and requested orientation. The three routes used different output dimensions.

Independent typography evidence

A Contra Labs blind study used 10 working designers, four models, 60 prompts, and 240 images. Ideogram 4 won 47.9% of typography matchups and received a 3.55/5 client-work score. The same study found Gemini stronger in detailed-scene accuracy and stylized work, so the result is a typography lead—not a universal quality crown.

Independent designer study
Contra Labs Ideogram 4 typography study summary and headline findings
Contra Labs reported a typography lead for Ideogram while also recording category wins for competing models. The result belongs to that study’s methodology.

How Ideogram 4 JSON Prompts and Bounding Boxes Work

Ideogram’s official prompting guide says the model was trained on structured JSON captions. A caption can describe the background and elements separately, including text strings, object descriptions, normalized bounding boxes, and a style color palette.

Bounding boxes use normalized 0–1000 coordinates in the order shown by the official schema. This makes a prompt portable across image resolutions. It is still generative guidance rather than an editable design-canvas constraint.

Reusable JSON prompt

Poster layout with text, boxes, and palette

Official prompting guide
Ideogram 4 official prompting guide explaining plain text and JSON captions
The local/reference workflow is JSON-native. Plain text is expanded through Magic Prompt in supported routes.

What our gateway test actually proved

The Broly `/v1/tasks` gateway required a `prompt` field. When we placed JSON-looking text there, the service treated it as ordinary input and rewrote it through Magic Prompt. The returned caption removed every submitted bounding box and palette field.

A second diagnostic placed the complete caption in `metadata.json_prompt`; the gateway ignored it and generated an unrelated LUMA AI-video poster. The practical conclusion is narrow but important: this Broly route does not expose Ideogram’s native json_prompt field.

The overlay below compares intended boxes with the rewritten output. The CTA fell inside its planned region, while the headline and bottle only partly matched. This visual diagnoses the gateway path; it is not a native Ideogram bbox-precision score.

Gateway route diagnostic
Intended headline, bottle and CTA bounding boxes over the T2B route output
Route diagnostic only. The gateway removed the submitted bbox fields before generation, so this cannot measure native Ideogram bounding-box precision.

For ordinary prompt tuning before you move into JSON, these methods to improve AI image generation accuracy help reduce conflicting subject, layout, and style instructions.

How to Use Ideogram 4.0

Option 1: hosted Ideogram products

The hosted route is the simplest choice for creators who want Ideogram’s current product experience without managing model files. It also avoids treating the public non-commercial weights as a production license.

Option 2: the official Ideogram v4 API

The official endpoint is POST https://api.ideogram.ai/v1/ideogram-v4/generate. Its API reference documents mutually exclusive text_prompt и json_prompt request fields, plus rendering speed and image-size controls.

Option 3: local NF4 or FP8 weights

The repository table lists NF4 for CUDA and Diffusers support. It labels FP8 for all hardware in the official implementation and does not list Diffusers support for that row. Both model downloads are gated and both remain subject to Ideogram’s model agreement.

Option 4: ComfyUI

ComfyUI publishes an official Ideogram 4 text-to-image workflow with the required model files and workflow entry. ComfyUI makes the graph accessible; it does not replace the weight license or guarantee that every GPU has enough memory.

If you prefer exploring comparable local or hosted workflows, the FLUX 2 Pro pricing and prompt guide offers a useful contrast between model access, prompt style, and production cost.

Hardware, Model Size, and Practical Limits

The honest answer to “How much VRAM does Ideogram 4 need?” is that there is no single verified number in the official material we reviewed. Variant, implementation, offloading, resolution, and system memory all change the practical floor.

Our local machine has a GeForce MX450 with 2GB of VRAM, which is not a sensible test bed for a 9.3B image model. We did not download large weights or run local inference, so this review does not turn a guess into a minimum requirement.

  • Start with the official NF4/FP8 compatibility table, not a third-party one-number promise.
  • Budget memory for the text encoder, denoiser, VAE, resolution, and workflow overhead.
  • Use a tested ComfyUI or Diffusers configuration that names the exact variant and offloading settings.
  • Separate “it loaded once” from a stable production configuration.

The same caution applies to output pixels. The official local guide defines valid resolutions, but our hosted gateway did not treat the requested `size` as a strict output contract. A route may map your request to a supported preset or orientation.

Ideogram 4 API Pricing and Commercial Routes

On August 14, 2026, the official Ideogram 4.0 model page listed three hosted API prices: Turbo at $0.03 per image, Default at $0.06, and Quality at $0.10. Pricing is dynamic, so reopen the page before publishing a budget.

Turbo$0.03

Per image on the official hosted API page.

По умолчанию$0.06

Per image on the official hosted API page.

Качество$0.10

Per image on the official hosted API page.

Checked August 14, 2026. Platform or gateway prices are a separate layer.

Our six Broly tasks totaled $1.33 at the gateway’s task prices: $0.20 for each Default task and $0.33 for the Quality task. The completed task records separately reported $0.40 in upstream Ideogram generation cost. Those two totals measure different billing layers and should not be substituted for one another.

Official hosted pricing
Ideogram 4 hosted API pricing with Turbo, Default and Quality tiers
Ideogram listed three per-image API tiers when checked August 14, 2026. Recheck dynamic prices before publication or purchasing.

Ideogram 4 vs GPT Image 2 vs Nano Banana 2

The useful decision is not “Which model won one image?” It is “Which workflow reduces the kind of correction I actually do?” Ideogram’s edge is visible design structure; GPT Image 2 and Nano Banana 2 can be easier when the brief is conversational or the task depends on another ecosystem.

Решающий факторIdeogram 4Образ GPT 2Нано-банан 2
Best reason to chooseStructured design, text, local controlConversational iteration and premium compositionReadable layouts and Google-centered workflows
Our T1 runStrongest visible gridDensest editorial hierarchyLargest centered text
Our T5 runStrong materials; count/orientation missTight premium close-upClearest brewing action
Local weightsYes, gated and non-commercialNo comparable public weight route hereNo comparable public weight route here

For another controlled choice framework, see the Сравнение Flux 2 и Nano Banana Pro. It shows why local control, text accuracy, editing, and convenience often produce different winners.

Наш сайт Seedream 5.0 Pro vs GPT Image 2 tests add a second photorealism and instruction-following reference point without pretending those results transfer directly to Ideogram.

Limitations and Safety Behavior

The first limit is consistency. A clean one-run poster does not prove every seed will spell six text groups correctly. A strong steel texture does not compensate for the wrong number of controls when the object inventory matters.

The second limit is route behavior. Native JSON, bounding boxes, and palette conditioning remain untested in our Broly venue because the gateway did not pass those fields through. The T4 palette proxy reached 87.45% of sampled pixels near a requested color, but a large ivory background dominated that percentage.

The third limit is production rights. Downloadable weights improve inspectability and local research, but the non-commercial agreement adds a deliberate decision step for commercial teams.

Banani’s named hands-on review praised speed and text while noting prompt-fidelity and credit-efficiency issues in its own UI-asset workflow. That is useful external context, not a substitute for your prompt set or proof of community consensus.

Ideogram’s local safety guide documents blocked-image behavior and Hive integrations. Hosted gateways may apply additional policies, so do not assume local, official hosted, and third-party routes reject or allow exactly the same inputs.

Before building a client pipeline, test one real poster, packaging concept, or ad at final display size. Short, explicit briefs usually outperform crowded instructions; these image-prompt ideas and examples can help you simplify the first draft.

Who Should Use Ideogram 4?

Choose Ideogram 4

Typography, brand concepts, poster structure, JSON-native control, or local research.

Choose another model

Conversational editing, a specific platform ecosystem, or a workflow where local weights do not matter.

Use the hosted API

Commercial generation without negotiating self-hosted model operations.

Pause before local use

You need a one-number VRAM guarantee or unrestricted commercial rights from public weights.

Ideogram 4 is easiest to recommend to designers and developers who know why they want structure. If you only want an attractive image from one sentence, JSON and local deployment may add friction without adding value.

Creators still choosing a general platform can compare our tested Midjourney alternatives before committing to one model family.

Ideogram 4.0 Review Verdict

Ideogram 4.0 is one of the most interesting image-model releases for people who treat text and layout as part of the image, not as decoration added later. Its official JSON format, bounding boxes, palette controls, native 2K range, and gated local weights give it a distinct design-tool identity.

Our tests support a measured verdict. The bilingual poster was excellent in one run, and the Quality route produced convincing materials. But exact object count, gauge readability, and requested orientation failed in the product task. The gateway also prevented a fair native JSON test.

Choose it for typography-led exploration, structured briefs, and local research. Use the official API or separate commercial licensing when the work is commercial. Keep another model in the shortlist when conversational iteration, editing, or ecosystem convenience matters more than JSON-native control.

The simple point is that Ideogram 4.0 earns a serious test, not an automatic crown. Run your own final-size asset, proofread every word, count every required object, and record the exact route that produced the output.

Ideogram 4.0 FAQ

Is Ideogram 4.0 open source or open weight?

Ideogram 4.0 is open weight, not an unrestricted open-source model. The code repository is public, but the downloadable NF4 and FP8 model weights link to Ideogram’s Non-Commercial Model Agreement. Read the code license and weight license separately before deployment.

Can I use Ideogram 4.0 commercially?

Yes, but the route matters. Ideogram’s hosted API is presented as a commercial route, and Ideogram offers separate commercial self-hosting licensing. The public downloadable weights are governed by a non-commercial agreement, so they should not be treated as automatic production permission.

Is Ideogram 4.0 free?

The public weights can be accessed for uses allowed by their agreement, but local operation still requires suitable hardware and setup. Hosted generation is paid per image. On August 14, 2026, Ideogram listed Turbo at $0.03, Default at $0.06, and Quality at $0.10.

How much VRAM does Ideogram 4 need?

Ideogram’s repository labels NF4 for CUDA and Diffusers and FP8 for broader hardware support, but it does not provide one universal minimum-VRAM number. We did not run local inference, so this review does not invent a requirement. Check the current repository and your chosen workflow.

How do I run Ideogram 4 in ComfyUI?

Use the official ComfyUI Ideogram 4 workflow and download the model files it specifies. Keep the Ideogram model agreement in view even when ComfyUI provides the workflow. Hardware support, practical speed, and memory use depend on the weight variant and system.

Why does Ideogram 4 use JSON prompts?

Ideogram 4 was trained with structured JSON captions. JSON separates background, objects, in-image text, layout boxes, and palette instructions more explicitly than a long paragraph. Plain text is still possible through routes that expand it, but structured input is the native control language.

How much does the Ideogram 4 API cost?

Ideogram’s model page listed $0.03 per Turbo image, $0.06 per Default image, and $0.10 per Quality image on August 14, 2026. Gateway or platform prices can be higher. Recheck the official page immediately before estimating a production budget.

Is Ideogram 4 better than GPT Image 2 or Nano Banana 2 for text?

No universal winner follows from our two one-run prompts. All three models rendered the required bilingual poster copy correctly. Ideogram used the clearest visible grid, GPT Image 2 used the densest editorial hierarchy, and Nano Banana 2 favored large centered readability.

What resolution and aspect ratios does Ideogram 4 support?

The official inference guide documents dimensions from 256 to 2048 pixels per side, in multiples of 16, with aspect ratios up to 6:1 or 1:6. A third-party gateway may map size requests differently; our hosted route did not return the requested pixel dimensions exactly.

Can I use Ideogram 4 on GlobalGPT?

We did not verify a public GlobalGPT Ideogram 4 product page on August 14, 2026. The Broly Anywhere test endpoint is not the same thing as a public product route. GlobalGPT does provide other verified image-model workflows, which can be compared without implying they are Ideogram 4.

Поделиться сообщением:

Похожие посты

Обзор Eleven Multilingual v2: языки, качество и API

Обзор ElevenLabs Multilingual v2 (2026): языки, качество голоса и API

Стоит ли использовать Eleven Multilingual v2 в 2026 году? Ознакомьтесь с обзором 29 поддерживаемых языков, качеством голоса, компромиссами при работе с длинными текстами, возможностями управления через API и текущими ценами.

Читать далее
eleven-music-v2-vs-qwen-audio-3-vsseed-audio-1-cover.webp

Eleven Music v2, Qwen Audio 3.0 и Seed Audio 1.0: какая модель искусственного интеллекта для обработки звука подходит для вашей задачи?

Сравните сегодня Eleven Music v2, Qwen Audio 3.0 TTS Flash и Seed Audio 1.0, оценив их по музыкальному сопровождению, дикторскому тексту, звуковому фону, стоимости и результатам тестов воспроизведения аудио.

Читать далее
Обзор DeepSeek V4 Pro: тесты, цены и как попробовать

Обзор DeepSeek V4 Pro: что изменилось, тесты производительности, цены и как попробовать устройство

DeepSeek V4 Pro или Flash? Ознакомьтесь с нашими тестами по программированию, обработке длинных контекстов и принятию решений, сравните актуальные цены на API, изучите ограничения и выберите подходящую модель.

Читать далее