<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>AI – IMG.LY Blog</title><description>Posts tagged AI on the IMG.LY blog.</description><link>https://img.ly/blog/tag/ai/</link><language>en-us</language><image><url>https://img.ly/apple-touch-icon.png</url><title>AI – IMG.LY Blog</title><link>https://img.ly/blog/tag/ai/</link></image><atom:link href="https://img.ly/blog/tag/ai/rss.xml" rel="self" type="application/rss+xml"/><generator>Astro</generator><lastBuildDate>Tue, 01 Sep 2026 09:42:42 GMT</lastBuildDate><ttl>60</ttl><item><title>Evaluating CoDesign MCP performance in resizing</title><link>https://img.ly/blog/evaluating-codesign-mcp-performance-in-resizing/</link><guid isPermaLink="true">https://img.ly/blog/evaluating-codesign-mcp-performance-in-resizing/</guid><description>One approved design, five ad formats, six AI models, the same brief. All six came back with ads you could ship, and Gemini 3.7 Flash did it for thirty cents. Here is what our automated checks can prove about that work, and what they cannot.</description><pubDate>Fri, 21 Aug 2026 09:00:00 GMT</pubDate><content:encoded>&lt;p&gt;You designed a banner. The same day, you need it in five sizes, and none of them can be wrong.&lt;/p&gt;
&lt;p&gt;We handed that job to six AI models through &lt;a href=&quot;https://img.ly/codesign/&quot;&gt;CoDesign&lt;/a&gt; MCP: take this approved design and rebuild it for five ad formats. All six came back with ads you could ship. Gemini 3.7 Flash did it for thirty cents and passed every check we wrote.&lt;/p&gt;
&lt;p&gt;That is what CoDesign MCP is for. You point an AI assistant such as Claude at your approved design, name the formats you need, and it opens the file, rebuilds the layout for each one and exports production files. What comes back is an editable design file, not a picture of one.&lt;/p&gt;
&lt;p&gt;Speed and price are the easy things to measure. We wanted to measure quality too, so we built an eval that runs the job automatically, opens every result and reports what passed. This is the first scenario out of that suite.&lt;/p&gt;
&lt;h2 id=&quot;three-banners-one-brief&quot;&gt;Three banners, one brief&lt;/h2&gt;
&lt;p&gt;Three different models were handed the same 320×50 slot, cut from the same master creative against the same brief, and all three came back with an ad you could traffic tomorrow. Every one of them keeps the wordmark and the button, nothing is clipped, and no type is set below 11px.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;Three 320×50 mobile banners built from the same ATLAS master by three different models&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1200px) 1200px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1200&quot; height=&quot;207&quot; src=&quot;https://img.ly/_astro/hook-three-mobile-strips.Q8QXs3k1_2XYzx.webp&quot; srcset=&quot;/_astro/hook-three-mobile-strips.Q8QXs3k1_Z1pTCV1.webp 640w, /_astro/hook-three-mobile-strips.Q8QXs3k1_2oKvSC.webp 750w, /_astro/hook-three-mobile-strips.Q8QXs3k1_27uKLv.webp 828w, /_astro/hook-three-mobile-strips.Q8QXs3k1_ZNKOe4.webp 1080w, /_astro/hook-three-mobile-strips.Q8QXs3k1_2XYzx.webp 1200w&quot;&gt;&lt;/p&gt;
&lt;p&gt;The 300×250 they came from carried five things: the wordmark, the headline, a retention chart, the offer line and the button. At 320×50 the wordmark and the button survive and the chart cannot, which leaves room for exactly one line of copy, so each model had to choose between the headline and the offer.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;One dropped the headline and kept the offer, “14-day free trial”.&lt;/li&gt;
&lt;li&gt;One kept the headline but cut “your users” down to “users” to buy itself the line.&lt;/li&gt;
&lt;li&gt;One kept the headline whole across two lines and dropped the offer.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;We have opinions about which of those reads best, the way any design team would. What we do not have is a way to prove one of them correct, because the choice depends on whether this campaign is selling the product or the trial, and that lives in the media brief rather than in the design. Our checks can confirm that all three are valid and that none of them are broken; they cannot tell you which one the campaign actually needed. Grading the half that has a right answer and leaving the other half to the people who own the brief is the distinction our evals are built around.&lt;/p&gt;
&lt;h2 id=&quot;what-is-gradeable-and-what-is-not&quot;&gt;What is gradeable and what is not&lt;/h2&gt;
&lt;p&gt;The obvious thing to evaluate is the thing that demos well, prompt in and poster out, but we were more interested in the work that actually fills a designer’s week, which is mostly derivative. The design decisions were made once, at the master, and everything after that is a matter of carrying them faithfully into thirty other shapes. That is more technical work than creative work, which is exactly why it is worth handing to an agent.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;Three valid 300×600 versions of the same ATLAS banner, differing in type size and spacing&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1200px) 1200px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1200&quot; height=&quot;702&quot; src=&quot;https://img.ly/_astro/gradeable-three-halfpage-variants.D3Bbuya5_2lpIzm.webp&quot; srcset=&quot;/_astro/gradeable-three-halfpage-variants.D3Bbuya5_Z1FdAzV.webp 640w, /_astro/gradeable-three-halfpage-variants.D3Bbuya5_Z1ihu0z.webp 750w, /_astro/gradeable-three-halfpage-variants.D3Bbuya5_FqKnt.webp 828w, /_astro/gradeable-three-halfpage-variants.D3Bbuya5_ZWQJb7.webp 1080w, /_astro/gradeable-three-halfpage-variants.D3Bbuya5_2lpIzm.webp 1200w&quot;&gt;&lt;/p&gt;
&lt;p&gt;Ask which of the finished designs is better and you have a question with no ground truth to appeal to, since human raters disagree with each other and LLM judges reliably prefer their own output. You could ask a lot of humans and rank the results, but that is not something you can run on every build. Wrong designs, though, are wrong in ways you can query: text sitting behind another block, content off the page, type under the legibility floor. Part of design quality is judgement and part of it is fact, and what we have built grades the second part, which means everything below is an automated result rather than a design review. A model can pass every check we wrote and still make an ugly ad.&lt;/p&gt;
&lt;h2 id=&quot;what-we-are-doing-right-now&quot;&gt;What we are doing right now&lt;/h2&gt;
&lt;p&gt;When a model finishes a scenario, programmatic checks open the result and ask whether it is plausibly good. That starts with whether the run produced every design artifact and .imgly file at the right sizes, and the rest depends on the scenario: in a localization run we also check that the words were actually replaced.&lt;/p&gt;
&lt;p&gt;This is a technical floor rather than a grade. LLM judges are still unreliable enough that we do not use one on output quality, because if it is hard for us as designers to say which design is more correct, it is harder for a model, and the result would be noise.&lt;/p&gt;
&lt;p&gt;In the future we plan to experiment more with such methodologies but we also do not want to rely on the false security of an extremely noisy grading process.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;The eval suite&amp;#39;s comparison view for the display-banner matrix&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 2396px) 2396px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;2396&quot; height=&quot;1344&quot; src=&quot;https://img.ly/_astro/evals-suite-comparison-resize-matrix.CV8e1T_u_8PXp1.webp&quot; srcset=&quot;/_astro/evals-suite-comparison-resize-matrix.CV8e1T_u_Znaqv1.webp 640w, /_astro/evals-suite-comparison-resize-matrix.CV8e1T_u_Z20zsc2.webp 750w, /_astro/evals-suite-comparison-resize-matrix.CV8e1T_u_Z26h51n.webp 828w, /_astro/evals-suite-comparison-resize-matrix.CV8e1T_u_Z1Eg8Vb.webp 1080w, /_astro/evals-suite-comparison-resize-matrix.CV8e1T_u_CxK9F.webp 1280w, /_astro/evals-suite-comparison-resize-matrix.CV8e1T_u_Z1lqFrF.webp 1668w, /_astro/evals-suite-comparison-resize-matrix.CV8e1T_u_ZgDr5i.webp 2048w, /_astro/evals-suite-comparison-resize-matrix.CV8e1T_u_8PXp1.webp 2396w&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The comparison view for the display-banner matrix: cost and time sitting directly above what each model actually produced. The 1/1 counts runs, one repeat per cell, not check scores.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&quot;one-master-five-ad-sizes&quot;&gt;One master, five ad sizes&lt;/h2&gt;
&lt;p&gt;The scenario is a display campaign, because that is the version of this problem customers actually have: the same size matrix, every campaign. The master is a 300×250 MPU for ATLAS, an imaginary product-analytics SaaS running a free-trial campaign, carrying a wordmark, the headline “See what your users &lt;strong&gt;actually&lt;/strong&gt; do”, a retention bar chart labelled “43% retained”, the line “14-day free trial. No card.”, and a “Start free trial” button.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;The ATLAS 300×250 master at the top, with an arrow fanning down to the same ad rebuilt at 300×600, 728×90, 320×50 and 160×600&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1200px) 1200px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1200&quot; height=&quot;675&quot; src=&quot;https://img.ly/_astro/master-and-five-sizes.CzF5wnY5_21RxX1.webp&quot; srcset=&quot;/_astro/master-and-five-sizes.CzF5wnY5_FUIkG.webp 640w, /_astro/master-and-five-sizes.CzF5wnY5_Z7tPBU.webp 750w, /_astro/master-and-five-sizes.CzF5wnY5_mzCbS.webp 828w, /_astro/master-and-five-sizes.CzF5wnY5_Zpv1pN.webp 1080w, /_astro/master-and-five-sizes.CzF5wnY5_21RxX1.webp 1200w&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The 300×250 master at the top, and the four other sizes on the media plan cut from it. ATLAS is our own fixture, not a real customer.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The brief asks for the full matrix: 300×250, 728×90, 160×600, 300×600 and 320×50. Those sizes span a 16:1 swing in aspect ratio, so the right answer at each one is a genuinely different composition, and at some of them it means taking content out rather than shrinking it.&lt;/p&gt;
&lt;p&gt;Ten checks grade the result and two of them do the real work. &lt;strong&gt;&lt;code&gt;728×90: the chart was dropped, not shrunk&lt;/code&gt;&lt;/strong&gt; kills the model that squashes a bar chart into a 90-pixel strip, which satisfies every geometric constraint in the brief while being completely wrong. &lt;strong&gt;&lt;code&gt;the five sizes are different layouts&lt;/code&gt;&lt;/strong&gt; kills the model that scales one composition five ways and calls it a matrix.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;Contact sheet of six 728×90 leaderboards, master first&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1280px) 1280px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1280&quot; height=&quot;676&quot; src=&quot;https://img.ly/_astro/sheet-display-leaderboard.BStENqrh_1X5bgP.webp&quot; srcset=&quot;/_astro/sheet-display-leaderboard.BStENqrh_2uIgrs.webp 640w, /_astro/sheet-display-leaderboard.BStENqrh_Z1sXCsn.webp 750w, /_astro/sheet-display-leaderboard.BStENqrh_Z1JQH3t.webp 828w, /_astro/sheet-display-leaderboard.BStENqrh_Z2gFqpY.webp 1080w, /_astro/sheet-display-leaderboard.BStENqrh_1X5bgP.webp 1280w&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The master first, then the six 728×90 leaderboards. Every model that delivered deleted the retention chart.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Every model that delivered arrived at the same composition, wordmark left, headline across the middle, offer underneath it, button on the right and the retention chart deleted, which is six models from five different vendors reaching the same answer independently. We did not expect that, and it is still the result we find hardest to stop thinking about.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;Contact sheet of six 160×600 skyscrapers, master first&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1280px) 1280px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1280&quot; height=&quot;676&quot; src=&quot;https://img.ly/_astro/sheet-display-skyscraper.C88KXRcA_Z1lHK5K.webp&quot; srcset=&quot;/_astro/sheet-display-skyscraper.C88KXRcA_Z2cPTfD.webp 640w, /_astro/sheet-display-skyscraper.C88KXRcA_1krJOz.webp 750w, /_astro/sheet-display-skyscraper.C88KXRcA_ZmDw0F.webp 828w, /_astro/sheet-display-skyscraper.C88KXRcA_Z1S3Ozm.webp 1080w, /_astro/sheet-display-skyscraper.C88KXRcA_Z1lHK5K.webp 1280w&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The same six at 160×600, master first. With vertical room they all kept the chart. The flat format is where the decision to drop it had to be made.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&quot;which-model-is-best-for-using-with-codesign-mcp&quot;&gt;Which model is best for using with CoDesign MCP?&lt;/h2&gt;





























































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align=&quot;left&quot;&gt;#&lt;/th&gt;&lt;th align=&quot;left&quot;&gt;model&lt;/th&gt;&lt;th align=&quot;left&quot;&gt;checks&lt;/th&gt;&lt;th align=&quot;left&quot;&gt;model spend&lt;/th&gt;&lt;th align=&quot;left&quot;&gt;time&lt;/th&gt;&lt;th align=&quot;left&quot;&gt;tool calls&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align=&quot;left&quot;&gt;1&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;gemini-3.7-flash&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;10 / 10&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;$0.30&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;6.7 min&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;37&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align=&quot;left&quot;&gt;2&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;glm-5.2&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;10 / 10&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;$0.40&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;8.2 min&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;25&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align=&quot;left&quot;&gt;3&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;gpt-5.6-sol&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;10 / 10&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;$0.52&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;4.0 min&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;35&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align=&quot;left&quot;&gt;4&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;grok-4.6&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;10 / 10&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;$1.20&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;20.9 min&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;39&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align=&quot;left&quot;&gt;5&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;claude-fable-5&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;10 / 10&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;$3.95&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;10.3 min&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;26&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align=&quot;left&quot;&gt;6&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;claude-opus-5&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;10 / 10&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;$3.97&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;18.8 min&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;40&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Running this scenario illustrates general numbers across our eval suite. The 6 models above are consistently able to follow the prompts to produce different output designs as requested in the scenario. These numbers also highlight quite the spread in model costs and time spend. What you do not see in these numbers is that glm 5.2, gpt-5.6-sol and grok 4.6 are producing suboptimal designs in our opinion. They are valid, but subjectively fable, opus and gemini 3.7 consistently across runs and across scenarios perform simply better.&lt;/p&gt;
&lt;p&gt;The biggest surprise for us is &lt;strong&gt;Gemini 3.7 Flash&lt;/strong&gt;, which produced the whole five-size matrix for $0.30 in model spend, in under seven minutes. At thirty cents a designer opens six composed variants instead of a blank artboard and takes the best one further, which changes what the tool is for.&lt;/p&gt;
&lt;p&gt;Meanwhile there were no surprises regarding &lt;strong&gt;Opus 5.0 and Fable 5.0&lt;/strong&gt;, both models perform really well - but are slow and expensive.&lt;/p&gt;
&lt;h2 id=&quot;what-does-it-mean-for-codesign-mcp&quot;&gt;What does it mean for CoDesign MCP&lt;/h2&gt;
&lt;p&gt;You can now resize a master design into 5 different sizes reliably for $0.30 using Gemini 3.7 Flash and the CoDesign MCP.&lt;/p&gt;
&lt;p&gt;If you are on a Claude Subscription plan, you might just want to continue using your Claude Code with the CoDesign MCP since it offers the best overall design quality - even if it’s slow and somewhat expensive.&lt;/p&gt;
&lt;p&gt;We will continue to work on upgrading and improving measurements of how well CoDesign MCP works with different models, looking into other scenarios like localization (translating and changing a design for a different locality and language) as well as rebranding existing design assets.&lt;/p&gt;
&lt;p&gt;You can install CoDesign MCP at &lt;a href=&quot;https://img.ly/codesign/&quot;&gt;&lt;strong&gt;img.ly/codesign&lt;/strong&gt;&lt;/a&gt;.&lt;/p&gt;</content:encoded><dc:creator>Dustin</dc:creator><dc:creator>Mirko</dc:creator><media:content url="/_astro/key-visual-resize-fanout.qeZDEyVs.png" medium="image"/><category>AI</category><category>Insights</category><category>CE.SDK</category></item><item><title>AI Design Agents in 2026: Which One Fits Your Work</title><link>https://img.ly/blog/ai-design-agents/</link><guid isPermaLink="true">https://img.ly/blog/ai-design-agents/</guid><description>Six AI design agents grouped by the work they suit: app screens, campaign assets, or design inside a larger task. Each entry says what you can still edit once the agent stops.</description><pubDate>Mon, 17 Aug 2026 09:00:00 GMT</pubDate><content:encoded>&lt;figure class=&quot;kg-embed-card&quot;&gt;
  &lt;iframe src=&quot;https://www.youtube.com/embed/T6jdkJMsi4Q?feature=oembed&quot; title=&quot;AI Design Agents Tested: Figma, Stitch, Lovart, Canva, CoDesign&quot; loading=&quot;lazy&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; referrerpolicy=&quot;strict-origin-when-cross-origin&quot; allowfullscreen&gt;&lt;/iframe&gt;
&lt;/figure&gt;
&lt;p&gt;Most roundups of AI design agents rank the tools one to six. That only works if every reader wants the same thing. A developer wiring up app screens and a brand manager producing a campaign are not shopping in the same market, even though the same six products keep showing up in both their search results.&lt;/p&gt;
&lt;p&gt;This survey sorts by the work instead. Two questions do most of the sorting: what you are making, and what you are left holding when the agent stops. If you want the wider tool category rather than agents specifically, we have &lt;a href=&quot;https://img.ly/blog/vibe-design-tools-compared/&quot;&gt;compared the AI design tools&lt;/a&gt; built around prompt-first workflows.&lt;/p&gt;
&lt;p&gt;We ran five of the six ourselves, on the same two prompts, and the findings below are what we saw rather than what the marketing pages say.&lt;/p&gt;
&lt;p&gt;Disclosure up front. IMG.LY publishes this survey, and two of the six tools here have a commercial relationship with us: CoDesign is our own product, and Manus is an IMG.LY customer. Both entries say so where they appear, we assess both against the same criteria as everything else, and anything we did not test ourselves rests on public sources.&lt;/p&gt;
&lt;h2 id=&quot;what-counts-as-a-design-agent&quot;&gt;What counts as a design agent&lt;/h2&gt;
&lt;p&gt;Three tests separate agents from the wider AI design tool space. The tool must plan multi-step work from a single goal. It must execute design operations rather than only advising. Its output must be usable design work: files, screens, or code.&lt;/p&gt;
&lt;p&gt;These tests exclude two familiar groups. Image generators such as Midjourney and Ideogram produce strong single images without planning or executing a task around them. Chat assistants configured with design prompts, such as Taskade’s design agent templates, return recommendations in text. Both are useful, but neither does design work on your behalf.&lt;/p&gt;
&lt;p&gt;Six products meet that definition in August 2026. We work through the definition in more detail in &lt;a href=&quot;https://img.ly/blog/what-is-a-design-agent/&quot;&gt;What Is an AI Design Agent?&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;how-we-tested&quot;&gt;How we tested&lt;/h2&gt;
&lt;p&gt;One prompt per category, given identically to every tool in that category.&lt;/p&gt;
&lt;p&gt;For the interface tools: design a three-screen onboarding flow for a habit-tracking app, the screens in order being a welcome screen, a goal-setting screen, and a first habit formation screen, styled clean, high contrast, large type, one accent color.&lt;/p&gt;
&lt;p&gt;For the graphic design tools: a launch campaign for a cold brew called Northwind, in three formats, an Instagram post, a story, and a DIN A5 flyer, warm and minimal, cream background, one product photo, headline supplied.&lt;/p&gt;
&lt;p&gt;Two things decide it, and neither is how good the first draft looks.&lt;/p&gt;
&lt;p&gt;Is the agent context aware? Can you feed it data, make it familiar with your brand, and does it hold that across every variant? We do not much care how fancy the design is. The table stakes have to be right before anything else counts.&lt;/p&gt;
&lt;p&gt;And does it fit the habits you carried over from the before-AI age? Making a small revision by going back and prompting an AI again is the wrong shape. You want to be the human in the loop who makes that edit manually, because by the time you are prompting for the third time your frustration is already high enough that you have stopped wanting to.&lt;/p&gt;
&lt;p&gt;One caveat on the Figma run, since it cuts against us: it executed the prompt inside IMG.LY’s own design system, so its icons, fonts, and patterns inherit work our designers had already done. That makes the head-to-head somewhat unfair in Figma’s favor.&lt;/p&gt;
&lt;p&gt;Manus is the exception. We have not run it hands-on, and its entry below rests on public sources.&lt;/p&gt;
&lt;h2 id=&quot;what-you-hold-afterward&quot;&gt;What you hold afterward&lt;/h2&gt;
&lt;p&gt;Generation quality converges fast. Every tool here produces a credible first draft, and the differences narrow every quarter. If you want to see where the underlying models differ, we &lt;a href=&quot;https://img.ly/ai-benchmarks/&quot;&gt;benchmark 15 of them&lt;/a&gt; against the same design prompts.&lt;/p&gt;
&lt;p&gt;What you keep does not converge. Some tools hand you a file inside an editor where you can select any element and change it. Others give you an export to finish in a different program. The rest keep the work inside their own platform, editable there and nowhere else.&lt;/p&gt;
&lt;h2 id=&quot;ai-design-agents-for-app-and-web-interfaces&quot;&gt;AI design agents for app and web interfaces&lt;/h2&gt;
&lt;p&gt;Screen design has the most mature agent tooling, partly because UI has structure a model can reason about and partly because the output can be code.&lt;/p&gt;
&lt;h3 id=&quot;google-stitch&quot;&gt;Google Stitch&lt;/h3&gt;
&lt;p&gt;Stitch is a free experimental tool from Google Labs that generates multi-screen mobile and web interfaces, plus frontend code, from text or image prompts. It launched at I/O 2025 and gained Gemini 3 and interactive prototype flows in December 2025.&lt;/p&gt;
&lt;p&gt;You describe an app and Stitch plans the screens, which you refine through further prompts. Screens render in the browser, with code export for the web stack and a paste-to-Figma path that preserves layers. Stitch has no source format of its own, so you can only keep editing in the exported code or the Figma file it hands off to. Reviewers recommend exploring in Stitch and refining elsewhere.&lt;/p&gt;
&lt;p&gt;On our onboarding prompt it did not adhere to the brief. Habit formation came before goal setting, so the screen order was wrong. Button styles were somewhat inconsistent, the look and feel differed between screens, visible in the progress indicator on the welcome screen against the one on goal setting, and the aspect ratios were weird.&lt;/p&gt;
&lt;p&gt;Editing is where it separates from Figma. Double-clicking does let you change the text, but you are apparently supposed to edit with AI, which is a nuisance when you only want to make a few design fixes. The brand kit is the part we liked: colors, accent colors, fonts, and spacing can all be changed globally. Though it looked like Stitch did not adhere to its own system that well.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;Google Stitch showing the generated onboarding screens, with habit formation appearing before goal setting and the progress indicators differing between screens&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1920px) 1920px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1920&quot; height=&quot;1158&quot; src=&quot;https://img.ly/_astro/stitch-ui.BpsuLGO1_21CgRX.webp&quot; srcset=&quot;/_astro/stitch-ui.BpsuLGO1_2mzJWp.webp 640w, /_astro/stitch-ui.BpsuLGO1_2c3zyR.webp 750w, /_astro/stitch-ui.BpsuLGO1_Z21Ahyw.webp 828w, /_astro/stitch-ui.BpsuLGO1_1RtJSh.webp 1080w, /_astro/stitch-ui.BpsuLGO1_Z1l1Hy3.webp 1280w, /_astro/stitch-ui.BpsuLGO1_2aSrbg.webp 1668w, /_astro/stitch-ui.BpsuLGO1_21CgRX.webp 1920w&quot;&gt;&lt;/p&gt;
&lt;p&gt;It is free while in Google Labs. Google does not publish the generation caps, and third-party sources describe the limits inconsistently. Best fit is getting a first version on screen fast, especially for developers who want screens and starter code in one pass. Its Labs status makes it a tool to explore with rather than build a production process on.&lt;/p&gt;
&lt;h3 id=&quot;figma-agent&quot;&gt;Figma Agent&lt;/h3&gt;
&lt;p&gt;Figma’s agent works directly on the design canvas. It entered beta in May 2026 and gained custom skills, web search, and external MCP connections at Config in June 2026.&lt;/p&gt;
&lt;p&gt;You prompt from any layer. The agent executes multi-step tasks: bulk edits across screens, dark-mode conversion, populating designs with real content, turning feedback threads into revisions. Custom skills written as markdown steer how it works.&lt;/p&gt;
&lt;p&gt;Output is native Figma layers with components, variables, and design tokens preserved. This is the strongest design-system integration of the six, because the agent works inside the system your team already maintains rather than approximating it. Code generation routes through Figma Make, a separate product.&lt;/p&gt;
&lt;p&gt;On our onboarding prompt it executed almost perfectly. The design was very clean, it adhered to the standards of modern apps, and it almost looked like something you would find browsing the App Store. Even the copy was decent, and it did not read like AI slop. The screens looked thought through: asked for a primary goal, it offered routines, breaking a bad habit, and staying consistent, which are the broad categories without too many assumptions on top. It stuck to the screen order, it baked in some interactivity, and it caught small details like “remind me every day at 8 a.m.” on the habit formation screen. As a novice you could take this and start prototyping.&lt;/p&gt;
&lt;p&gt;Remember the caveat above, though: this run inherited IMG.LY’s design system, so the icons, fonts, and patterns were already good before the agent started.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;The Figma Agent&amp;#39;s three-screen onboarding flow on the Figma canvas: welcome, goal setting and first habit, in order, with consistent buttons and type&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1920px) 1920px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1920&quot; height=&quot;1386&quot; src=&quot;https://img.ly/_astro/figma-ui.Y16RdKY8_Z268T6m.webp&quot; srcset=&quot;/_astro/figma-ui.Y16RdKY8_2v5Q6d.webp 640w, /_astro/figma-ui.Y16RdKY8_Z1qzg7w.webp 750w, /_astro/figma-ui.Y16RdKY8_ZXCx2r.webp 828w, /_astro/figma-ui.Y16RdKY8_MFTUh.webp 1080w, /_astro/figma-ui.Y16RdKY8_Z2rSWLd.webp 1280w, /_astro/figma-ui.Y16RdKY8_1Dj2Pf.webp 1668w, /_astro/figma-ui.Y16RdKY8_Z268T6m.webp 1920w&quot;&gt;&lt;/p&gt;
&lt;p&gt;The agent is free during beta, and Figma has said AI credits will apply at general availability without publishing amounts. It requires a full seat on Professional, Organization, or Enterprise; full seats start at $16 per person per month on the annual Professional plan. Early reviews report rough edges, including responsive layouts that broke on mobile.&lt;/p&gt;
&lt;p&gt;In the head-to-head, Figma beat Stitch by a wide margin, which was to be expected.&lt;/p&gt;
&lt;p&gt;Best fit is product and UI teams already working in Figma with an established design system.&lt;/p&gt;
&lt;h2 id=&quot;ai-design-agents-for-marketing-and-brand-assets&quot;&gt;AI design agents for marketing and brand assets&lt;/h2&gt;
&lt;p&gt;In campaign work you are producing dozens of assets at once, they all have to stay on brand, and some of them have to survive contact with a printer.&lt;/p&gt;
&lt;h3 id=&quot;lovart&quot;&gt;Lovart&lt;/h3&gt;
&lt;p&gt;Lovart is a dedicated design agent that turns a brief into a coordinated set of deliverables: logos, posters, packaging, social assets. It launched publicly in July 2025 after a closed beta that drew several hundred thousand users.&lt;/p&gt;
&lt;p&gt;Give it a campaign-level brief and it analyzes intent, researches references, then generates dozens of assets sharing one visual system. You and the agent iterate on ChatCanvas, an infinite canvas where you both edit the same design through conversation. Results land as layers with typography kept separately, and Lovart offers targeted element and text edits in place.&lt;/p&gt;
&lt;p&gt;On the Northwind campaign it was fairly quick, and the result was very aesthetically pleasing. We had very little to quibble with. The design subtleties were there, the product placement, the coffee beans, some cloth, and it stuck very well to the specification. Exactly what we had in mind.&lt;/p&gt;
&lt;p&gt;The one quibble is brand drift. Look at the logo, look at the bottle, and there is significant drift in brand identity across the three formats. That is fixable in a real scenario, because you can link a brand, so you would have the product image and the logo ready and we would not expect much trouble then, whether you use Lovart to create variations, produce marketing material to A/B test, or adjust for different formats.&lt;/p&gt;
&lt;p&gt;Editing is the problem. Ask to change the headline text and you cannot tell whether you are even looking at the same font. Quick edit means writing another prompt. There are no layers, and there is nothing you can do directly. So while the output looks nice, for revision work it is effectively useless.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;Lovart&amp;#39;s canvas with the Northwind campaign assets, where editing the headline means writing another prompt rather than selecting a text layer&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1920px) 1920px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1920&quot; height=&quot;1027&quot; src=&quot;https://img.ly/_astro/lovart-ui.DIzU1wlA_Zh3UPf.webp&quot; srcset=&quot;/_astro/lovart-ui.DIzU1wlA_248rGS.webp 640w, /_astro/lovart-ui.DIzU1wlA_3cGnN.webp 750w, /_astro/lovart-ui.DIzU1wlA_ZQVJHY.webp 828w, /_astro/lovart-ui.DIzU1wlA_2hTJac.webp 1080w, /_astro/lovart-ui.DIzU1wlA_Z18qvYg.webp 1280w, /_astro/lovart-ui.DIzU1wlA_UlQg1.webp 1668w, /_astro/lovart-ui.DIzU1wlA_Zh3UPf.webp 1920w&quot;&gt;&lt;/p&gt;
&lt;p&gt;Reviewers report that precise pixel-level work tends to move into Photoshop or Figma anyway, which matches what we found.&lt;/p&gt;
&lt;p&gt;Pricing runs on credit tiers from Starter to Ultimate, 2,000 to 27,000 credits per month. Lovart renders prices dynamically and third-party reports of the dollar amounts vary between roughly $15 and $90 monthly, so check the pricing page directly. Credit consumption is hard to predict on agent-driven tasks.&lt;/p&gt;
&lt;p&gt;Best fit is marketing teams whose unit of work is the campaign rather than the asset, and who accept a finishing pass elsewhere.&lt;/p&gt;
&lt;h3 id=&quot;canva-ai-assistant&quot;&gt;Canva AI Assistant&lt;/h3&gt;
&lt;p&gt;Canva’s assistant plans and executes design tasks by calling Canva’s own tools. It launched in April 2025 and received a tool-calling rebuild, announced as a research preview, in April 2026.&lt;/p&gt;
&lt;p&gt;It runs multi-step jobs such as producing a multi-channel campaign, and pulls context from connected sources including Slack, Gmail, Google Drive, and Notion. Scheduled tasks run in the background, with finished work arriving as drafts for review.&lt;/p&gt;
&lt;p&gt;Output is layered, fully editable Canva designs across presentations, social posts, documents, and spreadsheets. You can change any element without regenerating. The constraint is Canva itself. The work stays editable inside Canva, and getting it out means exporting.&lt;/p&gt;
&lt;p&gt;On the same Northwind prompt it was fairly quick, and it appears to match the brief against the vast template library Canva already has. From the get-go, though, it did not really stick to the specification. We asked for a cream background and one product photo. There is no cream background on the first asset, and we do not know why the product would be displayed on a laptop, which is plain weird. Only the last of the three is anywhere near acceptable.&lt;/p&gt;
&lt;p&gt;We also had no indication of whether it generated variations or ignored that part of the brief. Opening one of the designs gives you more variations, including the one you did not select. What you do not get is a multi-page layout you can change, or any way to specify changes precisely. At that point you are simply inside Canva, editing.&lt;/p&gt;
&lt;p&gt;As a starting point we would have been just as well off picking one of the Canva templates, uploading the product image, and adjusting the text. Compared with Lovart on the identical brief, Canva really does fall short.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;Canva&amp;#39;s assistant returning the Northwind assets, without the cream background the brief asked for and with the product shown on a laptop&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1920px) 1920px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1920&quot; height=&quot;1092&quot; src=&quot;https://img.ly/_astro/canva-ui.DpSjPHH5_Z1vYvTM.webp&quot; srcset=&quot;/_astro/canva-ui.DpSjPHH5_jD6q9.webp 640w, /_astro/canva-ui.DpSjPHH5_2cxk3p.webp 750w, /_astro/canva-ui.DpSjPHH5_Z1NMplW.webp 828w, /_astro/canva-ui.DpSjPHH5_Z1fW75G.webp 1080w, /_astro/canva-ui.DpSjPHH5_Z1aCzQe.webp 1280w, /_astro/canva-ui.DpSjPHH5_Z1vGnHn.webp 1668w, /_astro/canva-ui.DpSjPHH5_Z1vYvTM.webp 1920w&quot;&gt;&lt;/p&gt;
&lt;p&gt;The assistant is available on the free tier with rate limits and monthly credits. Pro is $18 monthly. Business is $25 per user monthly with larger allowances. Reviewers warn that letting the assistant run a job end to end uses premium credits faster than editing by hand.&lt;/p&gt;
&lt;p&gt;Best fit is teams already on Canva who want campaign production handled conversationally with brand kits applied automatically.&lt;/p&gt;
&lt;h3 id=&quot;imgly-codesign&quot;&gt;IMG.LY CoDesign&lt;/h3&gt;
&lt;p&gt;Disclosure: CoDesign is made by IMG.LY, which publishes this survey.&lt;/p&gt;
&lt;p&gt;CoDesign is a free local MCP server that gives an agent you already use a full design engine. It entered technical preview in June 2026.&lt;/p&gt;
&lt;p&gt;Instead of going to CoDesign, you install it into your own agent with one command, and the agent then performs design operations in conversation with you. IMG.LY documents the install for &lt;a href=&quot;https://img.ly/codesign/&quot;&gt;Claude Code and Codex&lt;/a&gt;, and any client that can run an MCP server locally works the same way. CoDesign can generate, rebrand, resize, edit, localize, import, and judge, the last checking a design against your brand rules. A representative chain: import a PSD, rebrand it, resize it to a dozen formats, check it against the rules, export.&lt;/p&gt;
&lt;p&gt;Output is a structured scene rather than a flat image. It opens in an editor where you select any element and change it without regenerating anything. Underneath is CE.SDK, IMG.LY’s production design SDK. It imports PSD, IDML, PDF, and PPTX, and exports PDF with CMYK, bleed, and ICC profiles when a printer needs them. The agent, your code, and a human editor all work on the same file.&lt;/p&gt;
&lt;p&gt;We ran the Northwind prompt through Claude Code. Before touching the canvas it asked where the product photo should come from, what accent color to use, what kind of cream we meant, and it drafted the copy for approval. In normal use you would run this inside your marketing documentation and brand assets, so your coding agent already carries a pile of context and can answer most of that itself. We gave it no brand context at all, to keep the comparison fair with the others.&lt;/p&gt;
&lt;p&gt;It left us waiting a bit longer, though that is not quite fair to say, since it works in tandem with the coding agent. It was thorough and diligent. It ran a self-check against a set of axes before handing the design back, and because we supplied no brand context those came back blank. It passed, and returned one editable master file.&lt;/p&gt;
&lt;p&gt;When it needed the product image it opened the IMG.LY dashboard, where the AI Gateway picked the model for it and credited the generation, which was frictionless.&lt;/p&gt;
&lt;p&gt;The design that came back is fairly conservative and does not assume too much: some copy, “now pouring, limited first batch”, a CTA. Fairly bare bones, but it works. What matters is that it stuck perfectly to the specification. The square asset, the story asset, and a PDF carrying its color space and a resolution we could hand to a printer. Edit one element, tell it to regenerate, and it propagates: same colors, same logo, same product image across every variant.&lt;/p&gt;
&lt;p&gt;All three formats came out of a single master file, generated in one pass from the same brief:&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;The three Northwind formats CoDesign returned from one brief: the DIN A5 flyer on the left, then the square Instagram post and the vertical story side by side, all carrying the same headline, colors and product photo&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 2168px) 2168px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;2168&quot; height=&quot;900&quot; src=&quot;https://img.ly/_astro/codesign-outputs.ByTlv-EZ_1l0LHA.webp&quot; srcset=&quot;/_astro/codesign-outputs.ByTlv-EZ_1WiJfB.webp 640w, /_astro/codesign-outputs.ByTlv-EZ_ZMeRqI.webp 750w, /_astro/codesign-outputs.ByTlv-EZ_Z2c0iCf.webp 828w, /_astro/codesign-outputs.ByTlv-EZ_Z2hT6nK.webp 1080w, /_astro/codesign-outputs.ByTlv-EZ_Z1t3588.webp 1280w, /_astro/codesign-outputs.ByTlv-EZ_1T2Tz8.webp 1668w, /_astro/codesign-outputs.ByTlv-EZ_y88Cz.webp 2048w, /_astro/codesign-outputs.ByTlv-EZ_1l0LHA.webp 2168w&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;The CoDesign editor with the Northwind master file open: the headline text layer selected for a direct edit, a shapes library on the left, and a Download PDF button in the toolbar&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1497px) 1497px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1497&quot; height=&quot;827&quot; src=&quot;https://img.ly/_astro/codesign-ui.Cr8GPPF2_Kvoq4.webp&quot; srcset=&quot;/_astro/codesign-ui.Cr8GPPF2_1S0RsV.webp 640w, /_astro/codesign-ui.Cr8GPPF2_Z1IJheT.webp 750w, /_astro/codesign-ui.Cr8GPPF2_1RszbR.webp 828w, /_astro/codesign-ui.Cr8GPPF2_ZhArp0.webp 1080w, /_astro/codesign-ui.Cr8GPPF2_1FlAe7.webp 1280w, /_astro/codesign-ui.Cr8GPPF2_Kvoq4.webp 1497w&quot;&gt;&lt;/p&gt;
&lt;p&gt;The local server is free to install and run with no account needed to start. An optional free account adds AI image generation, which routes out through IMG.LY’s AI Gateway. Paid licensing applies when you host the server for other people or embed it in a product.&lt;/p&gt;
&lt;p&gt;Best fit is anyone running an agent-first local workflow who needs production-grade editing, print-ready files, and brand consistency across every variant. As a technical preview it is early, and behavior can change between releases.&lt;/p&gt;
&lt;h2 id=&quot;delegating-design-inside-a-larger-task&quot;&gt;Delegating design inside a larger task&lt;/h2&gt;
&lt;h3 id=&quot;manus&quot;&gt;Manus&lt;/h3&gt;
&lt;p&gt;Disclosure: Manus is an IMG.LY customer. This is also the one entry we did not run ourselves, so unlike the five above it rests on public sources rather than hands-on testing.&lt;/p&gt;
&lt;p&gt;Manus is a general-purpose autonomous agent with a design workspace inside it. Design is one capability among research, app building, and document production.&lt;/p&gt;
&lt;p&gt;Manus decomposes a goal into subtasks, browses for context, runs for long stretches without supervision, and continues while you are offline. For design work it researches before designing, then produces assets you refine through object-level edits: click an element to change colors, swap a background, or reshape it, without regenerating the whole asset.&lt;/p&gt;
&lt;p&gt;Output covers images, video, 3D assets, and presentations. Its public documentation does not describe a layered source file of the kind a dedicated design tool exposes, so you edit objects rather than a design file. General agents are closing the gap on design output faster than anything else in this survey.&lt;/p&gt;
&lt;p&gt;The free tier includes 300 daily refresh credits. Paid plans run $20, $40, and $200 monthly for 4,000, 8,000, and 40,000 credits. Reviews describe unpredictable credit burn on complex tasks.&lt;/p&gt;
&lt;p&gt;Best fit is founders and operators who want one agent handling research, documents, and adequate design output, and who value autonomy over design-tool depth.&lt;/p&gt;
&lt;h2 id=&quot;which-tools-leave-you-a-real-editor&quot;&gt;Which tools leave you a real editor&lt;/h2&gt;
&lt;p&gt;Three tools in this survey give you a real editor after generation. Figma Agent puts native layers on the Figma canvas. Canva’s assistant produces fully editable Canva designs. CoDesign returns a structured scene that opens in a full editor. In all three you select an element and change it, instead of rewriting the prompt and hoping the next version keeps the parts you liked.&lt;/p&gt;
&lt;p&gt;Lovart sits between. It has a canvas and in-place editing, and precision work still tends to finish in Photoshop. Stitch hands off to Figma or to code. Manus edits at the object level, which suits an agent whose remit is much wider than design.&lt;/p&gt;
&lt;p&gt;Revision takes longer than the first draft, and tools that keep you in control of the file absorb those cycles.&lt;/p&gt;
&lt;h2 id=&quot;destination-app-or-capability-in-your-stack&quot;&gt;Destination app or capability in your stack&lt;/h2&gt;
&lt;p&gt;Five of the six tools here are destinations. You open Figma, Canva, Lovart, Stitch, or Manus, do the work there, and take the result away. That model is familiar and it works.&lt;/p&gt;
&lt;p&gt;The sixth, CoDesign, is a design engine that installs into the agent you already use, so the design step happens inside your existing workflow.&lt;/p&gt;
&lt;p&gt;That split is narrower than it first appears. Figma opened an MCP server in March 2026, so external coding agents can drive its canvas. Canva shipped an MCP server in February 2026 that exposes design generation inside ChatGPT, Claude, and Microsoft Copilot, and says more than 12 million designs have been created that way. Being reachable from your agent is not unique to CoDesign.&lt;/p&gt;
&lt;p&gt;The tools differ in what they hand back. Drive Figma’s MCP and you get a Figma file, which is useful if your team lives in Figma. Drive Canva’s and you get a Canva design. Drive CoDesign and you get a portable scene file, PDF with CMYK and bleed, PSD and IDML import, and no platform you have to keep an account with. For teams whose output ends at a printer or inside their own product, that portability decides it. It matters much less if your work already lives in Figma or Canva.&lt;/p&gt;
&lt;h2 id=&quot;how-to-choose&quot;&gt;How to choose&lt;/h2&gt;
&lt;p&gt;Start with what you are making. UI and app screens point to Stitch for speed or the Figma Agent for depth. Campaign and brand assets point to Lovart for volume from one brief, Canva for teams already there, or &lt;a href=&quot;https://img.ly/codesign/&quot;&gt;CoDesign&lt;/a&gt; for output that has to stay editable and reach print. If design is a side task inside something larger, Manus covers it.&lt;/p&gt;
&lt;p&gt;Then check the exit path before you commit, because it is the part you cannot change later. Ask where the file lives, whether you can edit it without regenerating, and what happens when it needs to leave the tool.&lt;/p&gt;
&lt;h2 id=&quot;frequently-asked-questions&quot;&gt;Frequently asked questions&lt;/h2&gt;
&lt;h3 id=&quot;what-is-an-ai-design-agent&quot;&gt;What is an AI design agent?&lt;/h3&gt;
&lt;p&gt;An AI design agent plans and executes multi-step design work from a goal, producing usable design output such as files, screens, or code. &lt;a href=&quot;https://img.ly/blog/what-is-a-design-agent/&quot;&gt;What Is an AI Design Agent?&lt;/a&gt; goes into where the line falls.&lt;/p&gt;
&lt;h3 id=&quot;how-is-an-ai-design-agent-different-from-an-image-generator&quot;&gt;How is an AI design agent different from an image generator?&lt;/h3&gt;
&lt;p&gt;An image generator returns a picture for a prompt. An agent plans a sequence of steps, performs design operations, and produces work you can carry forward.&lt;/p&gt;
&lt;h3 id=&quot;are-ai-design-agents-free&quot;&gt;Are AI design agents free?&lt;/h3&gt;
&lt;p&gt;Three of the six have a free way in: Stitch while it is in Labs, Manus on its free tier, and CoDesign’s local server. Figma’s agent is free during beta but needs a paid seat. Lovart and full Canva use are subscriptions.&lt;/p&gt;
&lt;h3 id=&quot;do-ai-design-agents-replace-designers&quot;&gt;Do AI design agents replace designers?&lt;/h3&gt;
&lt;p&gt;No. Every tool in this survey keeps a person reviewing the output, and all six advertise editing after generation.&lt;/p&gt;
&lt;h3 id=&quot;which-ai-design-agents-produce-editable-files&quot;&gt;Which AI design agents produce editable files?&lt;/h3&gt;
&lt;p&gt;Figma Agent, Canva’s assistant, and CoDesign keep element-level editability in a design file. Lovart offers in-canvas editing, though external finishing is common. Manus edits objects in place. Stitch relies on its Figma and code exports.&lt;/p&gt;
&lt;h3 id=&quot;can-i-use-a-design-agent-inside-claude-or-chatgpt&quot;&gt;Can I use a design agent inside Claude or ChatGPT?&lt;/h3&gt;
&lt;p&gt;Yes, through MCP. Canva exposes design generation in ChatGPT, Claude, and Copilot. Figma opened its MCP server to external coding agents. CoDesign runs as a local MCP server in coding agents such as Claude Code and Codex.&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="/_astro/hero.DBk-VO6j.webp" medium="image"/><category>AI</category><category>Insights</category></item><item><title>Introducing IMG.LY AI Benchmarks</title><link>https://img.ly/blog/introducing-imgly-ai-benchmarks/</link><guid isPermaLink="true">https://img.ly/blog/introducing-imgly-ai-benchmarks/</guid><description>We ran 15 image models through 37 production design prompts and published every score, every prompt and every generated image. Why we built it, why the scores are weighted per job, and the result that surprised us most.</description><pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/&quot;&gt;IMG.LY GenAI Benchmarks&lt;/a&gt; is live. 15 text-to-image models, 37 prompts taken from real design jobs, three fixed seeds each, 1,662 generated images, roughly 7,400 scores. Every prompt, every score, every generated image and the full &lt;a href=&quot;https://img.ly/ai-benchmarks/methodology/&quot;&gt;methodology&lt;/a&gt; are public.&lt;/p&gt;
&lt;h2 id=&quot;why-we-ran-it&quot;&gt;Why we ran it&lt;/h2&gt;
&lt;p&gt;Our customers embed image generation in products that ship physical, commercial output. Web-to-print storefronts, merch platforms, campaign tooling, template systems. When they asked us which model to use, the honest answer was a shrug and a link to a leaderboard that measures something else.&lt;/p&gt;
&lt;p&gt;The existing arenas rank aesthetic preference by crowd vote. That’s a signal, sure, but not the one most relevant to these use cases. They want to know whether the image has a real alpha channel, or whether users got the hex they asked for and whether a word is spelled right. A 10,000-piece personalization run needs to hold the character across every variant. Those are the failures that cost you a reprint or a support ticket, and none of them show up in a preference ranking.&lt;/p&gt;
&lt;p&gt;So we built the one that answers our customers’ question. Which model should you build on for production design work.&lt;/p&gt;
&lt;h2 id=&quot;why-some-of-it-is-judged-not-measured&quot;&gt;Why some of it is judged, not measured&lt;/h2&gt;
&lt;p&gt;Plenty of what matters here is measurable, and we measure it. Native output size, wall-clock latency and price per image come straight off the run. Transparency is a file-level check for a real alpha channel rather than a painted-on white rectangle. Color accuracy is the CIEDE2000 distance between the hex the prompt required and the color that came back.&lt;/p&gt;
&lt;p&gt;The rest needs a judgment call. Whether an image satisfies “three blue cubes stacked on the left and one red sphere on the right” is a checklist, and something has to look at the picture and tick it off. Whether a banner leaves a clean area for a headline has a right answer, but not a computable one.&lt;/p&gt;
&lt;p&gt;For those criteria a vision model does the judging, against a checklist written per prompt before the run. We added two rules to make the VLM’s judgment more transparent and control for arbitrariness. Every judgment records a written rationale beside the score, so you can read why an image lost a point and tell us we’re wrong. The criterion that really does need a person, design-readiness, the “would a professional ship this with under five minutes of cleanup” question, is surfaced to a human in the loop (that would be yours truly, the author).&lt;/p&gt;
&lt;p&gt;The rest of the protocol is the boring part that makes the numbers worth anything. Identical prompts for every model, default parameters, fixed seeds, no per-model prompt tuning ever, a versioned suite, and published scores that are never quietly restated. Partner models get no special treatment, and the Gateway links stay well away from the rankings.&lt;/p&gt;
&lt;h2 id=&quot;why-we-split-it-by-use-case&quot;&gt;Why we split it by use case&lt;/h2&gt;
&lt;p&gt;Average our eight criteria with equal weight and &lt;a href=&quot;https://img.ly/ai-benchmarks/models/nano-banana-2-lite/&quot;&gt;Nano Banana 2 Lite&lt;/a&gt; leads at 3.87 out of 5, while last place, at 2.55, goes to &lt;a href=&quot;https://img.ly/ai-benchmarks/models/nano-banana-pro/&quot;&gt;Nano Banana Pro&lt;/a&gt;, the most expensive flagship in the set. That’s not a data error. Nano Banana Pro has the best brand-color fidelity in the field and renders every required string in the suite exactly. It also costs 28 times more per image than the cheapest model and takes 23 seconds. An unweighted average treats all of that as equally important, so it buries the model with the best output.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;Cost per image plotted against the unweighted overall score for all 15 models. The cheap, fast models sit high on the left; the most expensive flagships sit low on the right. Stable Diffusion 1.5, FLUX.2 and Nano Banana 2 Lite form the Pareto frontier.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1200px) 1200px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1200&quot; height=&quot;660&quot; src=&quot;https://img.ly/_astro/cost-vs-quality.B8vY-NEd_2iGGAu.svg&quot; srcset=&quot;/_astro/cost-vs-quality.B8vY-NEd_4yLYC.svg 640w, /_astro/cost-vs-quality.B8vY-NEd_ZmGGij.svg 750w, /_astro/cost-vs-quality.B8vY-NEd_Z2txY4l.svg 828w, /_astro/cost-vs-quality.B8vY-NEd_KI4Im.svg 1080w, /_astro/cost-vs-quality.B8vY-NEd_2iGGAu.svg 1200w&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;This is the chart everyone asks for, and its vertical axis is the number you should not rank on. Explore the &lt;a href=&quot;https://img.ly/ai-benchmarks/models/&quot;&gt;live, interactive version&lt;/a&gt;, where every point links to that model’s scores.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Which is what leaderboards do. So we publish the same measurements &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/print-files/&quot;&gt;per use case&lt;/a&gt; as well, with a weight matrix per job and the reasoning behind each matrix written out. Print files weight resolution and color heavily, because a 1024px generation is a 3.4-inch print and an out-of-gamut red is a reprint. Personalization at scale weights cost and adherence, because unit economics decide the run. Merch weights the alpha channel above everything else, because a sticker without a clean cutout isn’t a sticker.&lt;/p&gt;
&lt;p&gt;Same numbers, different question, different answer. That’s why the section opens by asking what you’re building instead of handing you a ranked list.&lt;/p&gt;
&lt;h2 id=&quot;the-result-that-surprised-us-most&quot;&gt;The result that surprised us most&lt;/h2&gt;
&lt;p&gt;Reweighting moves models the length of the table.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/gpt-image-1-5/&quot;&gt;GPT Image 1.5&lt;/a&gt; sits twelfth of fifteen on the unweighted blend, dragged down by cost and latency. Reweight for &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/merch-and-stickers/&quot;&gt;merch and stickers&lt;/a&gt;, where transparency carries 35 percent of the matrix, and it wins outright at 3.72, nearly a full point clear of second place. It’s the only model in the run that reliably returns a genuine alpha channel. Thirteen of the fifteen return none at all, and a few paint a fake checkerboard into the pixels, which looks like transparency right up until it reaches a printer.&lt;/p&gt;
&lt;p&gt;On a general leaderboard, the best model for one of our customers’ most common jobs reads as a mid-table also-ran. That’s the argument for the whole project.&lt;/p&gt;
&lt;p&gt;The typography column surprised us a second time. Ideogram has the strongest public reputation for rendering text, and in our run it ranks twelfth of fifteen on measured text accuracy, behind several general-purpose flagships that now render every required string in the suite exactly. Route typographic work to a specialist on reputation alone and you can end up with worse type than the model you already call by default.&lt;/p&gt;
&lt;p&gt;Two more capability gaps got written up as standalone &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/&quot;&gt;findings&lt;/a&gt;. &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/brand-color/&quot;&gt;Brand color&lt;/a&gt;, where the best model manages 3.67 out of 5 and most of the field sits below 3. And &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/consistency/&quot;&gt;consistency&lt;/a&gt;, where no model holds a described character across three scenes and the ceiling is 3.89.&lt;/p&gt;
&lt;h2 id=&quot;what-this-means-for-an-editor&quot;&gt;What this means for an editor&lt;/h2&gt;
&lt;p&gt;We build editors, so our conclusion isn’t neutral. The data still points where it points.&lt;/p&gt;
&lt;p&gt;Read the four findings together and they describe the same problem four times. The cut-out is missing, the color is close but not the color in the brand book, the headline is right until it isn’t, and the character drifts between scenes. None of these are bugs waiting on the next model release. They’re what generation is, a probabilistic first draft. Getting from that draft to something a customer can order takes a handful of deterministic corrections, and somebody has to be able to make them.&lt;/p&gt;
&lt;p&gt;That’s an editor’s job. Background removal and edge cleanup close the transparency gap. Brand kits and exact color values close the color gap. Editable text layers turn a wrong character into a two-second fix instead of a regeneration. Reusable assets and templates carry identity that regeneration won’t. Our &lt;a href=&quot;https://img.ly/use-cases/ai-editor/&quot;&gt;AI Editor&lt;/a&gt; gives your users those controls over whatever model produced the image, and the &lt;a href=&quot;https://img.ly/ai-gateway/&quot;&gt;AI Gateway&lt;/a&gt; makes routing per job, which the data says you should be doing, a config change instead of another integration.&lt;/p&gt;
&lt;p&gt;We’d rather make that case with numbers anyone can check, including the ones that make models we partner with look bad.&lt;/p&gt;
&lt;h2 id=&quot;caveats-and-whats-next&quot;&gt;Caveats, and what’s next&lt;/h2&gt;
&lt;p&gt;This is a pilot dataset. Quality criteria are judged by a vision model, not yet by the blind expert panel, and design-readiness isn’t scored at all. The weight matrices are provisional. When the frozen suite lands, scores restate once and the suite version changes; after that, nothing gets restated without a clearly labeled new methodology version.&lt;/p&gt;
&lt;p&gt;The roster will grow. Image editing and video are the obvious next modalities, and we plan to re-run the suite whenever a notable model ships.&lt;/p&gt;
&lt;p&gt;Go &lt;a href=&quot;https://img.ly/ai-benchmarks/&quot;&gt;pick a use case&lt;/a&gt; and see which model wins your job. If the numbers disagree with your own experience, the prompts and the raw images are all sitting there, and we’d like to hear about it.&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="/_astro/hero.CyGIWvSf.webp" medium="image"/><category>AI</category><category>Image Gen</category><category>Insights</category></item><item><title>The Best GenAI Models for Web-to-Print: A Buyer&apos;s Guide</title><link>https://img.ly/blog/best-generative-ai-models-for-web-to-print/</link><guid isPermaLink="true">https://img.ly/blog/best-generative-ai-models-for-web-to-print/</guid><description>The best general image model can still fail at print. What decides it (vector output, print resolution, CMYK behavior, legible text) is exactly what general model rankings ignore. This guide gives you eleven print-specific evaluation criteria, a scoring rubric, and a model-by-model read on where each one wins.</description><pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;No single generative model is “best” for web-to-print. The right pick depends on the specific job, and print has jobs that screen-first leaderboards don’t measure. A model that tops a general image-quality benchmark can still be useless for print if it can’t render legible text, can’t output vector, or shifts hard out of CMYK gamut.&lt;/p&gt;
&lt;p&gt;This guide gives you the criteria that actually matter for print, a scoring rubric to apply them, and a model-by-model read on where each one wins. The goal is to route each task to the model that fits it, which is why a &lt;strong&gt;model-agnostic&lt;/strong&gt; (“bring your own model”) integration matters more than any one provider’s marketing.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How to read this guide.&lt;/strong&gt; The criteria and rubric are the durable part. The model roster moves monthly: treat the specific model notes as a snapshot, and re-run the rubric against current versions before you commit. The per-model benchmark scores come from IMG.LY’s &lt;a href=&quot;https://img.ly/ai-benchmarks/&quot;&gt;GenAI benchmark suite&lt;/a&gt;, which is now live: a pilot dataset you can explore today, re-run whenever new models drop. The gated PDF report that packages the weighted per-use-case totals and the full test set is the piece we’re opening to early-access subscribers first (sign-up at the end of this guide).&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;why-print-needs-its-own-criteria&quot;&gt;Why print needs its own criteria&lt;/h2&gt;
&lt;p&gt;A web-to-print product &lt;strong&gt;sells a physical object&lt;/strong&gt;. That changes what “good” means:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The output is measured at 300 DPI on a substrate, not at 72 PPI on a retina screen.&lt;/li&gt;
&lt;li&gt;Color is CMYK and spot, not sRGB.&lt;/li&gt;
&lt;li&gt;Logos and type must scale and stay crisp: vector, not pixels.&lt;/li&gt;
&lt;li&gt;The customer is a non-designer, so the model has to be controllable enough to stay inside a template.&lt;/li&gt;
&lt;li&gt;And because it’s a commercial product sold to a customer, the &lt;strong&gt;licensing&lt;/strong&gt; of the output carries real legal exposure.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;General benchmarks (LMArena-style human preference, aesthetic scores) miss most of this. The eleven criteria below are built around it.&lt;/p&gt;
&lt;h2 id=&quot;the-eleven-evaluation-criteria&quot;&gt;The eleven evaluation criteria&lt;/h2&gt;
&lt;h3 id=&quot;1-output-type-and-format-raster-vs-vector&quot;&gt;1. Output type and format: raster vs. vector&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters for print:&lt;/strong&gt; Logos, icons, line art, and type need to scale to any size and print crisply. Raster can’t; vector can. A model that outputs (or can be cleanly traced to) &lt;strong&gt;native SVG / vector&lt;/strong&gt; is uniquely valuable for the print-specific elements that screen apps don’t care about.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Can it produce vector natively? If not, how cleanly does its output vectorize? Does the vector survive into PDF/X as real paths?&lt;/p&gt;
&lt;h3 id=&quot;2-resolution-and-maximum-dimensions&quot;&gt;2. Resolution and maximum dimensions&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; A ~1024px generation is ~3.4” at 300 DPI. Large-format and even A4 need far more. Native max resolution, and how well the model holds detail when upscaled, determines whether you can print the output at size without a quality penalty.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Native max output dimensions; detail retention at 4×/8× upscale; effective printable size at 300 DPI.&lt;/p&gt;
&lt;h3 id=&quot;3-visual-quality-and-fidelity&quot;&gt;3. Visual quality and fidelity&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; The baseline. Realism, coherence, absence of artifacts, anatomical and structural correctness. Crucially, how it holds up under print’s unforgiving close inspection: banding, mushy detail, and plasticky textures show more on paper than on screen.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Side-by-side preference on print-representative prompts; artifact rate; detail at print scale.&lt;/p&gt;
&lt;h3 id=&quot;4-text-and-typography-rendering&quot;&gt;4. Text and typography rendering&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Print is full of words, and most models garble text inside images. For anything where text-in-image is unavoidable (a generated poster, a label motif), legibility and kerning are make-or-break. Best practice is still to set real type as a separate layer, but some jobs need it baked in. Our &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/text-reliability/&quot;&gt;text-reliability finding&lt;/a&gt; measures exactly this. The top of the field has largely solved Latin headlines: three models rendered every required string in our suite exactly. Small type and non-Latin scripts have not been solved, and the model most famous for typography ranks twelfth of fifteen. One wrong character still means regenerating the whole image, which is the real argument for keeping type on its own layer. Compare every model on the &lt;a href=&quot;https://img.ly/ai-benchmarks/prompts/t04-wordmark/&quot;&gt;same wordmark prompt&lt;/a&gt; to see the spread.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Legibility of short and long strings; correct spelling; kerning; multi-line; non-Latin scripts.&lt;/p&gt;
&lt;h3 id=&quot;5-prompt-adherence-and-controllability&quot;&gt;5. Prompt adherence and controllability&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Non-designers need predictable results, and templates need the output to land where it’s told. Controllability spans prompt adherence, &lt;strong&gt;image-to-image / inpainting / outpainting&lt;/strong&gt;, style and reference-image conditioning, and seed stability for repeatable results.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Does it honor constraints? Quality of edits (generative fill, expand, object removal)? Reference-image fidelity? Same-seed reproducibility?&lt;/p&gt;
&lt;h3 id=&quot;6-color-accuracy-and-cmyk-readiness&quot;&gt;6. Color accuracy and CMYK-readiness&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Vivid RGB outputs clip hard in CMYK and shift on press. Models that stay closer to printable gamut, and outputs that convert predictably, mean fewer surprises and less soft-proofing friction. Even before the CMYK conversion, hitting an exact brand hex in RGB is hard: our &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/brand-color/&quot;&gt;brand-color finding&lt;/a&gt; measures color accuracy with CIEDE2000 against the required hex, and no model reliably lands the color in your brand book.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Gamut coverage vs. CMYK; shift magnitude after ICC conversion; behavior on brand-critical colors and near-whites/blacks.&lt;/p&gt;
&lt;h3 id=&quot;7-transparency-and-background-handling&quot;&gt;7. Transparency and background handling&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Merch, stickers, die-cuts, and product compositing need clean subjects on transparent backgrounds and accurate edges (hair, glass, fine detail). Native transparent output and clean cutout edges save a manual masking step. In the models themselves this is close to unsolved: our &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/transparency/&quot;&gt;transparency finding&lt;/a&gt; shows 13 of 15 tested models emit no real alpha channel at all, which is exactly why a background-removal step belongs in the pipeline rather than a prayer that the model returns a clean cutout.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Native alpha output; edge quality on hard cases; cutout-path readiness.&lt;/p&gt;
&lt;h3 id=&quot;8-responsiveness--latency&quot;&gt;8. Responsiveness / latency&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Two very different bars. In an &lt;strong&gt;interactive editor&lt;/strong&gt;, anything over a few seconds breaks flow. In a &lt;strong&gt;&lt;a href=&quot;https://img.ly/use-cases/variable-data-printing/&quot;&gt;batch/VDP run&lt;/a&gt;&lt;/strong&gt;, throughput and concurrency matter more than per-image speed. The right model differs by context.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; P50/P95 latency at your resolution; throughput under concurrency; cold-start behavior.&lt;/p&gt;
&lt;h3 id=&quot;9-cost-per-generation&quot;&gt;9. Cost per generation&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Unit economics. A few cents per image is invisible for interactive one-offs and brutal across a 100k-recipient VDP run. Price has to be weighed against quality &lt;em&gt;for the specific job&lt;/em&gt;. You don’t pay flagship prices for a draft thumbnail.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Cost per image at your resolution/step settings; cost of edit vs. generate; volume pricing.&lt;/p&gt;
&lt;h3 id=&quot;10-consistency-and-repeatability&quot;&gt;10. Consistency and repeatability&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Brand work needs the same input to yield the same (or controllably similar) output, and product likenesses must not drift. Seed control, character and style consistency, and low variance separate “brand-safe” models from “slot-machine” ones. Our &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/consistency/&quot;&gt;consistency finding&lt;/a&gt; puts numbers on the drift: no model holds a described character across scenes, so identity has to come from templates and editing, not from a regeneration lottery.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Variance across runs at fixed seed; character/style consistency across a set; drift on iterative edits.&lt;/p&gt;
&lt;h3 id=&quot;11-commercial-licensing-and-ip-safety&quot;&gt;11. Commercial licensing and IP safety&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; You’re selling the printed output. Commercial-use rights, training-data provenance, indemnification, and content/safety filtering are legal exposure, not fine print. A model that’s brilliant but legally murky is a non-starter for a product customers resell.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Commercial-use terms; indemnity; provenance/transparency; enterprise and data-handling terms; content-filter behavior.&lt;/p&gt;
&lt;h2 id=&quot;the-scoring-rubric&quot;&gt;The scoring rubric&lt;/h2&gt;
&lt;p&gt;Score each model &lt;strong&gt;1–5&lt;/strong&gt; on each criterion, then weight by your use case. Suggested weights for the three most common web-to-print contexts:&lt;/p&gt;













































































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Criterion&lt;/th&gt;&lt;th&gt;Interactive design tool&lt;/th&gt;&lt;th&gt;High-volume VDP / automation&lt;/th&gt;&lt;th&gt;Merch / POD&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;1. Vector / SVG output&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;Medium&lt;/td&gt;&lt;td&gt;Medium&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;2. Resolution / max size&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;3. Visual quality&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;Medium&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;4. Text rendering&lt;/td&gt;&lt;td&gt;Medium&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;Medium&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;5. Controllability / editing&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;6. CMYK-readiness&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;7. Transparency / cutout&lt;/td&gt;&lt;td&gt;Medium&lt;/td&gt;&lt;td&gt;Low&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;8. Responsiveness&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;Medium (throughput)&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;9. Cost per generation&lt;/td&gt;&lt;td&gt;Medium&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;Medium&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;10. Consistency&lt;/td&gt;&lt;td&gt;Medium&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;11. Licensing / IP&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Scoring scale:&lt;/strong&gt; 1 = unusable for print · 2 = weak · 3 = workable with mitigation · 4 = strong · 5 = best-in-class.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Drop your own numbers into this rubric. The measured scores in the &lt;a href=&quot;https://img.ly/blog/best-generative-ai-models-for-web-to-print//#the-imgly-genai-benchmark-measured-results&quot;&gt;results table later in this guide&lt;/a&gt; give you a starting hypothesis to test; the qualitative notes are not a substitute for your own runs on your prompts.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;models-by-job&quot;&gt;Models by job&lt;/h2&gt;
&lt;p&gt;New models ship faster than any ranking can keep up with, so this section is organized &lt;strong&gt;by job&lt;/strong&gt;, the way you actually route requests. Models are drawn from the roster currently wired into CE.SDK’s AI plugins.&lt;/p&gt;
&lt;h3 id=&quot;job-a-text-to-image-generate-a-new-asset&quot;&gt;Job A: Text-to-image (generate a new asset)&lt;/h3&gt;
&lt;p&gt;Candidates: &lt;strong&gt;Recraft V3&lt;/strong&gt;, &lt;strong&gt;Recraft 20B&lt;/strong&gt;, &lt;strong&gt;Seedream V4&lt;/strong&gt;, &lt;strong&gt;Nano Banana&lt;/strong&gt;, &lt;strong&gt;Nano Banana Pro&lt;/strong&gt;, &lt;strong&gt;GPT Image 1&lt;/strong&gt;, &lt;strong&gt;Ideogram V3&lt;/strong&gt;.&lt;/p&gt;















































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Stand-out for print&lt;/th&gt;&lt;th&gt;Watch-outs&lt;/th&gt;&lt;th&gt;Benchmark&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Recraft V3&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;The print specialist: &lt;strong&gt;native vector/SVG generation&lt;/strong&gt;, strong text rendering, brand-style controls. Often the single most print-relevant text-to-image model.&lt;/td&gt;&lt;td&gt;Verify max raster resolution for large-format jobs.&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/recraft-v3/&quot;&gt;3.29 / 5&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Recraft 20B&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Faster, cheaper Recraft tier. Good for interactive iteration before a final Recraft V3 render.&lt;/td&gt;&lt;td&gt;Quality step-down vs. V3.&lt;/td&gt;&lt;td&gt;not yet in the suite&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Seedream V4&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;High visual fidelity and resolution; strong general-purpose hero imagery.&lt;/td&gt;&lt;td&gt;Text rendering and vector are not its strength.&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/seedream-4-5/&quot;&gt;3.80 / 5&lt;/a&gt; (measured on Seedream 4.5)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Nano Banana / Pro&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Fast, controllable, strong prompt adherence; Pro for higher quality. Good interactive default.&lt;/td&gt;&lt;td&gt;Confirm commercial-licensing terms for resale.&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/nano-banana-2/&quot;&gt;3.24 / 5&lt;/a&gt; · Pro &lt;a href=&quot;https://img.ly/ai-benchmarks/models/nano-banana-pro/&quot;&gt;2.55 / 5&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;GPT Image 1&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Strong instruction-following and &lt;strong&gt;comparatively reliable in-image text&lt;/strong&gt;; good for layout-aware generation.&lt;/td&gt;&lt;td&gt;Latency and cost on the higher side for interactive use.&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/gpt-image-1-5/&quot;&gt;3.07 / 5&lt;/a&gt; (measured on GPT Image 1.5)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Ideogram V3&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Built around typography and still a credible pick for poster and typographic art, but our measured text column no longer puts it on top.&lt;/td&gt;&lt;td&gt;Ranks twelfth of fifteen on text accuracy in our run.&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/ideogram-v3/&quot;&gt;2.97 / 5&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Routing heuristic:&lt;/strong&gt; logos, icons, and typographic art go to &lt;strong&gt;Recraft V3&lt;/strong&gt; (vector; the &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/print-files/&quot;&gt;print files use case&lt;/a&gt;). Words-in-image goes to whichever flagship currently tops the measured text column, which in our July 2026 run means &lt;strong&gt;GPT Image 1.5&lt;/strong&gt;, &lt;strong&gt;FLUX.2 [pro]&lt;/strong&gt; or &lt;strong&gt;Nano Banana Pro&lt;/strong&gt; rather than a typography specialist (the &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/text-heavy-designs/&quot;&gt;text-heavy designs use case&lt;/a&gt;). Photoreal hero imagery goes to &lt;strong&gt;Seedream V4&lt;/strong&gt; or &lt;strong&gt;Nano Banana Pro&lt;/strong&gt; (the &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/product-ads/&quot;&gt;product ads use case&lt;/a&gt;). Fast interactive drafts go to &lt;strong&gt;Recraft 20B / Nano Banana&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id=&quot;job-b-image-editing-adapt-a-customers-asset&quot;&gt;Job B: Image editing (adapt a customer’s asset)&lt;/h3&gt;
&lt;p&gt;Candidates: &lt;strong&gt;Flux Pro Kontext&lt;/strong&gt;, &lt;strong&gt;Flux Pro Kontext Max&lt;/strong&gt;, &lt;strong&gt;Nano Banana Edit&lt;/strong&gt;, &lt;strong&gt;Nano Banana Pro Edit&lt;/strong&gt;, &lt;strong&gt;Qwen Image Edit&lt;/strong&gt;, &lt;strong&gt;Gemini Flash Edit&lt;/strong&gt;, &lt;strong&gt;Seedream V4 Edit&lt;/strong&gt;, &lt;strong&gt;GPT Image 1&lt;/strong&gt;, &lt;strong&gt;Ideogram V3 Remix&lt;/strong&gt;.&lt;/p&gt;









































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Stand-out for print&lt;/th&gt;&lt;th&gt;Watch-outs&lt;/th&gt;&lt;th&gt;Benchmark&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Flux Pro Kontext / Max&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Strong &lt;strong&gt;instruction-based editing&lt;/strong&gt; with high subject preservation: change one thing without wrecking the rest. Good for generative fill/expand and targeted edits.&lt;/td&gt;&lt;td&gt;Max tier costs more; budget per edit.&lt;/td&gt;&lt;td&gt;editing suite planned&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Nano Banana Edit / Pro Edit&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Fast, controllable edits; good interactive default for object removal and swaps.&lt;/td&gt;&lt;td&gt;Verify edge quality on fine detail.&lt;/td&gt;&lt;td&gt;editing suite planned&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Qwen Image Edit&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Strong editing with good text handling during edits.&lt;/td&gt;&lt;td&gt;Validate commercial terms.&lt;/td&gt;&lt;td&gt;editing suite planned&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Gemini Flash Edit&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Low latency&lt;/strong&gt;: the responsiveness pick for interactive editing.&lt;/td&gt;&lt;td&gt;Quality trade-off vs. heavier models.&lt;/td&gt;&lt;td&gt;editing suite planned&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Seedream V4 Edit&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;High-fidelity edits matching its generation quality.&lt;/td&gt;&lt;td&gt;Heavier; weigh latency.&lt;/td&gt;&lt;td&gt;editing suite planned&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The benchmark currently measures text-to-image only; the editing models above join the leaderboard when the frozen suite adds an editing track, so their cells read “editing suite planned” rather than a score.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Routing heuristic:&lt;/strong&gt; precise “change X, keep everything else” goes to &lt;strong&gt;Flux Kontext&lt;/strong&gt;. Fast interactive edits go to &lt;strong&gt;Gemini Flash Edit / Nano Banana Edit&lt;/strong&gt;. Edits that must preserve in-image text go to &lt;strong&gt;Qwen Image Edit&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id=&quot;job-c-specialized-print-operations&quot;&gt;Job C: Specialized print operations&lt;/h3&gt;
&lt;p&gt;These aren’t single “models” so much as task pipelines, but they belong in the routing table:&lt;/p&gt;



































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Task&lt;/th&gt;&lt;th&gt;Approach / model&lt;/th&gt;&lt;th&gt;Note&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Vectorize (raster → SVG)&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Vectorize plugin / Recraft vector&lt;/td&gt;&lt;td&gt;The print-critical one: protects logos and PDF/X vector output.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Background removal&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;@imgly/background-removal&lt;/code&gt; (in-browser)&lt;/td&gt;&lt;td&gt;On-device, instant, privacy-preserving. Feeds cutout paths.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Image correction &amp;#x26; DPI&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Perfectly Clear plugin + DPI validation&lt;/td&gt;&lt;td&gt;Auto-corrects exposure, color, sharpness, and noise; validation flags undersized images. Prefer high-res native generation over upscaling.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Text generation / copy&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Claude Sonnet 4.6&lt;/strong&gt;, &lt;strong&gt;GPT-4.1 Nano&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Headlines, rewrites, per-recipient VDP copy.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Text-to-speech / sound&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;ElevenLabs (where relevant)&lt;/td&gt;&lt;td&gt;Less common in print; relevant for multi-channel campaigns.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;h2 id=&quot;the-quick-reference-picks&quot;&gt;The quick-reference picks&lt;/h2&gt;
&lt;p&gt;If you skip the rubric and just want today’s defaults, this is the snapshot we’d start from. Treat it as the hypothesis your own benchmarks confirm or overturn, not a verdict.&lt;/p&gt;


















































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Print job&lt;/th&gt;&lt;th&gt;Pick&lt;/th&gt;&lt;th&gt;Why&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Logos, icons, line art, anything that must scale&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Recraft V3&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;The only major model with native SVG/vector output. Everything else hands you pixels to trace. Vector survives into PDF/X as real paths and prints sharp at any size.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Legible text inside the image (posters, labels, packaging motifs)&lt;/td&gt;&lt;td&gt;&lt;strong&gt;GPT Image 1.5&lt;/strong&gt;, &lt;strong&gt;FLUX.2 [pro]&lt;/strong&gt;, &lt;strong&gt;Nano Banana Pro&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;These three rendered every required string in our suite exactly. The typography specialists did not: Ideogram ranks twelfth of fifteen on the measured column. Still, bake text in only when you must: real type as a separate editor layer beats all of them.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Photoreal hero imagery at print size&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Seedream 4.5&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Resolution is the print bottleneck, and it is the one criterion where the field genuinely splits. At default parameters Seedream returns 2048px (about 6.8” at 300 DPI); most of the field returns 1024px, which is 3.4” and a bet on the upscaler.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Transparent cut-outs for stickers, die-cuts, merch&lt;/td&gt;&lt;td&gt;&lt;strong&gt;GPT Image 1.5&lt;/strong&gt;, and a background-removal step regardless&lt;/td&gt;&lt;td&gt;It is the only model in our run that reliably emits a real alpha channel. Thirteen of fifteen emit none at all, so the cut-out has to be a pipeline step, not a prompt.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;”Change X, keep everything else” edits&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Flux Kontext Pro / Max&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Best-in-class instruction-following edits with subject preservation, which is critical when the asset is the customer’s product photo and “the AI improved it” means a reprint. The Nano Banana family is the consistency pick for variants of one subject.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Interactive drafts (speed and cost)&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Recraft 20B&lt;/strong&gt;, &lt;strong&gt;Gemini Flash Edit&lt;/strong&gt;, base &lt;strong&gt;Nano Banana&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;In an editor, a 20-second generation is a bounce. Draft cheap and fast, then re-render the final with the heavyweight. The customer never needs to know two models were involved.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Data sovereignty / self-hosting&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Qwen-Image / Qwen Image Edit&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Apache 2.0 open weights. For European operators and government buyers who won’t send customer uploads to a US API, “we can run it ourselves” beats a quality delta.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;VDP copy and headlines&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Claude Sonnet 4.6&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Cheap enough per recipient at 100k-mailer volume, and strong at length-constrained rewriting: the “fit this text frame” problem.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Three of those rationales carry more weight than the rest.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vector is the criterion that decides printability.&lt;/strong&gt; Screen products never need vector, so general models never optimized for it, and Recraft sits almost alone on the one measure that determines whether a logo prints cleanly at size. If you adopt a single routing rule from this guide, make it this one.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Resolution and color are pipeline problems the model can only start.&lt;/strong&gt; No model outputs CMYK; every one generates sRGB. “CMYK-readiness” is really about how hard a model’s palette clips on conversion, and your ICC and soft-proofing pipeline owns the rest. Same with DPI: pair every generation path with validation and a dedicated upscaler rather than trusting native output. You can swap the model whenever you want; the print-correctness layer around it has to stay put.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Transparency is not a model feature you can shop for.&lt;/strong&gt; When we asked every model for a transparent background and measured the alpha channel that came back, thirteen of fifteen returned none, and a few painted a checkerboard into the pixels instead. If your product sells stickers, die-cuts, or anything composited onto a garment, budget for background removal and edge cleanup as a permanent pipeline stage. Choosing a different model does not remove that stage.&lt;/p&gt;
&lt;p&gt;One roster note: &lt;strong&gt;Adobe Firefly&lt;/strong&gt; isn’t wired into CE.SDK’s plugin roster today, but for licensing-sensitive enterprises its trained-on-licensed-content and indemnification story is the strongest answer to criterion 11. Include it in your own evaluation if criterion 11 dominates your weighting.&lt;/p&gt;
&lt;h2 id=&quot;a-worked-decision-which-model-wins-for-your-product&quot;&gt;A worked decision: which model wins for &lt;em&gt;your&lt;/em&gt; product?&lt;/h2&gt;
&lt;p&gt;The rubric resolves to a different answer per context. Three quick reads:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://img.ly/products/creative-sdk/&quot;&gt;Interactive design tool&lt;/a&gt; (non-designers personalizing templates).&lt;/strong&gt; Weight responsiveness, controllability, and vector. Likely stack: &lt;strong&gt;Recraft 20B / Nano Banana&lt;/strong&gt; for fast generation, &lt;strong&gt;Recraft V3&lt;/strong&gt; for final vector, &lt;strong&gt;Gemini Flash Edit&lt;/strong&gt; for interactive edits. The customer iterates fast and the final asset is print-clean. See the &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/in-app-ugc/&quot;&gt;in-app UGC use case&lt;/a&gt; for how the benchmark weights this context.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://img.ly/use-cases/variable-data-printing/&quot;&gt;High-volume VDP / automation&lt;/a&gt; (100k personalized mailers).&lt;/strong&gt; Weight cost, throughput, consistency, and licensing. Likely stack: a cost-efficient generation model at high concurrency, &lt;strong&gt;Claude Sonnet 4.6&lt;/strong&gt; for per-recipient copy, seed-locked for repeatability, all run headless. Per-unit cost dominates the decision. See the &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/personalization-at-scale/&quot;&gt;personalization-at-scale use case&lt;/a&gt; for the weighted read.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://img.ly/use-cases/ai-editor/print-on-demand/&quot;&gt;Merch / POD&lt;/a&gt; (user art on physical products).&lt;/strong&gt; Weight transparency/cutout, resolution, and quality. Likely stack: &lt;strong&gt;background removal&lt;/strong&gt; plus cutout for subjects, &lt;strong&gt;upscaling&lt;/strong&gt; for resolution, &lt;strong&gt;Seedream V4 / Nano Banana Pro&lt;/strong&gt; for generated designs, with hard licensing checks because the output is resold. See the &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/merch-and-stickers/&quot;&gt;merch and stickers use case&lt;/a&gt;, where the transparency gap decides the stack.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In every case, the conclusion is the same: &lt;strong&gt;no one model wins all three.&lt;/strong&gt; The advantage is in being able to route each task to the model that wins it, and to swap models as better ones appear, without re-architecting your editor.&lt;/p&gt;
&lt;h2 id=&quot;what-makes-per-job-routing-practical-native-plugins--the-ai-gateway&quot;&gt;What makes per-job routing practical: native plugins + the AI Gateway&lt;/h2&gt;
&lt;p&gt;A criteria-driven, route-per-job strategy only works if switching models is cheap. If wiring in a new model means a new API integration, new auth, new key management, new error handling, and new billing plumbing every time, you’ll quietly standardize on whatever you integrated first, and inherit its weaknesses on every job it loses.&lt;/p&gt;
&lt;p&gt;CE.SDK is built to avoid that:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Every image model in this guide is natively supported as an AI plugin.&lt;/strong&gt; Text-to-image and image editing run as drop-in providers inside the editor, not bespoke integrations you build and maintain. Background removal, vectorization, and image correction (&lt;a href=&quot;https://img.ly/demos/perfectlyclear-plugin/web/&quot;&gt;Perfectly Clear&lt;/a&gt;) ship as their own plugins, and text generation routes through built-in Anthropic and OpenAI providers. Adopting Recraft for vector, Ideogram for text, and Flux Kontext for edits is &lt;a href=&quot;https://img.ly/demos/ai-editor/&quot;&gt;configuration, not three separate engineering projects&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The &lt;a href=&quot;https://img.ly/ai-gateway/&quot;&gt;AI Gateway&lt;/a&gt; gives you one unified surface to every model.&lt;/strong&gt; Instead of stitching together a dozen provider APIs, key schemes, and rate-limit behaviors, you call models through a single managed layer. That’s what turns “route each task to the best model” from an architecture diagram into a one-line config change, and what lets you swap a model the day a better one ships, without touching your editor or your print pipeline.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Outputs land inside the print pipeline, not beside it.&lt;/strong&gt; Whatever model produces the asset, it flows into locked templates with bleed, safe-area, and DPI constraints, then through CMYK / spot-color / PDF/X-3 export. The model is interchangeable; print correctness is not. Print-on-demand company &lt;a href=&quot;https://img.ly/case-studies/print-bar/&quot;&gt;The Print Bar&lt;/a&gt; and direct-mail platform &lt;a href=&quot;https://img.ly/case-studies/postbuddy/&quot;&gt;PostBuddy&lt;/a&gt; both built on this pipeline, embedding CE.SDK rather than maintaining their own print editor.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;So the “best model” question is no longer a one-time, high-stakes bet on a single provider; it’s an ongoing optimization. The benchmark scores below tell you &lt;em&gt;which&lt;/em&gt; model to route each job to today, and the native-plugin and Gateway architecture is what lets you act on that answer and revisit it when the scores change next quarter.&lt;/p&gt;
&lt;h2 id=&quot;the-imgly-genai-benchmark-measured-results&quot;&gt;The IMG.LY GenAI benchmark: measured results&lt;/h2&gt;
&lt;p&gt;IMG.LY’s GenAI benchmark now scores these models on print-representative prompts, and the results are live and explorable at &lt;a href=&quot;https://img.ly/ai-benchmarks/&quot;&gt;img.ly/ai-benchmarks&lt;/a&gt;. The table below is the top-line read for the text-to-image models in this guide: the overall score out of 5, the four print-critical criteria (each out of 5), list price per image, and median (p50) latency. All numbers are from the pilot-0 dataset (July 2026); scores are sorted by overall.&lt;/p&gt;







































































































































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Overall&lt;/th&gt;&lt;th&gt;Text&lt;/th&gt;&lt;th&gt;Color&lt;/th&gt;&lt;th&gt;Transparency&lt;/th&gt;&lt;th&gt;Resolution&lt;/th&gt;&lt;th&gt;$ / image&lt;/th&gt;&lt;th&gt;p50&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/nano-banana-2-lite/&quot;&gt;Nano Banana 2 Lite&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.87 / 5&lt;/td&gt;&lt;td&gt;4.9&lt;/td&gt;&lt;td&gt;2.9&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;3.0&lt;/td&gt;&lt;td&gt;$0.02&lt;/td&gt;&lt;td&gt;4.0s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/seedream-4-5/&quot;&gt;Seedream 4.5&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.80 / 5&lt;/td&gt;&lt;td&gt;4.9&lt;/td&gt;&lt;td&gt;2.9&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;5.0&lt;/td&gt;&lt;td&gt;$0.048&lt;/td&gt;&lt;td&gt;12.4s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/flux-2/&quot;&gt;FLUX.2&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.64 / 5&lt;/td&gt;&lt;td&gt;4.9&lt;/td&gt;&lt;td&gt;2.8&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;2.0&lt;/td&gt;&lt;td&gt;$0.013&lt;/td&gt;&lt;td&gt;2.0s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/flux-2-turbo/&quot;&gt;FLUX.2 Turbo&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.61 / 5&lt;/td&gt;&lt;td&gt;4.9&lt;/td&gt;&lt;td&gt;2.7&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;2.0&lt;/td&gt;&lt;td&gt;$0.015&lt;/td&gt;&lt;td&gt;2.0s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/gemini-25-flash-image/&quot;&gt;Gemini 2.5 Flash Image&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.60 / 5&lt;/td&gt;&lt;td&gt;5.0&lt;/td&gt;&lt;td&gt;1.7&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;3.0&lt;/td&gt;&lt;td&gt;$0.039&lt;/td&gt;&lt;td&gt;7.4s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/recraft-v3/&quot;&gt;Recraft V3&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.29 / 5&lt;/td&gt;&lt;td&gt;4.0&lt;/td&gt;&lt;td&gt;1.9&lt;/td&gt;&lt;td&gt;2.0&lt;/td&gt;&lt;td&gt;3.0&lt;/td&gt;&lt;td&gt;$0.04&lt;/td&gt;&lt;td&gt;7.4s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/nano-banana-2/&quot;&gt;Nano Banana 2&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.24 / 5&lt;/td&gt;&lt;td&gt;5.0&lt;/td&gt;&lt;td&gt;3.2&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;3.0&lt;/td&gt;&lt;td&gt;$0.10&lt;/td&gt;&lt;td&gt;13.2s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/flux-2-pro/&quot;&gt;FLUX.2 Pro&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.23 / 5&lt;/td&gt;&lt;td&gt;5.0&lt;/td&gt;&lt;td&gt;2.3&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;2.0&lt;/td&gt;&lt;td&gt;$0.04&lt;/td&gt;&lt;td&gt;11.5s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/qwen-image/&quot;&gt;Qwen-Image&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.21 / 5&lt;/td&gt;&lt;td&gt;4.8&lt;/td&gt;&lt;td&gt;2.0&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;2.0&lt;/td&gt;&lt;td&gt;$0.03&lt;/td&gt;&lt;td&gt;6.9s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/gpt-image-1-5/&quot;&gt;GPT Image 1.5&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.07 / 5&lt;/td&gt;&lt;td&gt;5.0&lt;/td&gt;&lt;td&gt;3.4&lt;/td&gt;&lt;td&gt;4.2&lt;/td&gt;&lt;td&gt;3.0&lt;/td&gt;&lt;td&gt;$0.12&lt;/td&gt;&lt;td&gt;34.0s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/ideogram-v3/&quot;&gt;Ideogram 3.0&lt;/a&gt;&lt;/td&gt;&lt;td&gt;2.97 / 5&lt;/td&gt;&lt;td&gt;4.7&lt;/td&gt;&lt;td&gt;2.8&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;3.0&lt;/td&gt;&lt;td&gt;$0.06&lt;/td&gt;&lt;td&gt;17.9s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/nano-banana-pro/&quot;&gt;Nano Banana Pro&lt;/a&gt;&lt;/td&gt;&lt;td&gt;2.55 / 5&lt;/td&gt;&lt;td&gt;5.0&lt;/td&gt;&lt;td&gt;3.7&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;3.0&lt;/td&gt;&lt;td&gt;$0.36&lt;/td&gt;&lt;td&gt;23.3s&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Measured on the pilot-0 suite (July 2026): 37 prompts times 3 seeds per model, every image scored 0 to 5 per criterion by an automated vision judge. Two models carry newer versions than the guide names them by: Seedream V4 is measured as Seedream 4.5, and GPT Image 1 as GPT Image 1.5. The suite also covers Seedream 5.0 Lite, Luma Photon, and Stable Diffusion 1.5, which this guide’s roster does not name; their scores are on the &lt;a href=&quot;https://img.ly/ai-benchmarks/models/&quot;&gt;rankings page&lt;/a&gt;. The overall column is an unweighted blend, so read it alongside the per-use-case weighting, not as a single leaderboard: the cheap fast model tops the raw average while the priciest lands last. Full methodology is at &lt;a href=&quot;https://img.ly/ai-benchmarks/methodology/&quot;&gt;img.ly/ai-benchmarks/methodology&lt;/a&gt;, and every model, prompt, and generated image is explorable at &lt;a href=&quot;https://img.ly/ai-benchmarks/&quot;&gt;img.ly/ai-benchmarks&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;A few things the numbers make plain, and that the criteria above predicted. Transparency is close to unsolved: 13 of 15 tested models emit no real alpha channel at all (&lt;a href=&quot;https://img.ly/ai-benchmarks/findings/transparency/&quot;&gt;the transparency finding&lt;/a&gt;), and GPT Image 1.5 is the only one that scores meaningfully on it. Brand color is broadly weak: measured with CIEDE2000, no model reliably hits an exact hex (&lt;a href=&quot;https://img.ly/ai-benchmarks/findings/brand-color/&quot;&gt;the brand-color finding&lt;/a&gt;). Text is the one column where the field mostly clears the bar: on our per-required-string measure several general-purpose flagships render every string in the suite exactly, while the model with the biggest reputation for text, Ideogram, actually lands near the bottom of the column, not the top (&lt;a href=&quot;https://img.ly/ai-benchmarks/findings/text-reliability/&quot;&gt;the text-reliability finding&lt;/a&gt;). The gaps open exactly where print work gets hard. Averaged across the field, the neon sign in Japanese scores 3.91 and the dense packaging label 4.40, against 4.91 for a greeting-card cover. Latin headlines are close to solved; small type and non-Latin scripts are not.&lt;/p&gt;
&lt;h3 id=&quot;the-same-numbers-weighted-for-print&quot;&gt;The same numbers, weighted for print&lt;/h3&gt;
&lt;p&gt;The overall column is the wrong lens for a print product, which is the whole reason the benchmark carries use-case weights. Reweighted for print files (resolution 35%, color 30%, text 15%, transparency 10%), &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/print-files/&quot;&gt;Seedream 4.5 leads at 3.64&lt;/a&gt;, ahead of Seedream 5.0 Lite at 3.50 and GPT Image 1.5 at 3.45: native output size at default parameters does most of the work, and the two models that return 2048px start the job at twice the printable width of the 1024px field.&lt;/p&gt;
&lt;p&gt;Reweight again for &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/merch-and-stickers/&quot;&gt;merch and stickers&lt;/a&gt;, where alpha carries 35%, and the order breaks completely: GPT Image 1.5 wins at 3.72, nearly a full point ahead of Recraft V3 at 2.78, because it is the only model that reliably returns a real alpha channel. Same measurements, different job, different answer. That inversion is the argument for routing per job rather than standardizing on one provider.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The gated report is coming to subscribers first.&lt;/strong&gt; &lt;em&gt;AI Image Models for Print Production: The Benchmark Report&lt;/em&gt; packages the weighted per-use-case totals, the full test set, and the print-specific analysis into one PDF. &lt;a href=&quot;https://share.hsforms.com/1IgAOV1wASXGPnFG4ZPLejg1hk3i&quot;&gt;Subscribe to get it before it is public&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The criteria and the routing logic above stand on their own. The scored numbers are what the suite adds: a measured, repeatable answer for each cell, refreshed as models change. For why the benchmark is built this way, and what else the run turned up, see &lt;a href=&quot;https://img.ly/blog/introducing-imgly-ai-benchmarks/&quot;&gt;&lt;strong&gt;Introducing IMG.LY AI Benchmarks&lt;/strong&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;bottom-line&quot;&gt;Bottom line&lt;/h2&gt;
&lt;p&gt;Pick criteria before you pick models. For &lt;a href=&quot;https://img.ly/use-cases/web-to-print-design-tool/&quot;&gt;web-to-print&lt;/a&gt;, the criteria that screen benchmarks ignore (vector output, print resolution, CMYK behavior, text legibility, licensing) are exactly the ones that decide whether a beautiful generation is a sellable print. Score the current models against those criteria for &lt;em&gt;your&lt;/em&gt; context, route each job to the model that wins it, and run it all through CE.SDK’s &lt;a href=&quot;https://img.ly/use-cases/ai-editor/&quot;&gt;native AI plugins&lt;/a&gt; and &lt;a href=&quot;https://img.ly/ai-gateway/&quot;&gt;AI Gateway&lt;/a&gt; so re-routing tomorrow, after the next leaderboard reshuffle, is a config change rather than a rebuild.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Companion piece: &lt;a href=&quot;https://img.ly/blog/how-to-leverage-generative-ai-in-web-to-print/&quot;&gt;&lt;strong&gt;How to Leverage Generative AI in Web-to-Print&lt;/strong&gt;&lt;/a&gt; covers the use cases, integration patterns, and print-specific pitfalls behind these model choices.&lt;/em&gt;&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="/_astro/cover.ZLW7wLRy.webp" medium="image"/><category>AI</category><category>Image Gen</category><category>Web-to-Print</category><category>Print</category></item><item><title>How to Leverage Generative AI in Web-to-Print</title><link>https://img.ly/blog/how-to-leverage-generative-ai-in-web-to-print/</link><guid isPermaLink="true">https://img.ly/blog/how-to-leverage-generative-ai-in-web-to-print/</guid><description>Generative AI changes the economics of web-to-print by doing the part customers were never qualified to do: generating images, fixing resolution, removing backgrounds, vectorizing logos. Here are the use cases worth building, and the print-specific realities to design around.</description><pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Web-to-print sells self-service: your customer designs a postcard, a flyer, a t-shirt, a packaging label in the browser and orders it without a designer in the loop. For years, though, the self-service stopped at the hard part. Someone still had to supply a usable image. Someone still had to know what 300 DPI meant. When they didn’t, the file landed on a prepress desk to be fixed by hand.&lt;/p&gt;
&lt;p&gt;One operator we spoke to put it bluntly: his designers spend roughly 40% of their time fixing customer artwork rather than creating templates. Tom Rowe of &lt;a href=&quot;https://img.ly/case-studies/print-bar/&quot;&gt;The Print Bar&lt;/a&gt; described what that looks like at the design step: &lt;em&gt;“Our brand didn’t match our experience. You could see it in our bounce rates.”&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Generative AI changes the economics of that hard part. It can turn a prompt into a usable asset, lift a 72-DPI logo to something printable, and strip a background in the browser before a human ever sees the file, removing the specific friction that kept self-service from being self-service in the first place.&lt;/p&gt;
&lt;p&gt;It can also go wrong: aimed carelessly, the same models produce beautiful screen images that fall apart on press. This post is about the difference between the two: the use cases worth building, and the print-specific realities you have to design around.&lt;/p&gt;
&lt;h2 id=&quot;the-shift-from-upload-a-print-ready-file-to-describe-what-you-want&quot;&gt;The shift: from “upload a print-ready file” to “describe what you want”&lt;/h2&gt;
&lt;p&gt;The old &lt;a href=&quot;https://img.ly/use-cases/web-to-print-design-tool/&quot;&gt;web-to-print&lt;/a&gt; contract asked the customer to arrive with a finished, technically correct asset. That filter excluded most of the market: the CRM manager, the franchise owner, the small-business buyer who &lt;em&gt;“knows Canva because they make their wedding invitations on it”&lt;/em&gt; but has never opened InDesign and never will.&lt;/p&gt;
&lt;p&gt;Generative AI moves the burden of production off the customer. Instead of &lt;em&gt;“upload a 300 DPI CMYK image with 3mm bleed,”&lt;/em&gt; the ask becomes &lt;em&gt;“type what you want on the card.”&lt;/em&gt; The editor, not the customer and not your prepress team, becomes responsible for turning intent into a production-safe file.&lt;/p&gt;
&lt;p&gt;That principle runs through everything below: &lt;strong&gt;AI is most valuable in web-to-print where it removes a step the customer was never qualified to do.&lt;/strong&gt; The closer a use case sits to it, the more it returns.&lt;/p&gt;
&lt;h2 id=&quot;the-use-cases-worth-building&quot;&gt;The use cases worth building&lt;/h2&gt;
&lt;h3 id=&quot;1-generate-on-brand-imagery-from-a-prompt-text-to-image&quot;&gt;1. Generate on-brand imagery from a prompt (text-to-image)&lt;/h3&gt;
&lt;p&gt;This is the most obvious win. A customer needs a hero image, a background, a seasonal motif, or a product scene, and doesn’t have one. &lt;a href=&quot;https://img.ly/demos/ai-editor/&quot;&gt;Text-to-image generation&lt;/a&gt; lets them describe it and get options in seconds, without a stock-photo license hunt or a design request ticket.&lt;/p&gt;
&lt;p&gt;In a print context the point is &lt;strong&gt;unblocking the order&lt;/strong&gt;, not novelty for its own sake. The customer who would have abandoned at “I don’t have a good image” now keeps going. Models like Recraft V3, Seedream V4, Ideogram V3, and the Nano Banana family each have different strengths here (covered in depth in the companion guide, &lt;em&gt;The Best GenAI Models for Web-to-Print&lt;/em&gt;), but the integration pattern is the same: generation feeds a placeholder inside a locked template, so the output lands inside your bleed, safe-area, and brand constraints automatically.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; generated images are usually delivered at screen resolution, often around 1024px. That’s fine for a business card photo, marginal for an A4 flyer, and unusable for large-format. Pair generation with upscaling (below) and size validation before you let it reach export. Native output sizes vary widely by model: the &lt;a href=&quot;https://img.ly/ai-benchmarks/models/&quot;&gt;measured native output in the model rankings&lt;/a&gt; (pilot-0 dataset, July 2026) shows some models topping out near 1024px while others generate at 2K.&lt;/p&gt;
&lt;h3 id=&quot;2-adapt-and-extend-customer-supplied-images-image-to-image-generative-fillexpand&quot;&gt;2. Adapt and extend customer-supplied images (image-to-image, generative fill/expand)&lt;/h3&gt;
&lt;p&gt;Customers rarely arrive with an asset that fits your canvas. It’s the wrong aspect ratio, it has the wrong background, or it’s a portrait crop where you need a landscape banner. Image-to-image editing and &lt;strong&gt;generative expand&lt;/strong&gt; (outpainting) solve the most common case: extending an image to fill bleed or a different SKU’s dimensions instead of stretching or letterboxing it.&lt;/p&gt;
&lt;p&gt;This is quietly one of the highest-value print use cases, because aspect-ratio and bleed mismatches are a top source of prepress rework. Generative fill also handles object removal (“take the coffee cup off the desk”), background swaps, and clean-up: the edits a non-designer can’t do in any tool they own.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; outpainting invents pixels. On a brand asset or a product photo, “invented” can mean “wrong.” Keep generative expand to backgrounds and ambient areas. Never let it reconstruct a logo, a face, or a product the customer is actually selling.&lt;/p&gt;
&lt;h3 id=&quot;3-background-removal-and-cutouts-for-mockups-and-merch&quot;&gt;3. Background removal and cutouts for mockups and merch&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/demos/background-removal/&quot;&gt;Background removal&lt;/a&gt; is the workhorse, and it has to be a pipeline step because the models won’t do it for you: our &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/transparency/&quot;&gt;transparency finding&lt;/a&gt; shows 13 of 15 tested models emit no real alpha channel at all. For merch and print-on-demand it’s the difference between a customer’s snapshot and a clean subject that drops onto a t-shirt, mug, or sticker. IMG.LY runs this &lt;strong&gt;in the browser&lt;/strong&gt; (via the open-source &lt;code&gt;@imgly/background-removal&lt;/code&gt; package), which matters for two reasons: it’s instant enough for an interactive editor, and the customer’s image never has to leave the device.&lt;/p&gt;
&lt;p&gt;Paired with the &lt;strong&gt;&lt;a href=&quot;https://img.ly/demos/cutout-lines/&quot;&gt;cutout plugin&lt;/a&gt;&lt;/strong&gt;, background removal also feeds the literal cut lines for die-cut stickers, labels, and packaging, turning a removed background into a production path.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; browser-based removal is excellent on clear subjects and struggles on hair, glass, and fine fringes. For a print product that will be inspected up close, expose a manual refine step rather than trusting the mask blindly.&lt;/p&gt;
&lt;h3 id=&quot;4-vectorize-raster-art-into-print-scalable-graphics&quot;&gt;4. Vectorize raster art into print-scalable graphics&lt;/h3&gt;
&lt;p&gt;This is the use case print people care about most and screen-first builders forget. A customer uploads a logo as a small, jagged PNG. Printed at size, it’s mush. &lt;strong&gt;&lt;a href=&quot;https://img.ly/demos/vectorizer-plugin/&quot;&gt;Vectorization&lt;/a&gt;&lt;/strong&gt; (AI raster-to-SVG) traces it into clean, resolution-independent paths that scale to any output size and reproduce crisply on press.&lt;/p&gt;
&lt;p&gt;Vectorize is also how you get logos and simple graphics into a state your &lt;strong&gt;&lt;a href=&quot;https://img.ly/demos/export-print-ready-pdf/&quot;&gt;PDF/X pipeline&lt;/a&gt;&lt;/strong&gt; can preserve as vector, instead of rasterizing everything and throwing away the scalability you vectorized for. Recraft V3 is notable here because it can generate &lt;strong&gt;natively vector&lt;/strong&gt; output rather than only tracing after the fact.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; vectorization is great for logos, icons, and flat art. It’s the wrong tool for a photograph. Detect the asset type and route photos to upscaling, not tracing.&lt;/p&gt;
&lt;h3 id=&quot;5-resolution-and-image-quality-upscaling-correction-and-dpi-repair&quot;&gt;5. Resolution and image quality: upscaling, correction, and DPI repair&lt;/h3&gt;
&lt;p&gt;The single most common web-to-print failure is a low-resolution image: the 72-DPI photo that looks fine on screen and prints as a blurry mess. Two different AI operations fix it, and it pays to keep them straight, because they solve different problems.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Upscaling (super-resolution)&lt;/strong&gt; raises the actual pixel count, reconstructing detail to lift a small image toward a printable size. It’s a generative step you route to a dedicated model, the same way you route a text-to-image prompt; the strong upscalers can take a 1024px generation to 4K. Reach for it when the source is simply too small for its placed size.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Image correction&lt;/strong&gt; is the other half, and it ships in CE.SDK today as the &lt;a href=&quot;https://img.ly/demos/perfectlyclear-plugin/web/&quot;&gt;Perfectly Clear plugin&lt;/a&gt;. It auto-corrects exposure, contrast, color, tint, sharpness, and noise in a single pass. It won’t add pixels, but it gets the most printable result out of the pixels you have, which is what a dim, soft phone photo usually needs more than raw resolution.&lt;/p&gt;
&lt;p&gt;Tie both to your &lt;strong&gt;DPI validation&lt;/strong&gt;. When the editor detects an image below your minimum threshold for its placed size, offer to correct or upscale it instead of just blocking export. You convert a dead end into a one-click fix, the kind of &lt;em&gt;“they don’t need to know what DPI is”&lt;/em&gt; experience operators are chasing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; upscaling fabricates detail. It rescues marginal images; it can’t conjure a sharp 4-megapixel product shot from a thumbnail, and correction can’t add resolution at all. Set honest thresholds, and still warn when the source is hopeless.&lt;/p&gt;
&lt;h3 id=&quot;6-copy-and-variable-text-generation-text--vdp-at-scale&quot;&gt;6. Copy and variable text generation (text + VDP at scale)&lt;/h3&gt;
&lt;p&gt;Generative text earns its place in two spots. First, as a writing aid inside the editor: headline options, a tagline, a “make it shorter / more formal / fit this space” rewrite for the non-writer staring at an empty text box. Second, and more powerfully for print, in &lt;strong&gt;&lt;a href=&quot;https://img.ly/use-cases/variable-data-printing/&quot;&gt;Variable Data Printing&lt;/a&gt;&lt;/strong&gt;: generating or localizing per-recipient copy across thousands of personalized postcards, mailers, or labels from a single template.&lt;/p&gt;
&lt;p&gt;VDP is the highest-leverage place to use generative text: generate the variants, merge them into print-ready files in headless mode, and run the batch.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; generated copy needs guardrails. Length limits so it doesn’t overflow the text frame, tone and brand constraints, and a human-review gate for anything legally sensitive: pricing, claims, regulated industries.&lt;/p&gt;
&lt;h3 id=&quot;7-template-adaptation-and-auto-resize-across-skus&quot;&gt;7. Template adaptation and auto-resize across SKUs&lt;/h3&gt;
&lt;p&gt;A print catalog is the same design across many sizes and substrates: the same campaign as a postcard, a flyer, a poster, a social tile. &lt;a href=&quot;https://img.ly/demos/automated-resizing/&quot;&gt;AI-assisted resize and re-layout&lt;/a&gt; adapt one master design to every SKU’s dimensions, reflowing content intelligently instead of scaling it blindly. Combined with locked templates, you expand SKUs without re-authoring each one.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; reflow decisions still need brand rules. Lock what must stay fixed (logo size, safe area, mandatory legal text) so the AI rearranges within your constraints, not over them.&lt;/p&gt;
&lt;h3 id=&quot;8-guided-prepress-correction&quot;&gt;8. Guided prepress correction&lt;/h3&gt;
&lt;p&gt;Instead of rejecting a bad file after the fact, use AI to catch and &lt;em&gt;fix&lt;/em&gt; problems inside the editor: flag the low-res image and offer to upscale it, &lt;a href=&quot;https://img.ly/demos/design-validation/&quot;&gt;detect text outside the safe area&lt;/a&gt; and nudge it in, notice a near-white “white” that won’t print and correct it. The same problems your prepress team used to fix by hand get caught here, before the order is placed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; auto-correction must be transparent and reversible. Show the customer what changed and let them undo it. Silent “helpful” edits to someone’s artwork erode trust fast.&lt;/p&gt;
&lt;h3 id=&quot;9-localization-and-market-variants&quot;&gt;9. Localization and market variants&lt;/h3&gt;
&lt;p&gt;For franchise networks and multi-market brands, &lt;a href=&quot;https://img.ly/demos/language/&quot;&gt;AI translation&lt;/a&gt; plus regeneration of localized imagery and copy turns one approved master into 12 market versions, each print-ready and on-brand, without a translation agency or 12 design tickets. A 200-location franchise can ship 200 localized flyers with the same brand guideline enforced by the editor on every one.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; machine translation in regulated or legal copy needs human sign-off, and text expansion (German runs roughly 30% longer than English) will break tight layouts unless your template frames can flex.&lt;/p&gt;
&lt;h3 id=&quot;10-realistic-product-previews&quot;&gt;10. Realistic product previews&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/demos/mockup-editor/&quot;&gt;Generative and 3D-assisted mockups&lt;/a&gt; show the design &lt;strong&gt;on the actual product&lt;/strong&gt;: the shirt, the mug, the folded brochure, the box, not an abstract artboard. Operators consistently tie this to higher conversion and fewer “will this look right?” support tickets, because the preview matches what arrives on the doorstep.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; a preview that’s prettier than the print sets up a disappointed customer and a reprint. Calibrate mockups to real output, including substrate color and finish.&lt;/p&gt;
&lt;h2 id=&quot;the-print-specific-realities&quot;&gt;The print-specific realities&lt;/h2&gt;
&lt;p&gt;Most generative models are built for screens. Print adds requirements the web doesn’t have, and the ones below are where projects most often break.&lt;/p&gt;
&lt;h3 id=&quot;resolution-and-dpi&quot;&gt;Resolution and DPI&lt;/h3&gt;
&lt;p&gt;Generated images typically arrive at around 1024px. At 300 DPI that’s about a 3.4-inch image: fine for a business card, not for a poster. &lt;strong&gt;Always validate output resolution against placed size&lt;/strong&gt;, and route undersized assets through upscaling before export. Treat raw generation resolution as a starting point, not a deliverable. Native output size is one of the axes we measure: see the &lt;a href=&quot;https://img.ly/ai-benchmarks/models/&quot;&gt;measured native output per model&lt;/a&gt; to know which models start closer to print size and which need the most upscaling.&lt;/p&gt;
&lt;h3 id=&quot;color-srgb-in-cmyk-out&quot;&gt;Color: sRGB in, CMYK out&lt;/h3&gt;
&lt;p&gt;Models generate in RGB. Print is CMYK, plus spot colors. The vivid blues, greens, and oranges that look great on screen sit outside the CMYK gamut and will shift on press. And even inside RGB, hitting an exact brand hex is hard: our &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/brand-color/&quot;&gt;brand-color finding&lt;/a&gt; measures color accuracy with CIEDE2000 and no model reliably lands the color in your brand book, which is one more reason correction belongs on the canvas. Your pipeline, not the model, owns this: convert with the right ICC profile, validate against the print provider’s color requirements, and preserve &lt;strong&gt;named spot colors&lt;/strong&gt; through to the PDF/X. Show the customer a soft-proof so the shift isn’t a surprise on delivery.&lt;/p&gt;
&lt;h3 id=&quot;vector-vs-raster-and-the-everything-got-rastered-trap&quot;&gt;Vector vs. raster (and the “everything got rastered” trap)&lt;/h3&gt;
&lt;p&gt;Most models output raster. For logos, type, and line art that need to scale and print crisply, raster is the wrong format, and a PDF that rasterizes all text and vectors is, in one operator’s words, &lt;em&gt;“death for the printer.”&lt;/em&gt; Use vector-native generation (e.g. Recraft) and vectorization for the elements that need it, and make sure your export keeps &lt;strong&gt;vector and text as vector&lt;/strong&gt; with CMYK values preserved (PDF/X-3), not flattened to pixels.&lt;/p&gt;
&lt;h3 id=&quot;text-rendering-inside-images&quot;&gt;Text rendering inside images&lt;/h3&gt;
&lt;p&gt;Generative models have a reputation for garbling text inside images, and our &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/text-reliability/&quot;&gt;text-reliability finding&lt;/a&gt; shows the top of the field has largely fixed it: a handful of flagships now render every required string in our suite exactly. The catch for print is that reputation and measurement no longer line up. Ideogram, the model best known for typography, ranks twelfth of fifteen on the measured column; several general-purpose flagships beat it. Reach for a specialist on reputation and you can end up with worse type than the default model you already route to. Put every model on the &lt;a href=&quot;https://img.ly/ai-benchmarks/prompts/t04-wordmark/&quot;&gt;same wordmark prompt&lt;/a&gt; to see the spread. For anything that must be readable in print, &lt;strong&gt;don’t bake text into a generated image.&lt;/strong&gt; Generate the imagery, then set real, editable, vectorizable type as a separate layer in the editor. If you must bake text in, pick the model from the measured column rather than from the marketing; the companion guide has the current ranking.&lt;/p&gt;
&lt;h3 id=&quot;brand-safety-and-guardrails&quot;&gt;Brand safety and guardrails&lt;/h3&gt;
&lt;p&gt;Let customers generate anything and you lose brand control. Box generation inside locked templates instead: it fills a defined placeholder, within fixed margins, alongside a logo and palette the customer can’t move. A franchisee can change the headline and the photo; the logo size, position, and safe area stay locked. AI expands what the customer can create without expanding what they can break.&lt;/p&gt;
&lt;h3 id=&quot;commercial-and-licensing-rights&quot;&gt;Commercial and licensing rights&lt;/h3&gt;
&lt;p&gt;A web-to-print product is sold. That makes the IP status of generated assets a question you have to answer before you ship. Does your model provider grant commercial-use rights to outputs? Are there indemnities? Does training-data provenance create exposure for a customer reselling the printed product? Pick models and providers whose commercial terms you’ve actually read, and surface usage terms where they matter.&lt;/p&gt;
&lt;h3 id=&quot;cost-and-latency-the-unit-economics&quot;&gt;Cost and latency (the unit economics)&lt;/h3&gt;
&lt;p&gt;Every generation costs money and time. In an interactive editor, a 30-second generation is a broken experience. In a high-volume VDP run, a few cents per asset multiplied by 100,000 recipients is a real line item. The spread is not small: the &lt;a href=&quot;https://img.ly/ai-benchmarks/models/&quot;&gt;measured $/image and p50 latency in the model rankings&lt;/a&gt; run from a fraction of a cent to tens of cents, and from a couple of seconds to over thirty. Match the model to the job: a fast, cheap model for interactive iteration, a higher-quality (slower, pricier) model for the final render, and use the &lt;a href=&quot;https://img.ly/ai-benchmarks/compare/&quot;&gt;compare tool&lt;/a&gt; to weigh the trade-off head-to-head. Cache aggressively. The companion guide scores models on that trade-off.&lt;/p&gt;
&lt;h3 id=&quot;data-privacy-and-sovereignty&quot;&gt;Data privacy and sovereignty&lt;/h3&gt;
&lt;p&gt;Customer-uploaded images can be sensitive: faces, IDs, confidential product designs. Browser-based processing (background removal, some editing) keeps data on-device. API-based generation sends it to a third party, which has GDPR and data-residency implications, especially for European customers and government buyers who &lt;em&gt;“prioritize a European vendor for data sovereignty.”&lt;/em&gt; Be explicit about what runs locally vs. in the cloud, and choose providers accordingly.&lt;/p&gt;
&lt;h3 id=&quot;hallucination-and-consistency&quot;&gt;Hallucination and consistency&lt;/h3&gt;
&lt;p&gt;Generative output is non-deterministic. The same prompt yields different results, and details drift. Our &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/consistency/&quot;&gt;consistency finding&lt;/a&gt; measures the worst of it: no model holds a described character across scenes, so series work needs identity as a reusable asset, not a regeneration lottery. For brand assets, product likenesses, and anything a customer will compare against reality, constrain hard (reference images, seeds, image-to-image rather than free generation) and keep a human gate on the outputs that matter.&lt;/p&gt;
&lt;h2 id=&quot;how-imgly-approaches-it&quot;&gt;How IMG.LY approaches it&lt;/h2&gt;
&lt;p&gt;All of these constraints point the same way: &lt;strong&gt;AI is a step inside a print pipeline, not a feature bolted onto an editor.&lt;/strong&gt; A few principles shape how CE.SDK handles it:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Model-agnostic by design.&lt;/strong&gt; The &lt;a href=&quot;https://img.ly/use-cases/ai-editor/&quot;&gt;AI plugins&lt;/a&gt; connect to any third-party model or API: bring your own model. You’re not locked to one provider’s quality, price, or licensing terms. You route each task to the model that wins it, and swap models as better ones ship every few weeks.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The &lt;a href=&quot;https://img.ly/ai-gateway/&quot;&gt;AI Gateway&lt;/a&gt;&lt;/strong&gt; provides managed access to models for editors, so you’re not stitching together a dozen API integrations and key-management schemes yourself.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Generation lands inside the print pipeline.&lt;/strong&gt; Outputs flow into locked templates with bleed, safe areas, DPI thresholds, and brand constraints already enforced, then through a &lt;strong&gt;CMYK / spot-color / PDF/X-3&lt;/strong&gt; export that keeps vectors and text as vectors. The model never gets to bypass print correctness.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Drop-in generative features.&lt;/strong&gt; Background removal, generative fill, image-to-image, vectorize, and text-to-image are available as plugins, so you adopt the use cases above incrementally rather than rebuilding your editor.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Headless mode&lt;/strong&gt; runs the same generation and export server-side for batch and variable-data runs.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The result is the experience operators actually want to sell: the customer describes what they want, the AI does the part they were never qualified to do, and the file that reaches the press is print-correct because the pipeline enforced it, not because the customer got lucky. For the product-level view of this stack, see &lt;a href=&quot;https://img.ly/use-cases/ai-editor/print-on-demand/&quot;&gt;AI for Print on Demand&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;what-customers-are-doing-with-it&quot;&gt;What customers are doing with it&lt;/h2&gt;
&lt;p&gt;High-volume print and direct-mail platforms are already building on this foundation. IMG.LY powers print and &lt;a href=&quot;https://img.ly/use-cases/print-personalization/&quot;&gt;personalization workflows&lt;/a&gt; for operators including &lt;strong&gt;&lt;a href=&quot;https://img.ly/case-studies/postbuddy/&quot;&gt;Postbuddy&lt;/a&gt;&lt;/strong&gt; (personalized direct mail), &lt;strong&gt;Swiss Post&lt;/strong&gt;, &lt;strong&gt;&lt;a href=&quot;https://img.ly/case-studies/digitas/&quot;&gt;Digitas&lt;/a&gt;&lt;/strong&gt;, and &lt;strong&gt;HP&lt;/strong&gt;, alongside hundreds of smaller print, merch, and franchise platforms.&lt;/p&gt;
&lt;p&gt;The pattern that recurs in those conversations is narrow and specific: AI removed a bottleneck. The prepress team that fixed 40% of incoming files now lets the editor catch problems upstream. The customers who bounced at “I don’t have a good image” finish the order. The franchise network that needed a designer for every local variant generates them inside guardrails. It comes back to the point from the top: AI pays off where it deletes a step the customer couldn’t do and the operator didn’t want to.&lt;/p&gt;
&lt;h2 id=&quot;where-to-start&quot;&gt;Where to start&lt;/h2&gt;
&lt;p&gt;You don’t need all ten use cases on day one. The highest-ROI starting points, in order:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Background removal&lt;/strong&gt;: instant, in-browser, immediately useful for any merch or photo product.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Image correction and DPI validation&lt;/strong&gt;: catches and cleans up weak uploads before they ever reach the printer.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Generative expand / fill&lt;/strong&gt;: kills aspect-ratio and bleed rework.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Text-to-image into locked placeholders&lt;/strong&gt;: unblocks the “I don’t have an image” abandoner.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Vectorize&lt;/strong&gt;: rescues logos and protects your PDF/X output.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Each one removes a documented source of friction or rework. Add them inside locked templates and a real CMYK/PDF/X pipeline, and self-service web-to-print finally works the way it was always sold: the customer really can do it themselves.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Next: which models to actually use. See the companion guide, &lt;a href=&quot;https://img.ly/blog/best-generative-ai-models-for-web-to-print/&quot;&gt;&lt;strong&gt;The Best GenAI Models for Web-to-Print: A Buyer’s Guide&lt;/strong&gt;&lt;/a&gt;, for a criteria-based comparison across output quality, vector/SVG support, text rendering, color, responsiveness, and cost. For the measured evidence behind these pitfalls, see the &lt;a href=&quot;https://img.ly/ai-benchmarks/&quot;&gt;&lt;strong&gt;IMG.LY GenAI Benchmarks&lt;/strong&gt;&lt;/a&gt;: every model on the same prompts, scored per criterion, with the method and headline results written up in &lt;a href=&quot;https://img.ly/blog/introducing-imgly-ai-benchmarks/&quot;&gt;&lt;strong&gt;Introducing IMG.LY AI Benchmarks&lt;/strong&gt;&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="/_astro/cover.CHrckJHN.webp" medium="image"/><category>AI</category><category>Web-to-Print</category><category>Print</category><category>Creative Workflows</category></item><item><title>AI Design Agents and Creative Automation: How to Ship a Full Campaign Without a Designer</title><link>https://img.ly/blog/ai-design-agents-and-creative-automation-how-to-ship-a-full-campaign-without-a-designer/</link><guid isPermaLink="true">https://img.ly/blog/ai-design-agents-and-creative-automation-how-to-ship-a-full-campaign-without-a-designer/</guid><description>Most marketing workflows are fully automated now - except one. Design is still a handoff, a wait, a revision loop, and another wait. Here&apos;s how AI design agents are closing that gap, and how to run a full campaign production session without leaving a single conversation.</description><pubDate>Wed, 01 Apr 2026 10:14:03 GMT</pubDate><content:encoded>&lt;p&gt;You have your files, hooks, Google Ads data but still can’t actually launch a campaign because you have to go to a designer for the final assets.&lt;/p&gt;
&lt;p&gt;That gap is exactly where campaign momentum dies. Not at the strategy stage or in copy but at the last mile, when everything is ready except the thing people will actually see.&lt;/p&gt;
&lt;p&gt;Luckily, it no longer has to.&lt;/p&gt;
&lt;h2 id=&quot;the-ai-marketing-stack-has-a-design-shaped-hole-in-it&quot;&gt;The AI Marketing Stack Has a Design-Shaped Hole in It&lt;/h2&gt;
&lt;p&gt;Most marketing workflows have been quietly transformed over the last two years. Copy generation, audience segmentation, keyword research, performance analysis - all of it runs faster now, with less manual input. A solo marketer can do what used to require a team. Except for one part.&lt;/p&gt;
&lt;p&gt;Design hasn’t moved. The workflow is still: write a brief, hand it to a designer, wait, review, revise, wait again. Then do the same thing in three more formats because the 1:1 you approved doesn’t fit Stories, Display, or LinkedIn. That loop can take days. And it doesn’t matter how good your AI-generated copy is if it’s sitting in a doc waiting for someone to have bandwidth.&lt;/p&gt;
&lt;p&gt;This isn’t a resourcing problem. Hiring more designers doesn’t fix the structural issue; it just adds capacity to a fundamentally slow process. The problem is that the tools built for design weren’t designed for the workflow a modern marketer actually runs. They expect a designer at the keyboard. They don’t expect a marketer with a campaign brief and a conversation window.&lt;/p&gt;
&lt;p&gt;The result: campaigns that are otherwise fully automated still stall before they ship. The bottleneck moved from copy to creative. And most teams haven’t noticed yet, because design delays feel normal. They’ve always been there.&lt;/p&gt;
&lt;h2 id=&quot;what-an-ai-design-agent-actually-changes&quot;&gt;What an AI Design Agent Actually Changes&lt;/h2&gt;
&lt;p&gt;An AI design agent is an autonomous AI system, not a prompt-response tool. It plans, reasons, and executes design tasks independently. Given a goal, it breaks that goal into steps, uses the tools available to it, retains context across the session, and can self-correct when the output isn’t right. That’s the category: a system that drives the workflow rather than waiting for instruction at each step.&lt;/p&gt;
&lt;p&gt;Most design agents deliver speed. The handoff that used to take days can happen in minutes. But the output is still a deliverable: a file you receive, use as-is, or move somewhere else. Most operate inside existing tools or generate assets you work with elsewhere. The speed is real, the editability usually isn’t. If something needs to change, you’re back to prompting from scratch.&lt;/p&gt;
&lt;p&gt;CoDesign sits differently. It doesn’t hand you a finished asset. It gives you a working starting point on a real canvas. Layers, text boxes, placeholders - real assets you can continue to work with, in the same conversation, without leaving the session.&lt;/p&gt;
&lt;p&gt;Brand consistency gets handled at the foundation. Load your brand kit once, including colors, fonts, logo, and layout rules, and every output that session respects those constraints. You’re not eyeballing hex codes or hoping the font looks right. The rules are applied from the start, not checked at the end.&lt;/p&gt;
&lt;p&gt;Multi-format adaptation is where the time savings become concrete. A campaign that runs across Instagram, Google Display, LinkedIn, and print doesn’t produce four separate briefs and four separate rounds of designer revisions. You describe the formats you need, and the agent adapts the work. The campaign stays consistent across all of them.&lt;/p&gt;
&lt;p&gt;The biggest shift isn’t speed, though that’s real. It’s that you stay in creative control throughout. There’s no handoff moment where you lose the thread and have to re-explain the brief to someone else. The context lives in the conversation.&lt;/p&gt;
&lt;h2 id=&quot;how-to-run-a-campaign-production-session-with-codesign&quot;&gt;How to Run a Campaign Production Session with CoDesign&lt;/h2&gt;
&lt;p&gt;This is the actual sequence. Walk through it once and the workflow becomes repeatable.&lt;/p&gt;
&lt;p&gt;1.&lt;strong&gt;Start with the brief.&lt;/strong&gt; Open CoDesign and describe the campaign in plain language. Audience, objective, platform, tone, any constraints. The more specific you are here, the less iteration you’ll need later. Treat it like briefing a senior designer who hasn’t worked with your brand before.&lt;/p&gt;
&lt;p&gt;2.&lt;strong&gt;Generate copy variations before you open the canvas.&lt;/strong&gt; Use whichever AI writing tool you already work with (ChatGPT, Claude, or whatever is in your stack) to produce multiple copy directions: headline hooks, body copy, CTAs. Get four or five versions per element. Copy and design are separate steps that feed into each other. Having real options ready before you start the design session means CoDesign has something specific to work with, not a blank brief waiting to be interpreted.&lt;/p&gt;
&lt;p&gt;3.&lt;strong&gt;Feed the brief to the AI design companion.&lt;/strong&gt; With copy variations ready, ask CoDesign to generate initial ad designs. Describe the format, the hierarchy you want, any layout preferences. You’ll get structured, editable designs on the canvas. Not a rendered image, but a working starting point with real layers.&lt;/p&gt;
&lt;p&gt;4.&lt;strong&gt;Apply your brand kit.&lt;/strong&gt; If you haven’t already, load your brand assets: logo, color palette, type system, approved imagery. The agent applies these rules to the designs. Every output from this point respects your brand standards automatically.&lt;/p&gt;
&lt;p&gt;5.&lt;strong&gt;Adapt across formats.&lt;/strong&gt; Tell the agent which formats you need. The social variant, the display variant, the vertical for Stories, the square for feed. Watch the layouts adapt to each context, maintaining the campaign idea and brand consistency across dimensions. If the hierarchy needs adjusting for a specific format, describe what’s not working and the agent fixes it in conversation.&lt;/p&gt;
&lt;p&gt;6.&lt;strong&gt;Refine in conversation.&lt;/strong&gt; This is where the canvas-based approach earns its value. You’re not generating new versions from scratch. You’re iterating on what’s already there, in the same session. “Move the logo to the bottom right. Try the headline in the lighter weight. Swap this layout for something with more white space.” Each exchange builds on the last, so the conversation stays grounded in what’s on the canvas rather than starting over from a new prompt.&lt;/p&gt;
&lt;p&gt;7.&lt;strong&gt;Export and ship.&lt;/strong&gt; When the designs are approved, export in the formats your media plan requires. The session context lives with the file, so if something needs to change post-launch, you’re not starting from zero.&lt;/p&gt;
&lt;p&gt;One honest note: the quality of the output is proportional to the quality of the brief. Vague prompts produce generic starting points. Teams that invest two minutes in a specific, structured brief consistently get more usable first outputs than teams that describe the campaign in one sentence and expect the agent to fill in the gaps.&lt;/p&gt;
&lt;h2 id=&quot;this-is-what-closing-the-loop-actually-looks-like&quot;&gt;This Is What Closing the Loop Actually Looks Like&lt;/h2&gt;
&lt;p&gt;Design agents don’t replace designers. They remove the bottleneck that sits between strategy and execution.&lt;/p&gt;
&lt;p&gt;The marketer who used to wait three days for ad assets can now produce a full campaign set in a single session. The designer who spent half their week on format resizes and small-copy tweaks can spend that time on the work that genuinely requires their judgment. That means brand-defining creative, campaign concepts, and anything where taste and experience are the actual input.&lt;/p&gt;
&lt;p&gt;It was never about willingness to collaborate. The tools just didn’t allow for anything else. Every design change, no matter how small, had to go through a handoff. A headline adjustment on a banner resize does not need a creative director. A resize from 1:1 to 9:16 does not need a brief, a Slack message, and a 48-hour turnaround.&lt;/p&gt;
&lt;p&gt;Thanks to design agents conversation shifts. It moves from “can you make this” to “how should this look.” And that’s a way more interesting conversation.&lt;/p&gt;
&lt;p&gt;Interested in trying IMG.LY CoDesign? &lt;a href=&quot;https://img.ly/forms/contact-sales/&quot;&gt;Reach out&lt;/a&gt; to our team.&lt;/p&gt;</content:encoded><dc:creator>Klaudia</dc:creator><media:content url="https://blog.img.ly/2026/04/page-1-export.png" medium="image"/><category>AI</category><category>Insights</category><category>Creative Workflows</category><category>Creative Automation</category></item><item><title>What Is an AI Design Agent?</title><link>https://img.ly/blog/what-is-a-design-agent/</link><guid isPermaLink="true">https://img.ly/blog/what-is-a-design-agent/</guid><description>&quot;Design agent&quot; is being used to mean two completely different things right now. Here&apos;s what it actually is - and why the distinction matters for creative teams and non-designers in 2026. </description><pubDate>Mon, 30 Mar 2026 04:52:01 GMT</pubDate><content:encoded>&lt;p&gt;The phrase “design agent” is being used in two completely different ways right now, and the confusion is worth clearing up before either meaning becomes the default. Search for it today and you’ll mostly find AI design agencies - services firms that use AI in their process. That’s a reasonable business, but it has nothing to do with what this article is about.&lt;/p&gt;
&lt;p&gt;The second meaning, a design agent as a category of AI tool, is the one that matters for product builders, developers, and creative teams trying to understand where AI-assisted design is actually going. That’s what we’re defining here: what it is, how it works, how it differs from tools you already know, and why the distinction matters if you’re building or evaluating creative software in 2026.&lt;/p&gt;
&lt;h2 id=&quot;first-what-an-ai-design-agent-is-not&quot;&gt;First, What an AI Design Agent Is Not&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;An AI design agency.&lt;/strong&gt; This is a services firm that uses AI tools in its creative process. It’s a completely different category - a service, not a tool - and it’s the most common result you’ll find when you search “design agent” right now.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;An AI image generator.&lt;/strong&gt; Midjourney, DALL-E, and similar tools produce images from text prompts. They generate a single output from a single instruction.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;An AI feature inside a design tool.&lt;/strong&gt; Figma’s AI suggestions, auto-layout assistance, and background removal are AI features. They augment a manual workflow making specific tasks faster.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A design automation pipeline.&lt;/strong&gt; Server-side systems that batch-generate assets from templates are automation, not agency. They’re fast and scalable, but they don’t converse, can’t interpret an ambiguous brief, and don’t refine their output in response to feedback. The confusion here is understandable: modern automation systems use AI models internally and can produce polished, varied-looking results that are easy to mistake for something more intelligent. But give one an ambiguous brief and it either fails or produces something technically correct that misses the point entirely. Give a design agent the same brief and it asks a question. That distinction - executing a fixed process versus reasoning about intent - is what separates the two categories.&lt;/p&gt;
&lt;p&gt;The distinction matters because “design agent” is becoming a meaningful category term, and what it actually describes is different enough from all of the above that collapsing the distinctions creates real confusion, both for people evaluating tools and for people building products.&lt;/p&gt;
&lt;h2 id=&quot;ai-design-agent---a-working-definition&quot;&gt;AI Design Agent - a Working Definition&lt;/h2&gt;
&lt;p&gt;There is no settled definition of an AI design agent yet, the term is still early. Based on what we see in the tools that genuinely deliver on the promise, three properties distinguish a real design agent from something that merely resembles one:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;1. Autonomy.&lt;/strong&gt; The agent takes multi-step actions without requiring a human instruction at each step. Given “create a five-page product catalog in a Scandinavian minimal style,” it works out the layout, typography, and image placement independently and produces the result.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;2. Conversational interface.&lt;/strong&gt; The agent communicates in natural language. It can ask clarifying questions, explain what it’s done, and accept follow-up instructions. The interaction feels like briefing a designer, not operating software.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;3. Refinement loop.&lt;/strong&gt; The agent takes feedback across multiple conversation turns: “make the typography lighter,” “apply the brand’s warm color palette”, and updates the design accordingly. This iterative loop is what separates a design agent from a one-shot generation tool.&lt;/p&gt;
&lt;p&gt;A tool that has all three of these properties is a design agent. A tool that has one or two is something else; probably useful, but in a different category. &lt;a href=&quot;https://img.ly/codesign/&quot;&gt;IMG.LY CoDesign&lt;/a&gt; is built against exactly this bar: it takes a brief, works through the layout on its own, and refines the result in conversation.&lt;/p&gt;
&lt;h2 id=&quot;how-an-ai-design-agent-works-in-practice&quot;&gt;How an AI Design Agent Works in Practice&lt;/h2&gt;
&lt;p&gt;An abstract definition only takes you so far. Here’s what a fully implemented design agent workflow actually looks like, drawn from a real demo of IMG.LY CoDesign. CoDesign runs as a &lt;a href=&quot;https://img.ly/codesign/&quot;&gt;design MCP server&lt;/a&gt;, which means the same agent can also work from Claude, Cursor, or any other MCP client, not just a built-in chat panel.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Step 1.&lt;/strong&gt; The user opens CoDesign. The agent interface - a chat panel - sits alongside the canvas. They’re in the same window.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Step 2.&lt;/strong&gt; The user types a brief: &lt;em&gt;“Here’s a CSV of five products from our furniture brand. Generate one landscape catalog page per product — two-column grid, full-bleed photo left, typography-led layout right. Clean, minimal. Think Hay, Muuto, Frama.”&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Step 3.&lt;/strong&gt; The agent processes the brief, the data, and any brand context it has access to. It generates a five-page catalog in the editor: product names, descriptions, prices, photo placeholders, layout structure consistent across all five pages.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Step 4.&lt;/strong&gt; The user reviews the output and follows up in the chat: &lt;em&gt;“Nice. Pre-fill the photo placeholders with black-and-white product photography, soft contrast.”&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Step 5.&lt;/strong&gt; The agent updates the design. The user manually adjusts one headline, corrects a price, and exports.&lt;/p&gt;
&lt;p&gt;The full workflow, from brief to print-ready output, took minutes rather than hours. The user never left the tool. The agent handled all structural and stylistic decisions. The human handled final judgment and small corrections.&lt;/p&gt;
&lt;p&gt;But generating layouts from a brief is only part of what a design agent can do. In the same chat interface, the agent can build functional decision-making tools directly inside the editor. Ask it to create a Color Themes panel, and it produces a working component: five named presets, each one applying a complete color theme across the entire design in a single click. The user doesn’t switch to a settings screen or manually update individual elements. The agent has built the control they need, right where they need it, as part of the same conversation. That’s a meaningfully different capability: not just producing a designed output, but constructing the tools that let the user make better decisions about that output as they finalize it.&lt;/p&gt;
&lt;h2 id=&quot;design-agents-vs-related-tools&quot;&gt;Design Agents vs. Related Tools&lt;/h2&gt;
&lt;p&gt;This table is meant as a fair comparison, each column represents how these tool categories actually behave today:&lt;/p&gt;















































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Capability&lt;/th&gt;&lt;th&gt;AI Image Generator&lt;/th&gt;&lt;th&gt;AI Design Features&lt;/th&gt;&lt;th&gt;Design Agent&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Takes a brief&lt;/td&gt;&lt;td&gt;Prompt only&lt;/td&gt;&lt;td&gt;No&lt;/td&gt;&lt;td&gt;Yes&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Produces editable layouts&lt;/td&gt;&lt;td&gt;No&lt;/td&gt;&lt;td&gt;Partial&lt;/td&gt;&lt;td&gt;Depends on a tool&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Multi-step autonomy&lt;/td&gt;&lt;td&gt;No&lt;/td&gt;&lt;td&gt;No&lt;/td&gt;&lt;td&gt;Yes&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Conversational refinement&lt;/td&gt;&lt;td&gt;No&lt;/td&gt;&lt;td&gt;Limited&lt;/td&gt;&lt;td&gt;Yes&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Brand / context awareness&lt;/td&gt;&lt;td&gt;No&lt;/td&gt;&lt;td&gt;Partial&lt;/td&gt;&lt;td&gt;Yes&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Iterative across a session&lt;/td&gt;&lt;td&gt;No&lt;/td&gt;&lt;td&gt;No&lt;/td&gt;&lt;td&gt;Yes&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The pattern here isn’t that design agents are better at everything - it’s that they operate at a different level of the workflow. Image generators and AI design features are task-level tools. A design agent is a workflow-level tool.&lt;/p&gt;
&lt;p&gt;There’s one adjacent term worth separating out too: &lt;a href=&quot;https://img.ly/blog/what-is-vibe-design/&quot;&gt;vibe design&lt;/a&gt;, the workflow where you describe what you want instead of building it manually. A design agent is one way to run a vibe design workflow, but the tools behind it differ a lot in what they hand back. We’ve &lt;a href=&quot;https://img.ly/blog/vibe-design-tools-compared/&quot;&gt;compared the AI design tools&lt;/a&gt; built around this workflow if you want to see how they stack up.&lt;/p&gt;
&lt;h2 id=&quot;agentic-design-vs-ai-design-agents-whats-the-difference&quot;&gt;Agentic Design vs. AI Design Agents: What’s the Difference?&lt;/h2&gt;
&lt;p&gt;“Agentic design” and “AI design agent” look interchangeable, and they’re starting to get used that way, but they describe different things.&lt;/p&gt;
&lt;p&gt;Agentic design is a practice: the set of patterns and decisions involved in building systems where an AI acts with some autonomy. How much the agent does before checking in, when it asks a clarifying question, how it surfaces what it changed - these are agentic design questions, and they apply to coding agents and support agents just as much as to creative tools.&lt;/p&gt;
&lt;p&gt;An AI design agent is a product: a tool that applies that practice to design work specifically. The autonomy slider in the next section is a good example of the overlap. Deciding where that slider sits is agentic design; the tool you move it in is a design agent. If you’re evaluating tools, you’re shopping for a design agent. If you’re building one, you’re doing agentic design.&lt;/p&gt;
&lt;h2 id=&quot;creative-agents-the-wider-category&quot;&gt;Creative Agents: The Wider Category&lt;/h2&gt;
&lt;p&gt;Zoom out one level and design agents sit inside a wider family: creative agents. A creative agent is any AI agent that produces creative output autonomously - copy, video, music, imagery, or design. Design agents are the branch of that family that works in layouts, typography, and brand systems.&lt;/p&gt;
&lt;p&gt;The defining properties carry across the family: autonomy over multi-step work, a conversational interface, and a refinement loop. What changes is the medium and the tooling around it - an agent cutting video needs a timeline the way a design agent needs a canvas. Keeping the levels straight helps when you’re mapping the space: the creative agent is the category, the design agent is its design-specific branch, and the products above are competing to define what that branch looks like.&lt;/p&gt;
&lt;h2 id=&quot;the-autonomy-slider-why-human-control-still-matters&quot;&gt;The Autonomy Slider: Why Human Control Still Matters&lt;/h2&gt;
&lt;p&gt;One of the most important design decisions when building or evaluating a design agent is how much autonomy the AI exercises by default.&lt;/p&gt;
&lt;p&gt;Full autonomy is fast. The agent makes all decisions and presents a finished output. But it can produce results that don’t match the user’s actual intent, especially when the brief is ambiguous or the stakes are high.&lt;/p&gt;
&lt;p&gt;Minimal autonomy is safe. The agent only suggests, the human decides everything. But at that point, you’ve lost much of the value of having an agent in the first place.&lt;/p&gt;
&lt;p&gt;The best implementations give users what you might call an &lt;strong&gt;autonomy slider&lt;/strong&gt;: the ability to let the agent run with full independence on some tasks, take targeted direction on others, and step aside entirely when the user wants to edit manually. The right level of autonomy depends on the task and the user’s confidence, not a fixed setting applied uniformly across every interaction.&lt;/p&gt;
&lt;p&gt;For any design agent, this means the interface needs to let the user reach in and adjust manually, mid-workflow, without losing what the agent has already done. The agent and the editor have to exist in the same environment. Separate windows with an export step between them break the loop entirely.&lt;/p&gt;
&lt;h2 id=&quot;where-ai-design-agents-create-the-most-value&quot;&gt;Where AI Design Agents Create the Most Value&lt;/h2&gt;
&lt;p&gt;Not every creative workflow benefits equally from an AI agent. The scenarios where design agents consistently create the most value tend to fall into three categories:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;High-volume, high-variation creative production.&lt;/strong&gt; Marketing teams producing dozens of ad variants, e-commerce teams generating product imagery at scale, publishers creating template-based content in bulk. A team can brief the agent once: brand colors, copy, format specs, and get 40 properly formatted variants back in the time it would have taken to manually produce three, each one consistent with the last.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-designer users who need professional results.&lt;/strong&gt; Not everyone who needs to produce a designed output is a designer. Marketers, retailers, operations teams, small business owners - they know what they want but don’t have the tools or time to build it manually. A design agent gives them a way in that doesn’t require either.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Expert designers who want to explore faster.&lt;/strong&gt; Experienced designers using the agent to generate starting points, explore multiple directions quickly, or offload the time-consuming production work while retaining full control over the final result. A designer can spend 20 minutes reviewing six distinct layout directions the agent generated from a single brief, rather than a full afternoon building each one from scratch.&lt;/p&gt;
&lt;p&gt;These aren’t exhaustive categories, but they represent the use cases where the workflow-level shift actually changes what’s possible, not just what’s faster.&lt;/p&gt;
&lt;h2 id=&quot;a-different-kind-of-design-tool&quot;&gt;A Different Kind of Design Tool&lt;/h2&gt;
&lt;p&gt;A design agent makes creative work faster and more accessible. For designers who want to explore more directions in less time, and for non-designers who have always known what they wanted but lacked the tools to build it.&lt;/p&gt;
&lt;p&gt;That’s what IMG.LY CoDesign is built around: &lt;a href=&quot;https://img.ly/codesign/&quot;&gt;editable designs straight from your agent&lt;/a&gt;, inside a fully featured design editor. The local MCP server is free to install and run, or &lt;a href=&quot;https://img.ly/forms/contact-sales/&quot;&gt;get in touch&lt;/a&gt; if you want to talk it through with the team.&lt;/p&gt;
&lt;h2 id=&quot;frequently-asked-questions&quot;&gt;Frequently Asked Questions&lt;/h2&gt;
&lt;h3 id=&quot;what-is-an-ai-design-agent&quot;&gt;What is an AI design agent?&lt;/h3&gt;
&lt;p&gt;An AI design agent is an AI tool that takes a design brief in natural language, produces an editable design on its own, and refines it through conversation. Three properties define the category: multi-step autonomy, a conversational interface, and a refinement loop. A tool that has all three is a design agent; a tool with one or two belongs to a different category.&lt;/p&gt;
&lt;h3 id=&quot;how-is-an-ai-design-agent-different-from-an-image-generator&quot;&gt;How is an AI design agent different from an image generator?&lt;/h3&gt;
&lt;p&gt;An image generator produces a single flat image from a single prompt. A design agent works at the workflow level: it interprets a brief, makes multi-step layout and typography decisions, produces structured output you can edit element by element, and takes feedback across a session.&lt;/p&gt;
&lt;h3 id=&quot;do-ai-design-agents-replace-designers&quot;&gt;Do AI design agents replace designers?&lt;/h3&gt;
&lt;p&gt;No. The most effective setups keep a human in charge of direction and final judgment. Non-designers use agents to reach professional results they couldn’t build manually; experienced designers use them to explore more directions in less time while keeping full control over the result.&lt;/p&gt;
&lt;h3 id=&quot;are-ai-design-agents-free&quot;&gt;Are AI design agents free?&lt;/h3&gt;
&lt;p&gt;It varies by tool. Google Stitch is free during its beta with monthly generation limits, Figma Make requires AI credits, and IMG.LY CoDesign’s local design MCP server is free to install and run, with no account needed to start.&lt;/p&gt;</content:encoded><dc:creator>Klaudia</dc:creator><media:content url="https://blog.img.ly/2026/03/what-s-a-design-agent.png" medium="image"/><category>AI</category><category>Insights</category></item><item><title>Best AI Design Tools Compared: 6 Vibe Design Platforms Tested</title><link>https://img.ly/blog/vibe-design-tools-compared/</link><guid isPermaLink="true">https://img.ly/blog/vibe-design-tools-compared/</guid><description>Vibe design lets you describe what you want and get a polished result. But not all vibe design tools produce the same kind of output, and that difference shapes everything downstream. Here&apos;s how Stitch, Figma Make, Lovart, Adobe Firefly Boards, Canva Magic Studio and    IMG.LY CoDesign compare. </description><pubDate>Sun, 29 Mar 2026 03:51:56 GMT</pubDate><content:encoded>&lt;p&gt;The best AI design tools no longer just generate images: they take a brief and hand back a finished design. The newest of them are &lt;a href=&quot;https://img.ly/blog/what-is-a-design-agent/&quot;&gt;AI design agents&lt;/a&gt;, tools you direct in conversation rather than operate click by click. The workflow has a name, &lt;a href=&quot;https://img.ly/blog/what-is-vibe-design/&quot;&gt;vibe design&lt;/a&gt;: vibe coding gave non-technical people the ability to build software by describing what they wanted, and the same shift is now happening in design. For the first time, a marketer, a retailer, or a small business owner can create a professional design exactly to their specification just by explaining it to an AI. No design skills required. No external tool. No waiting for a designer to become available.&lt;/p&gt;
&lt;p&gt;This article compares six tools leading that shift - what they do well, where they differ, and how to choose between them.&lt;/p&gt;
&lt;h2 id=&quot;tools-built-for-creators&quot;&gt;Tools Built for Creators&lt;/h2&gt;
&lt;p&gt;These tools represent the current state of standalone vibe design software. Each is genuinely good at what it does. The differences come down to what kind of work they’re built for.&lt;/p&gt;
&lt;h3 id=&quot;google-stitch&quot;&gt;Google Stitch&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://stitch.withgoogle.com&quot;&gt;Google Stitch&lt;/a&gt; is an AI-native design canvas from Google Labs, launched for public beta in early 2026. It accepts text, voice, images, and code as input and produces high-fidelity UI designs and interactive prototypes on an infinite canvas. Its design agent can reason across an entire project rather than just the current frame - a meaningful step beyond tools that only operate on one screen at a time.&lt;/p&gt;
&lt;p&gt;Stitch introduces a DESIGN.md file format for exporting and importing design rules across tools, and it connects to Cursor, Claude Code, and Gemini CLI via an MCP server. It’s available free during beta with monthly generation limits, and it’s clearly aimed at UI/UX designers and developers building product interfaces - not marketing or graphic design work.&lt;/p&gt;
&lt;h3 id=&quot;figma-make&quot;&gt;Figma Make&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://figma.com/make&quot;&gt;Figma Make&lt;/a&gt; takes natural language prompts or existing Figma designs and produces working prototypes and web apps - functional code, not just mockups, with optional Supabase backend integration. Because it lives inside Figma, it has access to your existing design libraries, component systems, and tokens from the start.&lt;/p&gt;
&lt;p&gt;It’s best understood as a rapid ideation and prototyping tool. It generates interactive, working outputs fast, but the code benefits from developer review before going to production. It’s distinct from Figma Sites, which handles actual publishing. AI credits are required for usage.&lt;/p&gt;
&lt;h3 id=&quot;lovart&quot;&gt;Lovart&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://lovart.ai&quot;&gt;Lovart&lt;/a&gt; is a standalone AI design agent founded by a former ByteDance senior product director. Its scope goes well beyond graphic asset creation: from a single prompt, it can generate brand identity systems, UI mockups, video content, packaging designs, and full marketing campaigns. It uses a proprietary MCoT (Mind Chain of Thought) reasoning engine designed to mimic how a creative director thinks - analyzing business context, target audience, and brand requirements, not just aesthetic style.&lt;/p&gt;
&lt;p&gt;Its infinite canvas continuously analyzes all assets present to maintain visual consistency across an entire project, and outputs are compatible with Figma, Photoshop, and After Effects. Lovart is positioned for marketing creatives, solo designers, and small creative teams who need to produce full campaigns without a large agency.&lt;/p&gt;
&lt;h3 id=&quot;adobe-firefly-boards&quot;&gt;Adobe Firefly Boards&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://firefly.adobe.com&quot;&gt;Adobe Firefly Boards&lt;/a&gt; is Adobe’s AI-first collaborative ideation canvas, built for creative professionals who need to move quickly from inspiration to concept. It lives inside the Adobe Firefly web app and mobile apps (iOS and Android), and syncs with Creative Cloud so work can move from moodboard to production without switching platforms.&lt;/p&gt;
&lt;p&gt;The core of Boards is an infinite multimedia canvas where you bring in text prompts, reference images, video clips, and Adobe Stock assets, then generate across all of them at once. What makes it unusual is the model breadth: Boards gives you access to Adobe’s own Firefly models alongside partner models from Runway, Pika, Luma AI, Google Veo, Black Forest Labs, and others, all inside the same workspace. Outputs are images, video clips, text overlays, and assembled mood boards or storyboards, and every generated asset carries Content Credentials that automatically record which AI model produced it.&lt;/p&gt;
&lt;h3 id=&quot;canva-magic-studio&quot;&gt;Canva Magic Studio&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://www.canva.com/magic/&quot;&gt;Canva Magic Studio&lt;/a&gt; is a suite of more than 15 AI-powered tools embedded directly into the Canva editor on web, mobile, and desktop. It’s not a separate app. If you’re already a Canva user, you’re already inside Magic Studio.&lt;/p&gt;
&lt;p&gt;The suite covers writing, image generation, video creation, format switching, and multi-language translation, all accepting text prompts, uploaded media, documents, and existing canvas elements as input. Brand Kit integration is what distinguishes it from more isolated AI generation tools: when you generate a design through Magic Design, it pulls from your saved Brand Kit (logo, colors, fonts) rather than producing something generically styled. That keeps AI-generated output on-brand without manual cleanup.&lt;/p&gt;
&lt;p&gt;The trade-off is structural. Canva outputs are flat exports within a consumer-grade editor, not structured design objects you can manipulate at the element level. Generations are fast and on-brand, but the editing ceiling is Canva’s editor. For users who need granular creative control or complex asset structures, that’s a real constraint, not just a preference.&lt;/p&gt;
&lt;h3 id=&quot;imgly-codesign&quot;&gt;IMG.LY CoDesign&lt;/h3&gt;
&lt;p&gt;IMG.LY CoDesign is AI-powered design companion built for creators and teams producing marketing materials, branded content, social assets, and multi-format campaigns. What sets it apart from other standalone tools is how it treats AI output: every element CoDesign generates is a structured, editable design object on the canvas - not a flat image to export and rebuild elsewhere. Move a headline, swap a font, resize a block, change a color - directly in the same tool that generated it.&lt;/p&gt;
&lt;p&gt;The agent and the editor aren’t two modes. They’re the same tool. You can prompt a direction, refine it in conversation, and take over the canvas directly at any point - without switching tools, exporting, or starting over. CoDesign also goes beyond the canvas: ask it to build a brand panel, a form-based template, or a custom configurator, and it builds that too - directly inside the editor, from the same prompt interface.&lt;/p&gt;
&lt;p&gt;CoDesign also ships as a &lt;a href=&quot;https://img.ly/codesign/&quot;&gt;free local MCP server for design&lt;/a&gt;, so the same agent works from Claude, Cursor, or any other MCP client, not just its own editor.&lt;/p&gt;
&lt;h2 id=&quot;vibe-design-tools-at-a-glance&quot;&gt;Vibe Design Tools at a Glance&lt;/h2&gt;































































































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;&lt;/th&gt;&lt;th&gt;&lt;strong&gt;Google Stitch&lt;/strong&gt;&lt;/th&gt;&lt;th&gt;&lt;strong&gt;Figma Make&lt;/strong&gt;&lt;/th&gt;&lt;th&gt;&lt;strong&gt;Lovart&lt;/strong&gt;&lt;/th&gt;&lt;th&gt;&lt;strong&gt;Adobe Firefly Boards&lt;/strong&gt;&lt;/th&gt;&lt;th&gt;&lt;strong&gt;Canva Magic Studio&lt;/strong&gt;&lt;/th&gt;&lt;th&gt;&lt;strong&gt;IMG.LY CoDesign&lt;/strong&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Primary use&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;UI/UX design from prompts&lt;/td&gt;&lt;td&gt;Prompt-to-prototype / app&lt;/td&gt;&lt;td&gt;Full campaign and brand creation&lt;/td&gt;&lt;td&gt;AI-first concepting and moodboarding&lt;/td&gt;&lt;td&gt;AI design generation inside Canva editor&lt;/td&gt;&lt;td&gt;AI-assisted design for creators and teams&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Where it lives&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Google (standalone web app)&lt;/td&gt;&lt;td&gt;Inside Figma&lt;/td&gt;&lt;td&gt;Lovart (standalone)&lt;/td&gt;&lt;td&gt;Adobe Firefly web + mobile; syncs with Creative Cloud&lt;/td&gt;&lt;td&gt;Inside Canva (web, mobile, desktop)&lt;/td&gt;&lt;td&gt;IMG.LY&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Input types&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Text, voice, images, code&lt;/td&gt;&lt;td&gt;Text, existing Figma designs&lt;/td&gt;&lt;td&gt;Text, images, brand briefs&lt;/td&gt;&lt;td&gt;Text, images, video, Adobe Stock assets&lt;/td&gt;&lt;td&gt;Text, images, video, documents, existing design elements&lt;/td&gt;&lt;td&gt;Text, images, CSV, brand kits&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Output type&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;UI designs and prototypes&lt;/td&gt;&lt;td&gt;Interactive prototypes and web apps&lt;/td&gt;&lt;td&gt;Brand assets, campaigns, video, packaging&lt;/td&gt;&lt;td&gt;Images, video, text overlays, mood boards, storyboards&lt;/td&gt;&lt;td&gt;Images, video, presentations, styled text, format-converted designs&lt;/td&gt;&lt;td&gt;Editable multi-page designs, videos, animations&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Output format&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Flat / exportable&lt;/td&gt;&lt;td&gt;Code / exportable&lt;/td&gt;&lt;td&gt;Flat / exportable&lt;/td&gt;&lt;td&gt;Flat / exportable; Content Credentials attached&lt;/td&gt;&lt;td&gt;Flat / exportable within Canva editor&lt;/td&gt;&lt;td&gt;Structured editable objects on canvas&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Brand context&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Manual per session&lt;/td&gt;&lt;td&gt;Via Figma design system&lt;/td&gt;&lt;td&gt;Canvas-aware within session&lt;/td&gt;&lt;td&gt;Enterprise Custom Models; personal library import&lt;/td&gt;&lt;td&gt;Brand Kit applied automatically at generation&lt;/td&gt;&lt;td&gt;Loaded at session start&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;In-chat UI generation&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;No&lt;/td&gt;&lt;td&gt;No&lt;/td&gt;&lt;td&gt;No&lt;/td&gt;&lt;td&gt;No&lt;/td&gt;&lt;td&gt;No&lt;/td&gt;&lt;td&gt;Yes&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Multi-model AI&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;No&lt;/td&gt;&lt;td&gt;No&lt;/td&gt;&lt;td&gt;No&lt;/td&gt;&lt;td&gt;Yes (Runway, Pika, Luma, Google Veo, BFL, and more)&lt;/td&gt;&lt;td&gt;No (Dream Lab uses Leonardo.ai Phoenix)&lt;/td&gt;&lt;td&gt;No&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Best for&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Product designers, developers&lt;/td&gt;&lt;td&gt;Product teams in Figma ecosystem&lt;/td&gt;&lt;td&gt;Marketing creatives, solo designers&lt;/td&gt;&lt;td&gt;Creative professionals doing concepting and ideation&lt;/td&gt;&lt;td&gt;Non-designers and content teams needing fast, on-brand output&lt;/td&gt;&lt;td&gt;Creators and teams needing editable AI-generated design&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;For a detailed breakdown of how Canva Magic Studio compares to CoDesign specifically, see &lt;a href=&quot;https://img.ly/imgly-codesign-vs-canva-magic-studio/&quot;&gt;Canva Magic Studio vs. CoDesign&lt;/a&gt;. If you’re evaluating from the developer side, CoDesign is also the one that doubles as an &lt;a href=&quot;https://img.ly/codesign/&quot;&gt;AI design tool for developers&lt;/a&gt;: it registers as a local MCP server, so design generation becomes a capability of whatever agent stack you already run.&lt;/p&gt;
&lt;h2 id=&quot;how-to-choose&quot;&gt;How to Choose&lt;/h2&gt;
&lt;p&gt;The criteria here aren’t about feature counts. They’re about the kind of work you’re doing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Choose Google Stitch if&lt;/strong&gt; you’re a product designer or developer building UI and want an AI agent that reasons across your whole project, not just individual screens. It’s the most technically connected of the six (MCP server, code input, DESIGN.md export), and it’s clearly built for product interface work.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Choose Figma Make if&lt;/strong&gt; you’re already in the Figma ecosystem and want to go from prompt to working prototype fast. It inherits your design system automatically, which removes a lot of setup friction, and it outputs functional code rather than static designs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Choose Lovart if&lt;/strong&gt; you need to produce full creative campaigns (brand identity, packaging, video, marketing assets) from a single prompt. Its MCoT reasoning engine and canvas-level consistency analysis make it particularly strong for high-volume marketing creative work.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Choose Adobe Firefly Boards if&lt;/strong&gt; you’re a creative professional working inside the Adobe ecosystem who needs to generate and align on large volumes of visual concepts before moving them into production. It’s the right call for designers, photographers, and creative directors doing serious concepting work that will ultimately land in Photoshop, Premiere, or Illustrator.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Choose Canva Magic Studio if&lt;/strong&gt; you need fast, on-brand creative output and your users aren’t designers. Magic Studio’s Brand Kit integration means generated designs pull from your saved colors, fonts, and logos automatically: no cleanup, no manual adjustment. It’s not the tool for pixel-level creative control, but for marketing teams and content creators producing high volumes across multiple formats and channels, that’s rarely the priority.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Choose CoDesign if&lt;/strong&gt; you need AI-generated output you can actually edit. Its output is &lt;a href=&quot;https://img.ly/codesign/&quot;&gt;editable design files&lt;/a&gt;, not flat exports: every element lands on the canvas as a structured design object (headline, image block, color field). That matters when the brief changes, the brand needs adjusting, or the output needs to become ten variations rather than one. For creators and teams producing branded content at scale, that editability is the difference between a starting point and a finished asset.&lt;/p&gt;
&lt;p&gt;CoDesign is in early access. Be among the first to prompt a direction and walk away with a design you can actually use. &lt;a href=&quot;https://img.ly/forms/contact-sales/&quot;&gt;Talk to our team&lt;/a&gt; to learn more.&lt;/p&gt;</content:encoded><dc:creator>Klaudia</dc:creator><media:content url="https://blog.img.ly/2026/03/vibedesign-tool-comparison-1.png" medium="image"/><category>AI</category><category>Insights</category></item><item><title>What Is Vibe Design? The Definitive Guide for Product Builders, Designers, and Creative Teams</title><link>https://img.ly/blog/what-is-vibe-design/</link><guid isPermaLink="true">https://img.ly/blog/what-is-vibe-design/</guid><description>Vibe design just got a name - but the shift it describes has been building for years. Here&apos;s what it actually means, where it came from, and what it means if you&apos;re building a product where users create things.</description><pubDate>Fri, 20 Mar 2026 17:23:51 GMT</pubDate><content:encoded>&lt;h2 id=&quot;vibe-design---a-concept-that-just-got-a-name&quot;&gt;Vibe Design - A Concept That Just Got a Name&lt;/h2&gt;
&lt;p&gt;Creative work has always involved describing an idea and having someone skilled make it real. Vibe design follows the same principle, except the “someone” is AI, ready whenever you are, no brief needed. You describe what you want through text, images, brand assets, or sketches, and the AI generates the design. You direct the process; the tool handles the making.&lt;/p&gt;
&lt;p&gt;The term itself is not new but was brought into mainstream use in early 2026 by Google via their &lt;a href=&quot;https://blog.google/innovation-and-ai/models-and-research/google-labs/stitch-ai-ui-design/&quot;&gt;Stitch announcement&lt;/a&gt;, and within a day it was appearing across The Register, CNBC, and TechRadar. The practice it describes had been building for well over a year, in Figma’s AI features, in generative design tools, in the growing number of products letting users create from a prompt rather than from a toolbox.&lt;/p&gt;
&lt;p&gt;Today, we’ll define what vibe design actually is, trace where the idea came from, and explain what it means for people building products.&lt;/p&gt;
&lt;h2 id=&quot;where-the-term-comes-from-vibe-codings-design-sibling&quot;&gt;Where the Term Comes From: Vibe Coding’s Design Sibling&lt;/h2&gt;
&lt;p&gt;To understand vibe design, it helps to start with the concept it’s directly descended from: vibe coding.&lt;/p&gt;
&lt;p&gt;Former Director of AI at Tesla and co-founder of OpenAI, &lt;a href=&quot;https://x.com/karpathy/status/1886192184808149383&quot;&gt;Andrej Karpathy coined “vibe coding” in early 2025&lt;/a&gt; to describe a shift in how developers work with AI-generated code. The idea: instead of writing implementation yourself, you describe what you want in natural language and an AI writes the code. The developer’s role shifts from implementation to direction. You stop thinking about &lt;em&gt;how&lt;/em&gt; things are built and focus entirely on &lt;em&gt;what&lt;/em&gt; you want to create.&lt;/p&gt;
&lt;p&gt;Vibe design is the same shift applied to visual creation. Instead of opening a design tool and manually placing elements, adjusting colors, choosing fonts and spacing, you describe what you want (in words, or by uploading a reference image, a photo, a sketch, a brand kit), and the design takes shape. The phrase “vibe coding design” captures this lineage neatly: it’s the same intent-first principle, moved from engineering into the visual layer. The tool on the other side of that description is, increasingly, an &lt;a href=&quot;https://img.ly/blog/what-is-a-design-agent/&quot;&gt;AI design agent&lt;/a&gt;: software that takes the brief, does the multi-step work itself, and stays in the conversation for refinement.&lt;/p&gt;
&lt;p&gt;The underlying movement had been gathering momentum since AI image generators went mainstream, accelerating rapidly as tools like Figma Make, Lovart, and now Google Stitch brought the concept directly into design workflows. The label arrived late to a trend that had already changed how a lot of creative work gets done.&lt;/p&gt;
&lt;h2 id=&quot;what-vibe-design-actually-means-a-working-definition&quot;&gt;What Vibe Design Actually Means: A Working Definition&lt;/h2&gt;
&lt;p&gt;Here is a working definition that holds up across tools and use cases:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vibe design is a creative workflow in which the primary input is intent (described in natural language or visual references) rather than manual manipulation of design tools. The designer’s role becomes one of direction, curation, and refinement rather than construction.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Three things define a genuine vibe design workflow:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;1. Intent-first input.&lt;/strong&gt; The starting point is a brief, a description, a reference image, or a combination, not a blank canvas and a toolbox. You’re communicating what you want, not building it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;2. Generative execution.&lt;/strong&gt; An AI interprets that intent and produces a designed output: a layout, a color scheme, a complete page, a set of variations. The construction step is handled by the system.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;3. Human refinement in the loop.&lt;/strong&gt; The human stays involved throughout, approving directions, adjusting outputs, steering away from things that don’t work. The AI handles execution; the human handles judgment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What vibe design is not:&lt;/strong&gt; it’s not simply using an AI image generator to produce pictures. Image generation is one possible input &lt;em&gt;into&lt;/em&gt; a vibe design workflow, not the workflow itself. Vibe design produces editable, structured design outputs (layouts, components, documents, campaigns), not static images. The output is something you can work with, not just something you can look at.&lt;/p&gt;
&lt;h2 id=&quot;how-it-differs-from-ai-assisted-design&quot;&gt;How It Differs from AI-Assisted Design&lt;/h2&gt;
&lt;p&gt;“AI-assisted design” has covered a lot of ground over the past few years: autocomplete for design tokens, background removal, content generation within a layout. These are useful additions to a manual workflow. But in all of them, the designer still drives. AI is a tool called on for specific tasks while the human remains in the seat.&lt;/p&gt;
&lt;p&gt;Vibe design flips the ratio. The AI drives the initial creation; the human steers and refines. It’s a different relationship with the tool, not a faster version of the same one. The clearest expression of that relationship is a &lt;a href=&quot;https://img.ly/codesign/&quot;&gt;design agent that returns editable files&lt;/a&gt;: you describe the direction, the agent builds the layout, and everything it produces stays open to manual editing afterward.&lt;/p&gt;
&lt;p&gt;The distinction matters because it changes three things:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;What skills are most useful.&lt;/strong&gt; Writing clear, directed prompts and making fast curatorial judgments matters more than knowing every keyboard shortcut.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;What the workflow looks like.&lt;/strong&gt; You’re reviewing and steering outputs rather than constructing from scratch.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;What software is relevant.&lt;/strong&gt; Vibe design tools are built around a different interaction model than tools built to accelerate traditional manual design.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;vibe-design-in-practice-three-scenarios&quot;&gt;Vibe Design in Practice: Three Scenarios&lt;/h2&gt;
&lt;p&gt;The best way to make this concrete is to show what an AI design agent workflow built around vibe design actually looks like. These three scenarios cover the range of contexts where it’s becoming relevant.&lt;/p&gt;
&lt;h3 id=&quot;scenario-1-a-marketing-team-no-designer-available&quot;&gt;Scenario 1: A Marketing Team, No Designer Available&lt;/h3&gt;
&lt;p&gt;A marketing manager at a mid-sized e-commerce brand needs a product launch campaign for social media. There’s no designer available this week. They’re tied up on a bigger project.&lt;/p&gt;
&lt;p&gt;She opens the creative tool embedded in their marketing platform, uploads the product photo and the brand guidelines, and types: &lt;em&gt;“Create a campaign for our summer collection, clean, minimal, white space heavy, headline-driven.”&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;She receives a set of formatted, brand-consistent assets sized for each social channel. The layouts are on-brand. The typography follows the guidelines she uploaded. She adjusts the headline copy on two of the assets and swaps one background color. The whole thing takes 12 minutes.&lt;/p&gt;
&lt;p&gt;No design skills required. No third-party tool. No waiting for a designer to become available. The campaign goes out on schedule.&lt;/p&gt;
&lt;h3 id=&quot;scenario-2-a-designer-exploring-variations&quot;&gt;Scenario 2: A Designer Exploring Variations&lt;/h3&gt;
&lt;p&gt;A senior designer is working on a brand campaign for a luxury lifestyle client. She’s settled on a layout she likes, but she can’t land on the right color direction. Everything she’s tried manually feels either too cold or too safe, and exporting variations to compare them side by side is eating time she doesn’t have.&lt;/p&gt;
&lt;p&gt;Instead, she types a single instruction into the agent chat: &lt;em&gt;“Add a panel on the left with five color theme presets I can click to instantly apply to my design.”&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The agent builds the panel directly inside the editor. Five named presets (Warm Sand, Midnight, Rose Quartz, Forest, Slate Blue), each applying a complete color theme across the entire design in one click: backgrounds, accents, headings, body text, all updated together. She works through all five in under a minute and finds the direction she was looking for without typing another prompt.&lt;/p&gt;
&lt;p&gt;The variation-exploration workflow, at its most useful, doesn’t just produce more outputs; it builds the tools you need to make the decision faster.&lt;/p&gt;
&lt;h3 id=&quot;scenario-3-a-product-team-embedding-creative-capability&quot;&gt;Scenario 3: A Product Team Embedding Creative Capability&lt;/h3&gt;
&lt;p&gt;A print-on-demand platform serves retailers and small brands who need to produce product catalogues regularly but don’t have in-house design resource. One of their customers (a retailer for a furniture brand) opens the editor, pastes a CSV of five products into the agent chat, and describes the layout style she wants: two-column landscape, typography-led, minimal, referencing the aesthetic of Hay, Muuto, and Frama.&lt;/p&gt;
&lt;p&gt;The agent generates a complete five-page catalogue inside the editor, one page per product, consistent layout throughout, with product names, descriptions, prices, and photo placeholders already in place. She follows up in plain language: &lt;em&gt;“Pre-fill the photo placeholders with elegant product photography, make them black and white, soft contrast.”&lt;/em&gt; The agent updates all five pages. She adjusts one headline manually and exports.&lt;/p&gt;
&lt;h2 id=&quot;the-vibe-design-tools-shaping-the-space-right-now&quot;&gt;The Vibe Design Tools Shaping the Space Right Now&lt;/h2&gt;
&lt;p&gt;Several tools are explicitly built around this workflow.&lt;/p&gt;



































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Tool&lt;/th&gt;&lt;th&gt;What it does&lt;/th&gt;&lt;th&gt;Where the agent lives&lt;/th&gt;&lt;th&gt;Best for&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Google Stitch&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Voice and text prompts to UI design&lt;/td&gt;&lt;td&gt;Google’s standalone tool&lt;/td&gt;&lt;td&gt;UI/UX designers, developers&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Figma Make&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Prompt to prototype inside Figma&lt;/td&gt;&lt;td&gt;Inside Figma (standalone)&lt;/td&gt;&lt;td&gt;Product designers working in Figma&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Lovart&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;AI design agent for graphic creation&lt;/td&gt;&lt;td&gt;Lovart’s standalone platform&lt;/td&gt;&lt;td&gt;Marketing creatives, solo designers&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;IMG.LY&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Design companion for all types of design works and tasks&lt;/td&gt;&lt;td&gt;IMG.LY Codesign studio&lt;/td&gt;&lt;td&gt;Designers, marketers, brand owners&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Google Stitch&lt;/strong&gt; is built around the idea that UI design should start with a conversation. You describe a screen (its purpose, the actions it needs to support, the general feel), and Stitch produces an interface design you can refine. It’s aimed at developers and UI/UX designers who want to move faster in the early stages of building a product interface. Where it works well is in getting from a rough idea to a structured screen layout without having to make every decision from scratch.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Figma Make&lt;/strong&gt; extends the environment that product designers already work in. Because it lives inside Figma, it has access to your existing components, tokens, and design system. The prompt-to-prototype workflow is useful for designers who want to explore how a brief might translate into a working layout without manually composing every frame. Its biggest advantage is that the output lands directly in a space where a full design team can take over.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Lovart&lt;/strong&gt; is focused on graphic and campaign creation rather than UI or product design. It’s built for the kind of work that marketing creatives and solo designers do a lot of, producing visual assets for social, campaigns, brand activations. The emphasis is on speed and aesthetic quality for graphic outputs rather than on structured, component-based design systems.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;IMG.LY&lt;/strong&gt; brings a design companion together with manual canvas edits through &lt;a href=&quot;https://img.ly/codesign/&quot;&gt;IMG.LY CoDesign&lt;/a&gt; - each element created by AI can be manually moved, changed, or adapted. It maintains brand and template context throughout the generation process, redefining what a template means: from a simple asset with placeholders to a true brand guideline. It also stands out for the breadth of its import and export options, which makes it a strong fit for graphic and asset design work.&lt;/p&gt;
&lt;h2 id=&quot;where-vibe-design-has-limits&quot;&gt;Where Vibe Design Has Limits&lt;/h2&gt;
&lt;p&gt;Vibe design works well when there’s something to work from, like a reference image, a brand kit, or an existing visual direction. When there’s genuinely nothing to draw from, outputs tend toward the generic. A completely new brand identity with no existing visual language is a poor fit for a vibe design workflow; that kind of work still benefits from the deliberate, decision-by-decision process of traditional design.&lt;/p&gt;
&lt;p&gt;It’s also less suited to accessibility-critical UI, where precise specification (contrast ratios, touch targets, interaction states) matters more than mood or aesthetic direction. A generated layout might look right without being accessible, and catching that requires careful manual review.&lt;/p&gt;
&lt;p&gt;Finally, the more technically constrained the brief, the more refinement the output will need. Vibe design compresses the path to a starting point; it doesn’t always compress the path to a final, production-ready output. Teams that go in expecting to iterate will get more out of it than teams that expect to export and ship.&lt;/p&gt;
&lt;h2 id=&quot;the-human-element-vibe-design-is-not-autonomous-design&quot;&gt;The Human Element: Vibe Design Is Not Autonomous Design&lt;/h2&gt;
&lt;p&gt;A common concern about vibe design workflows is that they reduce the role of skill and judgment in creative work. The evidence so far points the other way.&lt;/p&gt;
&lt;p&gt;The most effective workflows keep the human firmly in control of direction, curation, and final judgment. The AI generates; the human decides what’s good, what fits the brief, what needs to change. The setup that makes this work in practice is &lt;a href=&quot;https://img.ly/codesign/&quot;&gt;one file, three authors&lt;/a&gt;: the agent, your code, and you all working in the same editable design instead of passing exports back and forth. Removing the human layer doesn’t improve outcomes; it just produces more output with no quality filter.&lt;/p&gt;
&lt;p&gt;What changes is not &lt;em&gt;whether&lt;/em&gt; human judgment matters, but &lt;em&gt;at which stage&lt;/em&gt; it matters most. In a traditional design workflow, judgment is exercised continuously, at every click, every color choice, every alignment decision. In a vibe design workflow, judgment operates at a higher level: Is this the right direction? Does this match the intent? What needs to change?&lt;/p&gt;
&lt;p&gt;The craft is still there. The instruments are different.&lt;/p&gt;
&lt;h2 id=&quot;a-shift-thats-already-underway&quot;&gt;A Shift That’s Already Underway&lt;/h2&gt;
&lt;p&gt;Vibe design isn’t a trend arriving from the future. It’s a name for a shift that’s been building for several years and just became visible enough to label properly.&lt;/p&gt;
&lt;p&gt;The creative AI tools exist. The workflows are being adopted. The user expectation is forming. Naming the practice in 2026 didn’t create the movement; it just gave it a shared vocabulary that makes it easier to talk about and build toward.&lt;/p&gt;
&lt;p&gt;If what you’re looking for is a balance between manual control of the output and power of AI generation, IMG.LY Codesign might be the right fit for you. &lt;a href=&quot;https://img.ly/forms/contact-sales/&quot;&gt;Talk to our team&lt;/a&gt; to see how it can fit inside your stack.&lt;/p&gt;</content:encoded><dc:creator>Klaudia</dc:creator><media:content url="https://blog.img.ly/2026/03/vibedesign-tools-ai.jpg" medium="image"/><category>AI</category><category>Agent Skills</category><category>Insights</category></item><item><title>CE.SDK v1.69 Release Notes</title><link>https://img.ly/blog/creative-editor-sdk-v-1-69-0-release-notes/</link><guid isPermaLink="true">https://img.ly/blog/creative-editor-sdk-v-1-69-0-release-notes/</guid><description>Agent Skills turns your AI coding assistant into a CE.SDK expert. Ship editors in minutes, not days. Plus Starter Kits, PPTX import, and video editor updates.</description><pubDate>Fri, 27 Feb 2026 18:29:53 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;v1.69 is the biggest developer experience release we’ve shipped.&lt;/strong&gt; &lt;em&gt;Agent Skills&lt;/em&gt; collapse the time from idea to working editor from days to under a minute. Production-ready &lt;em&gt;Starter Kits&lt;/em&gt; eliminate days of boilerplate setup. &lt;em&gt;PPTX and Canva Importer&lt;/em&gt; lets users bring their existing designs directly into your editor. And a comprehensive video overhaul across Web, iOS, and Android closes the gap between CE.SDK and native professional editing apps.&lt;/p&gt;
&lt;p&gt;Let’s dive in!&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;introducing-imgly-agent-skills-for-cesdk&quot;&gt;Introducing IMG.LY Agent Skills for CE.SDK&lt;/h2&gt;
&lt;p&gt;&lt;video src=&quot;https://blog.img.ly/2026/02/agent-skill-design-editor-photo-editor-video-whitelabel-sdk.mp4&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The biggest change to CE.SDK’s developer experience since launch.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;AI agents have changed how developers build. With Agent &lt;em&gt;Skills&lt;/em&gt;, CE.SDK works natively with your AI agent.&lt;/p&gt;
&lt;p&gt;Install the plugin for Claude Code or Cursor and your AI coding agent becomes a CE.SDK expert with bundled access to guides, API references, and starter kits across 10 web frameworks. No external services, no MCP servers, no context-switching.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What your developers can now do:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;/cesdk:build&lt;/code&gt;&lt;/strong&gt; — Describe a use case in plain language. The agent detects the framework, pulls the right starter kit, and scaffolds a working project autonomously.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;/cesdk:explain&lt;/code&gt;&lt;/strong&gt; — Ask any CE.SDK implementation question and get answers adapted to the specific framework and experience level.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;/cesdk:docs-[framework]&lt;/code&gt;&lt;/strong&gt; — Retrieve guides and API references directly from the IDE, without leaving the editor.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1776px) 1776px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1776&quot; height=&quot;1073&quot; src=&quot;https://img.ly/_astro/build-photo-editor-video-editor-with-claude-codex-ai-chatgpt-1_Z2pLGsr.webp&quot; srcset=&quot;/_astro/build-photo-editor-video-editor-with-claude-codex-ai-chatgpt-1_ZSgJxt.webp 640w, /_astro/build-photo-editor-video-editor-with-claude-codex-ai-chatgpt-1_2mBBIp.webp 750w, /_astro/build-photo-editor-video-editor-with-claude-codex-ai-chatgpt-1_7zDq3.webp 828w, /_astro/build-photo-editor-video-editor-with-claude-codex-ai-chatgpt-1_1m1QkW.webp 1080w, /_astro/build-photo-editor-video-editor-with-claude-codex-ai-chatgpt-1_1wOErv.webp 1280w, /_astro/build-photo-editor-video-editor-with-claude-codex-ai-chatgpt-1_ZPvEli.webp 1668w, /_astro/build-photo-editor-video-editor-with-claude-codex-ai-chatgpt-1_Z2pLGsr.webp 1776w&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;For your product, this means:&lt;/strong&gt; developers prototype in minutes, not days. Non-technical team members can now build and extend functional editors independently. And onboarding new engineers to CE.SDK shrinks from a multi-day ramp to a single session.&lt;/p&gt;
&lt;p&gt;→ &lt;a href=&quot;https://img.ly/blog/img-ly-agent-skills-web/&quot;&gt;Quick Start &amp;#x26; Full Announcement&lt;/a&gt;&lt;br&gt;
→ &lt;a href=&quot;https://img.ly/docs/cesdk/js/get-started/agent-skills-f7g8h9/?ref=img.ly&quot;&gt;Explore Agent Skills Documentation&lt;/a&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;launch-in-minutes-with-production-ready-starter-kits-for-web&quot;&gt;Launch in Minutes with Production-Ready Starter Kits for Web&lt;/h2&gt;
&lt;p&gt;&lt;video src=&quot;https://blog.img.ly/2026/02/starterkits-web-photo-editor-video-editor-design-sdk.mp4&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;p&gt;Previously, getting CE.SDK running in a real product meant adapting complex demo projects: stripping features, adjusting architecture, and undoing assumptions baked into example code.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Starter Kits replace all of that with pre-configured, production-ready editor UIs.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Specialized kits for Photo, Video, and Design editors are scoped, clean, and immediately deployable. Each kit starts with a focused feature set. You explicitly opt in to what you need, so your UI stays performant and your users don’t see capabilities they can’t use.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;For your product, this means:&lt;/strong&gt; a functional, shippable editor in minutes rather than days of setup. Clean architecture from day one, without accumulating technical debt.&lt;/p&gt;
&lt;p&gt;→ &lt;a href=&quot;https://img.ly/docs/cesdk/js/get-started/vanilla-js/quickstart-yef23s/&quot;&gt;Explore Starter Kit Documentation&lt;/a&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;let-users-bring-their-pptx--canva-designs-to-cesdk&quot;&gt;Let Users Bring Their PPTX &amp;#x26; Canva Designs to CE.SDK&lt;/h2&gt;
&lt;p&gt;&lt;video src=&quot;https://blog.img.ly/2026/02/pptx-canva-importer-import-canva-designs.mp4&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;p&gt;Some customers don’t start from scratch. They have existing creative assets locked inside PowerPoint and Canva, and until now, bringing those into CE.SDK required manual recreation.&lt;/p&gt;
&lt;p&gt;The new PPTX Importer removes that barrier.&lt;/p&gt;
&lt;p&gt;Users can import native PowerPoint files or export &lt;strong&gt;Canva designs&lt;/strong&gt; as PPTX and bring them directly into your CE.SDK-powered editor, with layouts, elements, and structure intact.&lt;/p&gt;
&lt;p&gt;→ &lt;a href=&quot;https://img.ly/demos/pptx-template-import/web/&quot;&gt;Try the PPTX Import Demo&lt;/a&gt;&lt;br&gt;
→ &lt;a href=&quot;https://www.npmjs.com/package/@imgly/pptx-importer&quot;&gt;View PPTX npm&lt;/a&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;professional-grade-video-editing-across-platforms&quot;&gt;Professional-Grade Video Editing Across Platforms&lt;/h2&gt;
&lt;p&gt;This release expands CE.SDK’s video features across web &amp;#x26; mobile platforms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Control Video Speed (iOS &amp;#x26; Android)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;video src=&quot;https://blog.img.ly/2026/02/slow-speed-on-reel-tiktok-videos-sdk.mp4&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;p&gt;Users can now adjust the playback speed of video and audio clips via a dedicated Speed UI. This has been one of the most-requested mobile features and is now available on both platforms.&lt;br&gt;
→ &lt;a href=&quot;https://img.ly/docs/cesdk/ios/create-video/control-daba54/#playback-speed&quot;&gt;View Video Speed Documentation iOS&lt;/a&gt;&lt;br&gt;
→ &lt;a href=&quot;https://img.ly/docs/cesdk/android/create-video/control-daba54/#playback-speed&quot;&gt;View Video Speed Documentation Android&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Create Video Groups in the Timeline (Web)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;video src=&quot;https://blog.img.ly/2026/02/group-timeline-tracks.mp4&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;p&gt;Multiple clips can now be combined into a single manageable unit with synchronized timing, trimming, and a unified timeline representation.&lt;/p&gt;
&lt;p&gt;This enables complex, multi-layer content production and speeds up the design process.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Video Timeline Now Grows Automatically (Web)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;video src=&quot;https://blog.img.ly/2026/02/growing-timeline.mp4&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;p&gt;The timeline now automatically adjusts its height based on content. Manual track configuration is gone, and complex projects are easier to pick up for your new users.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Use a Clutter-Free iOS Timeline (iOS)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1480px) 1480px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1480&quot; height=&quot;643&quot; src=&quot;https://img.ly/_astro/cleaner-timeline-ios_Z25tajS.webp&quot; srcset=&quot;/_astro/cleaner-timeline-ios_Z1hQmLR.webp 640w, /_astro/cleaner-timeline-ios_2nXdsr.webp 750w, /_astro/cleaner-timeline-ios_jkXE4.webp 828w, /_astro/cleaner-timeline-ios_2in6s2.webp 1080w, /_astro/cleaner-timeline-ios_6rsyz.webp 1280w, /_astro/cleaner-timeline-ios_Z25tajS.webp 1480w&quot;&gt;&lt;/p&gt;
&lt;p&gt;Non-video tracks now show a single thumbnail per clip, reducing clutter in multi-layer timelines on your iOS app.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;full-changelog&quot;&gt;Full Changelog&lt;/h2&gt;
&lt;p&gt;See all technical details, breaking changes, and performance improvements in the &lt;a href=&quot;https://img.ly/docs/cesdk/changelog/v1-69-0/&quot;&gt;v1.69.0 Changelog.&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Thank you for building with IMG.LY.&lt;/em&gt;&lt;/p&gt;</content:encoded><dc:creator>Neslihan</dc:creator><media:content url="https://blog.img.ly/2026/02/creative-editor-sdk-v169-imgly-whitelabel-design-editor-video-editor-react.jpg" medium="image"/><category>Release Notes</category><category>CE.SDK</category><category>Agent Skills</category><category>Agentic Development</category><category>AI</category><category>Starter Kits</category></item><item><title>Introducing IMG.LY Agent Skills</title><link>https://img.ly/blog/img-ly-agent-skills-web/</link><guid isPermaLink="true">https://img.ly/blog/img-ly-agent-skills-web/</guid><description>Inject CE.SDK expertise into your IDE with Agent Skills, turning AI assistants into autonomous implementation partners, building editors in minutes.</description><pubDate>Mon, 16 Feb 2026 18:54:06 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;br&gt;
We are moving beyond static documentation. &lt;em&gt;IMG.LY Agent Skills&lt;/em&gt; is a specialized intelligence layer that transforms AI coding assistants (Claude Code, Cursor, Windsurf) into CE.SDK implementation experts. Instead of mapping APIs manually, you now inject our full SDK expertise directly into your IDE and coding agent.&lt;/p&gt;
&lt;p&gt;→ &lt;a href=&quot;https://img.ly/docs/cesdk/js/get-started/agent-skills-f7g8h9/&quot;&gt;Explore Agent Skill Documentation&lt;/a&gt;&lt;br&gt;
→ &lt;a href=&quot;https://github.com/imgly/agent-skills&quot;&gt;Access Agent Skills GitHub Repository&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;the-quick-start-paradox&quot;&gt;The “Quick Start” Paradox&lt;/h2&gt;
&lt;p&gt;In our recent analysis of client integrations, a consistent pattern emerged: even with a high-performance SDK, the “Time to First Edit” is often delayed by research. Developers spend a significant portion of their initial integration time context-switching between their IDE and documentation, manually mapping complex API hierarchies to their specific framework.&lt;/p&gt;
&lt;p&gt;We figured: You don’t just need better documentation; you need an &lt;strong&gt;expert partner&lt;/strong&gt; who is already inside your codebase.&lt;/p&gt;
&lt;h2 id=&quot;the-solution-imgly-agent-skills-for-web&quot;&gt;The Solution: IMG.LY Agent Skills for Web&lt;/h2&gt;
&lt;p&gt;We are moving beyond static documentation. Today, we are launching &lt;strong&gt;IMG.LY Agent Skills&lt;/strong&gt;, a specialized intelligence layer that transforms AI coding assistants like Claude Code, Cursor, and Gemini into CE.SDK implementation experts.&lt;/p&gt;
&lt;p&gt;Instead of searching for answers, you now give your agent our complete SDK knowledge with a single command.&lt;/p&gt;
&lt;h2 id=&quot;how-it-works-your-autonomous-implementation-partner&quot;&gt;How It Works: Your Autonomous Implementation Partner&lt;/h2&gt;
&lt;p&gt;By injecting versioned, live documentation and framework-specific starter kits directly into your agent’s context window, we’ve created three adaptive paths to launch:&lt;/p&gt;
&lt;h3 id=&quot;1-the-explain-path&quot;&gt;1. The Explain Path&lt;/h3&gt;
&lt;p&gt;&lt;video src=&quot;https://blog.img.ly/2026/02/0-explain-skill.mp4&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;p&gt;Architecture is complex. The Explain Skill acts as a digital implementation partner that adapts to you. It can provide a technical deep-dive on the export pipeline for a Senior Dev or a high-level summary for a Product Manager, in any language and at any level of detail.&lt;/p&gt;
&lt;h3 id=&quot;2-the-build-path&quot;&gt;&lt;strong&gt;2. The Build Path&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;video src=&quot;https://blog.img.ly/2026/02/1-build-skill.mp4&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;p&gt;Command your agent to “Add a photo editor to my React app.” The agent autonomously detects your environment, pulls the correct starter kit, and handles the boilerplate. It doesn’t just show you code; it builds the foundation for you.&lt;/p&gt;
&lt;h3 id=&quot;3-the-docs-path&quot;&gt;&lt;strong&gt;3. The Docs Path&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Stop the tab-switching. Your agent retrieves versioned API references offline, keeping your focus entirely within the terminal or IDE.&lt;/p&gt;
&lt;h2 id=&quot;the-shift-to-autonomous-engineering&quot;&gt;The Shift to Autonomous Engineering&lt;/h2&gt;
&lt;p&gt;This release marks a strategic shift for IMG.LY. We believe the next era of software development isn’t about “better documentation,” but about &lt;strong&gt;Portable Expertise&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;By providing that AI with a “skill” rather than a manual, we are empowering teams to move from a multi-month build cycle to a 30-second scaffold.&lt;/p&gt;
&lt;h2 id=&quot;supported-frameworks&quot;&gt;Supported Frameworks&lt;/h2&gt;
&lt;p&gt;The initial Web release supports 10 frameworks out of the box:&lt;br&gt;
&lt;strong&gt;React · Vue.js · Svelte · Angular · Next.js · Nuxt.js · SvelteKit · Electron · Node.js · Vanilla JS&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id=&quot;get-started&quot;&gt;Get Started&lt;/h2&gt;
&lt;p&gt;Run the command:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;npx&lt;/span&gt;&lt;span&gt; skills&lt;/span&gt;&lt;span&gt; add&lt;/span&gt;&lt;span&gt; imgly/agent-skills&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;em&gt;For example&lt;/em&gt;, &lt;strong&gt;Claude Code&lt;/strong&gt; users, can run:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;claude&lt;/span&gt;&lt;span&gt; plugin&lt;/span&gt;&lt;span&gt; marketplace&lt;/span&gt;&lt;span&gt; add&lt;/span&gt;&lt;span&gt; imgly/agent-skills&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;claude&lt;/span&gt;&lt;span&gt; plugin&lt;/span&gt;&lt;span&gt; install&lt;/span&gt;&lt;span&gt; cesdk@imgly&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;→ &lt;a href=&quot;https://img.ly/docs/cesdk/js/get-started/agent-skills-f7g8h9/&quot;&gt;Explore the full Agent Skill Documentation&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The “editor of your dreams” is now a conversation away.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;Join 3,000+ creative professionals who get early access to new features and updates.&lt;/strong&gt; &lt;a href=&quot;https://share.hsforms.com/1IgAOV1wASXGPnFG4ZPLejg1hk3i?ref=img.ly&quot;&gt;&lt;strong&gt;Subscribe&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;.&lt;/strong&gt;&lt;/p&gt;</content:encoded><dc:creator>Neslihan</dc:creator><media:content url="https://blog.img.ly/2026/02/Build-Design-Video-Photo-Editor-With-AI-Agent-sdk-imgly-2.jpg" medium="image"/><category>AI</category><category>CE.SDK</category></item><item><title>Build in a Day: AI Video Clipping with CE.SDK</title><link>https://img.ly/blog/build-in-a-day-ai-video-clipping-with-ce-sdk/</link><guid isPermaLink="true">https://img.ly/blog/build-in-a-day-ai-video-clipping-with-ce-sdk/</guid><description>Build a client-side AI video editor that turns long videos into short highlights using CE.SDK and AI.</description><pubDate>Thu, 05 Feb 2026 12:17:28 GMT</pubDate><content:encoded>&lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;We built a video shortener in a single day using Claude Code and CE.SDK. It extracts 3-4 short clips from long-form video, handles transcription, identifies the best moments via AI, detects speakers, and outputs vertical/horizontal/square formats, all running in the browser.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Features:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Extracts 3-4 clips per video (highlights, summaries, or cleaned-up edits)&lt;/li&gt;
&lt;li&gt;Outputs 9:16 (vertical), 16:9 (landscape), or 1:1 (square)&lt;/li&gt;
&lt;li&gt;Detects speakers and maps them to faces with user confirmation&lt;/li&gt;
&lt;li&gt;Auto-crops to follow the active speaker&lt;/li&gt;
&lt;li&gt;Adds captions and text hooks&lt;/li&gt;
&lt;li&gt;Non-destructive: change aspect ratio or template without re-processing&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Best suited for:&lt;/strong&gt; Videos with speech/dialogue (podcasts, interviews, presentations, vlogs)&lt;/p&gt;
&lt;h2 id=&quot;why-client-side&quot;&gt;Why Client-Side?&lt;/h2&gt;
&lt;p&gt;CE.SDK’s CreativeEngine runs in the browser via WebAssembly. Video decoding, timeline manipulation, effects, and preview all happen on the user’s device.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Benefits:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;No upload/download wait: edits preview instantly&lt;/li&gt;
&lt;li&gt;Non-destructive: switch aspect ratio or template without rendering&lt;/li&gt;
&lt;li&gt;Lower infrastructure costs: your costs don’t scale with video length or user count&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;tech-stack&quot;&gt;Tech Stack&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Frontend:&lt;/strong&gt; Next.js + React&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Video Engine:&lt;/strong&gt; CE.SDK (CreativeEngine)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Transcription:&lt;/strong&gt; ElevenLabs Scribe v2&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI Analysis:&lt;/strong&gt; Google Gemini&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;architecture-overview&quot;&gt;Architecture Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;High-Level Flow&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;AI Video Shortener&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 818px) 818px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;818&quot; height=&quot;841&quot; src=&quot;https://img.ly/_astro/ai-video-processing-build-a-video-shortener_u34XT.webp&quot; srcset=&quot;/_astro/ai-video-processing-build-a-video-shortener_Z25jXWD.webp 640w, /_astro/ai-video-processing-build-a-video-shortener_2uzdyU.webp 750w, /_astro/ai-video-processing-build-a-video-shortener_u34XT.webp 818w&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Required API Keys&lt;/strong&gt;&lt;/p&gt;

























&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Service&lt;/th&gt;&lt;th&gt;Purpose&lt;/th&gt;&lt;th&gt;Environment Variable&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;CE.SDK&lt;/td&gt;&lt;td&gt;Video editing engine&lt;/td&gt;&lt;td&gt;&lt;code&gt;NEXT_PUBLIC_CESDK_LICENSE&lt;/code&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;ElevenLabs&lt;/td&gt;&lt;td&gt;Speech-to-text transcription&lt;/td&gt;&lt;td&gt;&lt;code&gt;ELEVENLABS_API_KEY&lt;/code&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Gemini (via OpenRouter or direct)&lt;/td&gt;&lt;td&gt;AI highlight detection&lt;/td&gt;&lt;td&gt;&lt;code&gt;OPENROUTER_API_KEY&lt;/code&gt; or &lt;code&gt;GEMINI_API_KEY&lt;/code&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;h2 id=&quot;setting-up-cesdk&quot;&gt;Setting Up CE.SDK&lt;/h2&gt;
&lt;h3 id=&quot;what-is-cesdk&quot;&gt;What is CE.SDK?&lt;/h3&gt;
&lt;p&gt;CE.SDK (CreativeEngine SDK) is a browser-based engine for video, image, and design editing: a programmable video editor you can embed in your app.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Key Concepts:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Engine:&lt;/strong&gt; The runtime that manages the editing session&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scene:&lt;/strong&gt; The document/project containing all elements&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Blocks:&lt;/strong&gt; Individual elements (video clips, text, shapes, audio)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Timeline:&lt;/strong&gt; Time-based arrangement of blocks for video editing&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;installation&quot;&gt;Installation&lt;/h3&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;npm&lt;/span&gt;&lt;span&gt; install&lt;/span&gt;&lt;span&gt; @cesdk/cesdk-js&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;initializing-the-creativeengine&quot;&gt;Initializing the CreativeEngine&lt;/h3&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; CreativeEngine &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;@cesdk/cesdk-js&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; engine&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; await&lt;/span&gt;&lt;span&gt; CreativeEngine.&lt;/span&gt;&lt;span&gt;init&lt;/span&gt;&lt;span&gt;({&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  license: process.env.&lt;/span&gt;&lt;span&gt;NEXT_PUBLIC_CESDK_LICENSE&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;});&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Create a video scene&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; scene&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.scene.&lt;/span&gt;&lt;span&gt;createVideo&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Get the page (timeline container)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; pages&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.scene.&lt;/span&gt;&lt;span&gt;getPages&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; page&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; pages[&lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;];&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Configure page dimensions for your target aspect ratio&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setWidth&lt;/span&gt;&lt;span&gt;(page, &lt;/span&gt;&lt;span&gt;1080&lt;/span&gt;&lt;span&gt;); &lt;/span&gt;&lt;span&gt;// 9:16 vertical&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setHeight&lt;/span&gt;&lt;span&gt;(page, &lt;/span&gt;&lt;span&gt;1920&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;uploading-video-to-cesdk&quot;&gt;Uploading Video to CE.SDK&lt;/h3&gt;
&lt;p&gt;CE.SDK works with video through a fill-based system. The graphic block is the container, while the video fill holds the actual media source and playback properties.&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Create a video block&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; videoBlock&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;create&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;graphic&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; videoFill&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;createFill&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;video&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Set the video source&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setString&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  videoFill,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &apos;fill/video/fileURI&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  videoUrl &lt;/span&gt;&lt;span&gt;// Can be a blob URL or remote URL&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Apply fill to block&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setFill&lt;/span&gt;&lt;span&gt;(videoBlock, videoFill);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Add to timeline&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;appendChild&lt;/span&gt;&lt;span&gt;(page, videoBlock);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;extracting-audio-for-transcription&quot;&gt;Extracting Audio for Transcription&lt;/h3&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Configure audio-only export&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; mimeType&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; &apos;audio/mp4&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Export just the audio track&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; audioBlob&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; await&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;export&lt;/span&gt;&lt;span&gt;(page, mimeType, {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  targetWidth: &lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  targetHeight: &lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;});&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// audioBlob can now be sent to transcription API&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Setting both dimensions to 0 tells CE.SDK to skip video encoding entirely, making this export much faster than exporting the full video.&lt;/p&gt;
&lt;h3 id=&quot;getting-video-metadata&quot;&gt;Getting Video Metadata&lt;/h3&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Get video duration&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; duration&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;getDuration&lt;/span&gt;&lt;span&gt;(videoBlock);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Get dimensions from the fill&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; videoFill&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;getFill&lt;/span&gt;&lt;span&gt;(videoBlock);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; sourceWidth&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;getSourceWidth&lt;/span&gt;&lt;span&gt;(videoFill);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; sourceHeight&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;getSourceHeight&lt;/span&gt;&lt;span&gt;(videoFill);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;console.&lt;/span&gt;&lt;span&gt;log&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;`Video: ${&lt;/span&gt;&lt;span&gt;sourceWidth&lt;/span&gt;&lt;span&gt;}x${&lt;/span&gt;&lt;span&gt;sourceHeight&lt;/span&gt;&lt;span&gt;}, ${&lt;/span&gt;&lt;span&gt;duration&lt;/span&gt;&lt;span&gt;}s`&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;ai-powered-transcription--highlight-detection&quot;&gt;AI-Powered Transcription &amp;#x26; Highlight Detection&lt;/h2&gt;
&lt;h3 id=&quot;the-pipeline&quot;&gt;The Pipeline&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Audio → Transcription:&lt;/strong&gt; Send extracted audio to ElevenLabs Scribe&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Transcription → Analysis:&lt;/strong&gt; Feed word-level transcript to Gemini&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Analysis → Timestamps:&lt;/strong&gt; Map AI suggestions back to precise video times&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id=&quot;transcription-with-speaker-diarization&quot;&gt;Transcription with Speaker Diarization&lt;/h3&gt;
&lt;p&gt;ElevenLabs Scribe v2 provides:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Word-level timestamps (start/end time for each word)&lt;/li&gt;
&lt;li&gt;Speaker diarization (which speaker said what)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The output is a structured transcript where each word has a precise timestamp, enabling frame-accurate editing.&lt;/p&gt;
&lt;h3 id=&quot;ai-highlight-detection-with-gemini&quot;&gt;AI Highlight Detection with Gemini&lt;/h3&gt;
&lt;p&gt;The prompt structure matters. Here’s what works:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;You are analyzing a video transcript to identify segments for short-form content.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;TRANSCRIPT:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;[Word-by-word transcript with timestamps]&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;TASK:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Identify 3-4 segments that work as standalone short videos. For each segment:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;1. Find the exact starting and ending words&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;2. Ensure clean sentence boundaries (no mid-sentence cuts)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;3. Aim for 30-60 second segments&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;OUTPUT FORMAT (JSON):&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;{&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &quot;concepts&quot;: [&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      &quot;id&quot;: &quot;concept_1&quot;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      &quot;title&quot;: &quot;Hook title&quot;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      &quot;description&quot;: &quot;Why this segment works as a standalone clip&quot;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      &quot;trimmed_text&quot;: &quot;The exact transcript text to keep...&quot;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      &quot;estimated_duration_seconds&quot;: 45&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    }&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  ]&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;}&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;CRITERIA FOR SELECTION:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;- Strong hooks (surprising statements, questions, bold claims)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;- Complete thoughts (don&apos;t cut mid-explanation)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;- Emotional peaks (humor, insight, controversy)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;- Standalone value (makes sense without context)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Before finalizing each segment, ask: &quot;If someone started watching here,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;would they understand what&apos;s being discussed?&quot;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;mapping-back-to-timestamps&quot;&gt;Mapping Back to Timestamps&lt;/h3&gt;
&lt;p&gt;Once Gemini returns the &lt;code&gt;trimmed_text&lt;/code&gt;, we match it against our word-level transcript to find exact timestamps:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;AI returns:     &quot;The secret to success is actually quite simple...&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Transcript has: [{ word: &quot;The&quot;, start: 45.2 }, { word: &quot;secret&quot;, start: 45.4 }, ...]&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Result:         Trim video from 45.2s to 52.8s&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This text-matching approach is more reliable than asking the AI to output timestamps directly. LLMs can hallucinate timestamps or miscalculate offsets, but they’re excellent at identifying the right words, so we let the transcript data provide the ground truth for timing.&lt;/p&gt;
&lt;h2 id=&quot;working-with-the-cesdk-timeline&quot;&gt;Working with the CE.SDK Timeline&lt;/h2&gt;
&lt;h3 id=&quot;understanding-blocks&quot;&gt;Understanding Blocks&lt;/h3&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Video/Image content&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; graphic&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;create&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;graphic&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Audio track&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; audio&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;create&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;audio&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Text overlay&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; text&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;create&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;text&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Each block can be positioned on the timeline&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setTimeOffset&lt;/span&gt;&lt;span&gt;(block, startTimeInSeconds);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setDuration&lt;/span&gt;&lt;span&gt;(block, durationInSeconds);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;manipulating-trim-points&quot;&gt;Manipulating Trim Points&lt;/h3&gt;
&lt;p&gt;Trimming controls which portion of the source media is shown:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; videoFill&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;getFill&lt;/span&gt;&lt;span&gt;(videoBlock);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Set where in the source video to start (in seconds)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setTrimOffset&lt;/span&gt;&lt;span&gt;(videoFill, &lt;/span&gt;&lt;span&gt;45.2&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Set how long to play from that point&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setTrimLength&lt;/span&gt;&lt;span&gt;(videoFill, &lt;/span&gt;&lt;span&gt;30.0&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Also update the block&apos;s duration to match&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setDuration&lt;/span&gt;&lt;span&gt;(videoBlock, &lt;/span&gt;&lt;span&gt;30.0&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;working-with-fills-and-their-timing&quot;&gt;Working with Fills and Their Timing&lt;/h3&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Get the fill (contains the actual media)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; fill&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;getFill&lt;/span&gt;&lt;span&gt;(block);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Fills have their own timing properties&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; trimStart&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;getTrimOffset&lt;/span&gt;&lt;span&gt;(fill);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; trimDuration&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;getTrimLength&lt;/span&gt;&lt;span&gt;(fill);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// The block&apos;s duration should typically match the fill&apos;s trim length&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setDuration&lt;/span&gt;&lt;span&gt;(block, trimDuration);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Think of the fill as the media source (which part of the original video to use) and the block as the timeline placement (when and how long it appears). Both need to be updated together for clean edits.&lt;/p&gt;
&lt;h3 id=&quot;creating-time-based-edits-from-transcript-words&quot;&gt;Creating Time-Based Edits from Transcript Words&lt;/h3&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;typescript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;interface&lt;/span&gt;&lt;span&gt; TranscriptWord&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  word&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; string&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  start&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; number&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  end&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; number&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  speaker_id&lt;/span&gt;&lt;span&gt;?:&lt;/span&gt;&lt;span&gt; string&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;}&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;function&lt;/span&gt;&lt;span&gt; applyTranscriptTrim&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  engine&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; CreativeEngine&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  videoBlock&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; number&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  words&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; TranscriptWord&lt;/span&gt;&lt;span&gt;[]&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;) {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  if&lt;/span&gt;&lt;span&gt; (words.&lt;/span&gt;&lt;span&gt;length&lt;/span&gt;&lt;span&gt; ===&lt;/span&gt;&lt;span&gt; 0&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;return&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  const&lt;/span&gt;&lt;span&gt; startTime&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; words[&lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;].start;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  const&lt;/span&gt;&lt;span&gt; endTime&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; words[words.&lt;/span&gt;&lt;span&gt;length&lt;/span&gt;&lt;span&gt; -&lt;/span&gt;&lt;span&gt; 1&lt;/span&gt;&lt;span&gt;].end;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  const&lt;/span&gt;&lt;span&gt; duration&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; endTime &lt;/span&gt;&lt;span&gt;-&lt;/span&gt;&lt;span&gt; startTime;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  const&lt;/span&gt;&lt;span&gt; fill&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;getFill&lt;/span&gt;&lt;span&gt;(videoBlock);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  engine.block.&lt;/span&gt;&lt;span&gt;setTrimOffset&lt;/span&gt;&lt;span&gt;(fill, startTime);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  engine.block.&lt;/span&gt;&lt;span&gt;setTrimLength&lt;/span&gt;&lt;span&gt;(fill, duration);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  engine.block.&lt;/span&gt;&lt;span&gt;setDuration&lt;/span&gt;&lt;span&gt;(videoBlock, duration);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;}&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;generating-speaker-thumbnails&quot;&gt;Generating Speaker Thumbnails&lt;/h3&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;typescript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;async&lt;/span&gt;&lt;span&gt; function&lt;/span&gt;&lt;span&gt; generateSpeakerThumbnail&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  engine&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; CreativeEngine&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  videoBlock&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; number&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  timestampSeconds&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; number&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  size&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; number&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; 128&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;)&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; Promise&lt;/span&gt;&lt;span&gt;&amp;#x3C;&lt;/span&gt;&lt;span&gt;string&lt;/span&gt;&lt;span&gt;&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  const&lt;/span&gt;&lt;span&gt; fill&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;getFill&lt;/span&gt;&lt;span&gt;(videoBlock);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  // Seek to the specific timestamp&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  engine.block.&lt;/span&gt;&lt;span&gt;setTrimOffset&lt;/span&gt;&lt;span&gt;(fill, timestampSeconds);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  engine.block.&lt;/span&gt;&lt;span&gt;setTrimLength&lt;/span&gt;&lt;span&gt;(fill, &lt;/span&gt;&lt;span&gt;0.1&lt;/span&gt;&lt;span&gt;); &lt;/span&gt;&lt;span&gt;// Just a single frame&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  // Export as image&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  const&lt;/span&gt;&lt;span&gt; blob&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; await&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;export&lt;/span&gt;&lt;span&gt;(videoBlock, &lt;/span&gt;&lt;span&gt;&apos;image/jpeg&apos;&lt;/span&gt;&lt;span&gt;, {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    targetWidth: size,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    targetHeight: size,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  });&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  return&lt;/span&gt;&lt;span&gt; URL&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;createObjectURL&lt;/span&gt;&lt;span&gt;(blob);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;}&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;We sample multiple timestamps throughout each speaker’s talk time to show different facial angles and expressions. This helps users identify the right person even if they’re looking away or mid-gesture in one frame.&lt;/p&gt;
&lt;h2 id=&quot;speaker-detection--face-tracking&quot;&gt;Speaker Detection &amp;#x26; Face Tracking&lt;/h2&gt;
&lt;h3 id=&quot;why-semi-automatic&quot;&gt;Why Semi-Automatic?&lt;/h3&gt;
&lt;p&gt;Fully automatic speaker detection fails often enough that we added a confirmation step. Users verify detected faces against speaker names from the transcript. It takes a few seconds and prevents bad crops on the entire video.&lt;/p&gt;
&lt;h3 id=&quot;how-it-works&quot;&gt;How It Works&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;Sample frames throughout the video&lt;/li&gt;
&lt;li&gt;Detect &amp;#x26; cluster faces using face-api.js (runs in browser, no server needed)&lt;/li&gt;
&lt;li&gt;User confirms speaker identities via thumbnails&lt;/li&gt;
&lt;li&gt;Correlate with transcript diarization to map speakers → face locations&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This gives us verified speaker-to-face mapping for dynamic cropping and picture-in-picture layouts.&lt;/p&gt;
&lt;h2 id=&quot;multi-speaker-templates--dynamic-switching&quot;&gt;Multi-Speaker Templates &amp;#x26; Dynamic Switching&lt;/h2&gt;
&lt;h3 id=&quot;the-concept&quot;&gt;The Concept&lt;/h3&gt;
&lt;p&gt;When a video has multiple speakers, we can create layouts that show:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The active speaker prominently&lt;/li&gt;
&lt;li&gt;Other speakers in smaller picture-in-picture views&lt;/li&gt;
&lt;li&gt;Dynamic switching as the conversation flows&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;creating-picture-in-picture-with-cesdk&quot;&gt;Creating Picture-in-Picture with CE.SDK&lt;/h3&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Duplicate the video block for each speaker slot&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; pipBlock&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;duplicate&lt;/span&gt;&lt;span&gt;(originalVideoBlock);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Position and size the PiP&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setWidth&lt;/span&gt;&lt;span&gt;(pipBlock, &lt;/span&gt;&lt;span&gt;200&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setHeight&lt;/span&gt;&lt;span&gt;(pipBlock, &lt;/span&gt;&lt;span&gt;200&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setPositionX&lt;/span&gt;&lt;span&gt;(pipBlock, &lt;/span&gt;&lt;span&gt;20&lt;/span&gt;&lt;span&gt;); &lt;/span&gt;&lt;span&gt;// 20px from left&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setPositionY&lt;/span&gt;&lt;span&gt;(pipBlock, &lt;/span&gt;&lt;span&gt;20&lt;/span&gt;&lt;span&gt;); &lt;/span&gt;&lt;span&gt;// 20px from top&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Enable cropping&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setClipped&lt;/span&gt;&lt;span&gt;(pipBlock, &lt;/span&gt;&lt;span&gt;true&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setContentFillMode&lt;/span&gt;&lt;span&gt;(pipBlock, &lt;/span&gt;&lt;span&gt;&apos;Cover&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;key-technique-muting-duplicate-audio&quot;&gt;Key Technique: Muting Duplicate Audio&lt;/h3&gt;
&lt;p&gt;When duplicating video blocks for multi-speaker layouts, each copy has its own audio track. We must mute all but one. The &lt;code&gt;setMuted&lt;/code&gt; API operates on the video fill, not the block itself:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// For each speaker slot after the first, mute the video fill&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;if&lt;/span&gt;&lt;span&gt; (slotIndex &lt;/span&gt;&lt;span&gt;&gt;&lt;/span&gt;&lt;span&gt; 0&lt;/span&gt;&lt;span&gt;) {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  const&lt;/span&gt;&lt;span&gt; videoFill&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;getFill&lt;/span&gt;&lt;span&gt;(duplicatedBlock);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  if&lt;/span&gt;&lt;span&gt; (videoFill) {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    engine.block.&lt;/span&gt;&lt;span&gt;setMuted&lt;/span&gt;&lt;span&gt;(videoFill, &lt;/span&gt;&lt;span&gt;true&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  }&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;}&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;dynamic-speaker-switching&quot;&gt;Dynamic Speaker Switching&lt;/h3&gt;
&lt;p&gt;As the active speaker changes throughout the video, we:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Detect which speaker is talking (from transcript diarization)&lt;/li&gt;
&lt;li&gt;Swap speaker positions in the template&lt;/li&gt;
&lt;li&gt;Keep the active speaker in the prominent position&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The layout updates automatically as the conversation switches between speakers. We apply different trim offsets to each duplicated block based on the transcript timing, so the main speaker slot shows the person currently talking while PiP slots show the listeners.&lt;/p&gt;
&lt;h2 id=&quot;preview-playback--export&quot;&gt;Preview, Playback &amp;#x26; Export&lt;/h2&gt;
&lt;h3 id=&quot;setting-up-the-canvas&quot;&gt;Setting Up the Canvas&lt;/h3&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; container&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; document.&lt;/span&gt;&lt;span&gt;getElementById&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;cesdk-canvas&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.element.&lt;/span&gt;&lt;span&gt;attachTo&lt;/span&gt;&lt;span&gt;(container);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;playback-controls&quot;&gt;Playback Controls&lt;/h3&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.player.&lt;/span&gt;&lt;span&gt;play&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.player.&lt;/span&gt;&lt;span&gt;pause&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.player.&lt;/span&gt;&lt;span&gt;setPlaybackTime&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;30.5&lt;/span&gt;&lt;span&gt;); &lt;/span&gt;&lt;span&gt;// seek to 30.5 seconds&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; currentTime&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.player.&lt;/span&gt;&lt;span&gt;getPlaybackTime&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; isPlaying&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.player.&lt;/span&gt;&lt;span&gt;isPlaying&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;syncing-ui-state&quot;&gt;Syncing UI State&lt;/h3&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.player.&lt;/span&gt;&lt;span&gt;onPlaybackTimeChanged&lt;/span&gt;&lt;span&gt;(() &lt;/span&gt;&lt;span&gt;=&gt;&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  const&lt;/span&gt;&lt;span&gt; time&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.player.&lt;/span&gt;&lt;span&gt;getPlaybackTime&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  updateTimeDisplay&lt;/span&gt;&lt;span&gt;(time);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  updateProgressBar&lt;/span&gt;&lt;span&gt;(time &lt;/span&gt;&lt;span&gt;/&lt;/span&gt;&lt;span&gt; totalDuration);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;});&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.player.&lt;/span&gt;&lt;span&gt;onPlaybackStateChanged&lt;/span&gt;&lt;span&gt;(() &lt;/span&gt;&lt;span&gt;=&gt;&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  updatePlayButton&lt;/span&gt;&lt;span&gt;(engine.player.&lt;/span&gt;&lt;span&gt;isPlaying&lt;/span&gt;&lt;span&gt;());&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;});&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;export-options&quot;&gt;Export Options&lt;/h3&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; exportOptions&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  targetWidth: &lt;/span&gt;&lt;span&gt;1080&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  targetHeight: &lt;/span&gt;&lt;span&gt;1920&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  framerate: &lt;/span&gt;&lt;span&gt;30&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  videoBitrate: &lt;/span&gt;&lt;span&gt;8_000_000&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;// 8 Mbps&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;};&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; blob&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; await&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;export&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  page,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &apos;video/mp4&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  exportOptions,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  (&lt;/span&gt;&lt;span&gt;progress&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;=&gt;&lt;/span&gt;&lt;span&gt; updateProgressBar&lt;/span&gt;&lt;span&gt;(progress &lt;/span&gt;&lt;span&gt;*&lt;/span&gt;&lt;span&gt; 100&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Trigger download&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; url&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; URL&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;createObjectURL&lt;/span&gt;&lt;span&gt;(blob);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; a&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; document.&lt;/span&gt;&lt;span&gt;createElement&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;a&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;a.href &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; url;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;a.download &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; &apos;shortened-video.mp4&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;a.&lt;/span&gt;&lt;span&gt;click&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;For longer videos, consider showing estimated time remaining or allowing background export. Browser export is single-threaded and blocks the tab. A 5-minute export of a 60-second clip isn’t unusual on average hardware, so user feedback is critical.&lt;/p&gt;
&lt;h2 id=&quot;the-finished-app&quot;&gt;The Finished App&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;The user flow:&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Upload&lt;/strong&gt; → Drop a long-form video into the browser&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Configure&lt;/strong&gt; → Pick output mode (highlights/summary/cleanup) and aspect ratio (9:16, 16:9, 1:1)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Verify speakers&lt;/strong&gt; → Match detected faces to transcript speaker names&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Review clips&lt;/strong&gt; → Browse the 3-4 AI-suggested segments, adjust if needed&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Choose template&lt;/strong&gt; → Solo speaker, sidecar, stacked, etc.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Preview&lt;/strong&gt; → Scrub through the timeline, see exactly what you’ll get&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Export&lt;/strong&gt; → Download the final video directly from the browser&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&quot;whats-next&quot;&gt;What’s Next&lt;/h2&gt;
&lt;h3 id=&quot;ideas-for-extension&quot;&gt;Ideas for Extension&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Caption style controls:&lt;/strong&gt; Custom fonts, animations, and positioning for subtitles&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;B-roll insertion:&lt;/strong&gt; Automatically add relevant stock footage&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Music &amp;#x26; sound effects:&lt;/strong&gt; AI-selected background audio&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Brand templates:&lt;/strong&gt; Custom overlays, intros, outros&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Batch processing:&lt;/strong&gt; Process multiple videos in sequence&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;taking-it-server-side&quot;&gt;Taking It Server-Side&lt;/h3&gt;
&lt;p&gt;Client-side processing for large files strains browser memory, and users must keep the tab open during export. A hybrid approach works better for production: upload in the background while users edit, then render on a server. You can also offload just the export step: let users build their edits in the browser, then send the CE.SDK scene JSON to your backend for faster, background rendering.&lt;/p&gt;
&lt;p&gt;CE.SDK runs server-side with the same API. For batch processing, background jobs, or offloading rendering from user devices, see the &lt;a href=&quot;https://img.ly/docs/cesdk/renderer/cesdk-renderer-overview-7f3e9a/&quot;&gt;CE.SDK Renderer&lt;/a&gt; for creative automation.&lt;/p&gt;
&lt;h2 id=&quot;resources&quot;&gt;Resources&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://img.ly/docs/cesdk/&quot;&gt;CE.SDK Documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://img.ly/docs/cesdk/node/create-video-c41a08/&quot;&gt;CE.SDK Video Editing Guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/imgly/videoclipper&quot;&gt;GitHub: Video Shortener Source&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://elevenlabs.io/docs/api-reference&quot;&gt;ElevenLabs API Docs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://ai.google.dev/gemini-api/docs&quot;&gt;Gemini API Docs&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;p&gt;Made by &lt;a href=&quot;https://img.ly/&quot;&gt;IMG.LY&lt;/a&gt; with &lt;a href=&quot;https://img.ly/products/creative-sdk/&quot;&gt;CE.SDK&lt;/a&gt;&lt;/p&gt;</content:encoded><dc:creator>Eray</dc:creator><media:content url="https://blog.img.ly/2026/01/Gemini_Generated_Image_qj8gi7qj8gi7qj8g.png" medium="image"/><category>AI</category><category>Insights</category><category>Video Editing</category></item><item><title>From Prompt to Editor: Running CE.SDK Inside ChatGPT with the Apps SDK</title><link>https://img.ly/blog/from-prompt-to-editor-running-ce-sdk-inside-chatgpt-with-the-apps-sdk/</link><guid isPermaLink="true">https://img.ly/blog/from-prompt-to-editor-running-ce-sdk-inside-chatgpt-with-the-apps-sdk/</guid><description>We built a technical demo running CE.SDK directly inside ChatGPT using MCP, showing how AI chats can open real, interactive editors. It highlights strict MCP contracts, visual-first UX, and how chat becomes a coordination layer for doing creative work.</description><pubDate>Fri, 19 Dec 2025 15:43:16 GMT</pubDate><content:encoded>&lt;p&gt;With the new ChatGPT Apps SDK and Model Context Protocol (MCP), chat interfaces are starting to look less like Q&amp;#x26;A tools and more like places where work actually happens. To explore what that means for creative workflows, we built a small technical demo: &lt;a href=&quot;https://github.com/imgly/cesdk-web-examples/tree/main/cookbooks-chatgpt-app&quot;&gt;&lt;strong&gt;CE.SDK running directly inside ChatGPT&lt;/strong&gt;.&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;From a user’s perspective, the flow is almost trivial. You ask ChatGPT for something like an ecommerce template. ChatGPT searches our template catalog, selects a matching design, and opens it instantly in a fully interactive CE.SDK editor, right inside the chat interface. What looks like a preview is, in fact, a real editor loaded with a real template scene.&lt;/p&gt;
&lt;p&gt;This isn’t meant as a product announcement. It’s a technical proof of concept showing how creative SDKs can plug directly into AI-native interfaces.&lt;/p&gt;
&lt;h2 id=&quot;cesdk-as-a-chatgpt-app&quot;&gt;CE.SDK as a ChatGPT App&lt;/h2&gt;
&lt;figure class=&quot;kg-embed-card&quot;&gt;
&lt;iframe src=&quot;https://www.youtube.com/embed/AoDqUNVLLJ4?feature=oembed&quot; title=&quot;ChatGPT Opens a Real Design Editor: CE.SDK Inside MCP App (Tech Demo)&quot; loading=&quot;lazy&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; referrerpolicy=&quot;strict-origin-when-cross-origin&quot; allowfullscreen&gt;&lt;/iframe&gt;
&lt;/figure&gt;
&lt;p&gt;The integration is built around a custom MCP server that exposes CE.SDK to ChatGPT as a tool. The server speaks OpenAI’s JSON-RPC–style MCP and implements the standard lifecycle methods (&lt;code&gt;initialize&lt;/code&gt;, &lt;code&gt;tools.list&lt;/code&gt;, &lt;code&gt;tools.call&lt;/code&gt;, &lt;code&gt;resources.read&lt;/code&gt;). It knows about our premium template catalog and emits structured payloads that the frontend understands.&lt;/p&gt;
&lt;p&gt;On the client side, a Next.js app listens to tool output events streamed from ChatGPT, renders CE.SDK widgets, and hydrates them with the payloads returned by the tool, such as a scene URL, placeholder values, or export permissions. Templates are loaded via CE.SDK’s Template API, either from a URL or from a serialized scene string.&lt;/p&gt;
&lt;p&gt;Under the hood, the stack is fairly conventional:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Next.js 15 (App Router)&lt;/li&gt;
&lt;li&gt;CE.SDK Web / CreativeEngine&lt;/li&gt;
&lt;li&gt;A custom MCP handler to normalize JSON-RPC&lt;/li&gt;
&lt;li&gt;Vercel for hosting&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;What’s new is not the technology itself, but the context in which it runs.&lt;/p&gt;
&lt;h2 id=&quot;working-with-mcp-in-practice&quot;&gt;Working with MCP in Practice&lt;/h2&gt;
&lt;p&gt;The hardest part of the demo wasn’t CE.SDK. It was MCP.&lt;/p&gt;
&lt;p&gt;OpenAI’s MCP implementation is extremely strict. Even the smallest schema mismatch can trigger the infamous “TaskGroup 424” error, usually without any hint as to what went wrong. In many cases, the HTTP response is technically successful, but the JSON structure doesn’t match the expected schema closely enough.&lt;/p&gt;
&lt;p&gt;The key lesson here is to treat MCP responses as hard contracts:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Validate every response against a schema (for example with zod).&lt;/li&gt;
&lt;li&gt;Mirror OpenAI’s field names exactly, even for empty or optional capabilities.&lt;/li&gt;
&lt;li&gt;Assume that a 424 almost always means “your JSON shape is wrong”.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Another important insight was how critical visual context is in chat-based tools. If your MCP responses don’t include thumbnails or preview images, ChatGPT will often fall back to rendering links. For creative tools, that immediately breaks the experience. In a chat UI, visuals aren’t an enhancement. They &lt;em&gt;are&lt;/em&gt; the interface.&lt;/p&gt;
&lt;p&gt;State handling also requires a shift in mindset. ChatGPT can replay tool calls, and each prompt effectively creates a new widget instance. You can’t rely on mutating an existing editor. The frontend needs to be idempotent: load scenes from serialized state first, then apply changes. Every tool call should be treated as a fresh render.&lt;/p&gt;
&lt;h2 id=&quot;why-this-pattern-matters&quot;&gt;Why This Pattern Matters&lt;/h2&gt;
&lt;p&gt;This demo points to a broader change in how creative software may be accessed. Chat becomes a coordination layer, not just a conversational one. Instead of explaining how something could be designed, the AI opens the actual editor and lets the user continue from there.&lt;/p&gt;
&lt;p&gt;For CE.SDK, this fits naturally. Editors become embeddable capabilities rather than standalone applications, and AI systems become the entry point into creative workflows. Prompting turns into doing.&lt;/p&gt;
&lt;h2 id=&quot;beyond-openai-the-mcp-ui-standard&quot;&gt;Beyond OpenAI: The MCP UI Standard&lt;/h2&gt;
&lt;p&gt;Although this demo uses OpenAI’s MCP, the architecture maps cleanly to the new MCP UI standard recently introduced by Anthropic. That standard aims to make tool definitions and UI rendering more consistent across models and platforms.&lt;/p&gt;
&lt;p&gt;Because this integration already separates tool logic from UI rendering and relies on structured, explicit payloads, transferring it to Anthropic’s MCP UI model is conceptually straightforward. CE.SDK can act as a reusable creative surface across ChatGPT, Claude, and future AI app ecosystems.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://blog.modelcontextprotocol.io/posts/2025-11-21-mcp-apps/&quot;&gt;You can read more about Anthropic’s MCP UI direction here.&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;This demo is intentionally small and technical, but it highlights a meaningful shift: AI systems that don’t just describe creative outcomes, but open the tools to actually create them.&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><dc:creator>Sven</dc:creator><media:content url="https://blog.img.ly/2025/12/chatgpt-open-design-editor-tempalte.jpg" medium="image"/><category>AI</category><category>MCP</category><category>CE.SDK</category></item><item><title>Animate Between Images - AI-Native Video Workflows with CE.SDK and Veo 3</title><link>https://img.ly/blog/animate-between-images-ai-native-video-workflows-with-ce-sdk-and-veo-3/</link><guid isPermaLink="true">https://img.ly/blog/animate-between-images-ai-native-video-workflows-with-ce-sdk-and-veo-3/</guid><description>See how we integrated Veo 3.1 into CE.SDK to animate between images in seconds. With generation times as low as 9 s, this demo shows how easily you can embed AI-native workflows from stills to smooth video clips directly inside your Creative Editor.</description><pubDate>Wed, 22 Oct 2025 14:10:31 GMT</pubDate><content:encoded>&lt;p&gt;With the release of &lt;a href=&quot;https://aistudio.google.com/models/veo-3&quot;&gt;&lt;strong&gt;Veo 3.1&lt;/strong&gt;,&lt;/a&gt; we wanted to show just how effortless it is to embed generative AI capabilities directly into creative workflows. In this quick demo, we integrated Veo 3 into &lt;strong&gt;CreativeEditor SDK (CE.SDK)&lt;/strong&gt; enabling users to &lt;strong&gt;animate between two still images&lt;/strong&gt; in just a few clicks.&lt;/p&gt;
&lt;p&gt;&lt;video src=&quot;https://storage.googleapis.com/imgly-static-assets/static/blog/videos/ai-launch/IMG-LY-CE-SDK-veo-3-1-video-design-editor-sdk.mp4&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;h3 id=&quot;from-still-images-to-motion&quot;&gt;From Still Images to Motion&lt;/h3&gt;
&lt;p&gt;In our demo, we start with two images of the same person, one wearing a hat, the other without. Inside the editor, users simply select both images, click the &lt;strong&gt;AI context button&lt;/strong&gt;, and choose &lt;strong&gt;“Animate between images.”&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The images are loaded into a side panel where users can optionally add a &lt;strong&gt;text prompt&lt;/strong&gt; to guide the transition. Once generated, the resulting short video is placed directly on the canvas ready for editing, compositing, or export.&lt;/p&gt;
&lt;p&gt;What’s particularly impressive is the &lt;strong&gt;generation speed&lt;/strong&gt;. In this example, Veo 3.1 produced a smooth &lt;strong&gt;8-second transition in just 9 seconds&lt;/strong&gt; a major improvement compared to earlier versions. This speed makes iterative creative workflows feel fluid and responsive, bridging the gap between prompt-driven generation and real-time editing.&lt;/p&gt;
&lt;h3 id=&quot;try-it-out&quot;&gt;Try It Out&lt;/h3&gt;
&lt;p&gt;&lt;img alt=&quot;Try out Veo 3.1&amp;#39;s magic in CE.SDK.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1312px) 1312px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1312&quot; height=&quot;310&quot; src=&quot;https://img.ly/_astro/veo-fal-ai-video-editor-sdk-ce_sdk-imgly_sD9cv.webp&quot; srcset=&quot;/_astro/veo-fal-ai-video-editor-sdk-ce_sdk-imgly_fSI3j.webp 640w, /_astro/veo-fal-ai-video-editor-sdk-ce_sdk-imgly_Z2fext8.webp 750w, /_astro/veo-fal-ai-video-editor-sdk-ce_sdk-imgly_Z1SRm7j.webp 828w, /_astro/veo-fal-ai-video-editor-sdk-ce_sdk-imgly_Hlv7V.webp 1080w, /_astro/veo-fal-ai-video-editor-sdk-ce_sdk-imgly_1Vk21l.webp 1280w, /_astro/veo-fal-ai-video-editor-sdk-ce_sdk-imgly_sD9cv.webp 1312w&quot;&gt;&lt;/p&gt;
&lt;p&gt;You can check out the implementation on &lt;a href=&quot;https://github.com/imgly/plugins/blob/main/examples/ai/src/pages/Veo31Example.tsx&quot;&gt;GitHub&lt;/a&gt; and give Veo 3.1 a spin inside your CE.SDK instance (sign up for a &lt;a href=&quot;https://img.ly/docs/cesdk/&quot;&gt;free trial&lt;/a&gt; if you haven’t already).&lt;/p&gt;
&lt;h3 id=&quot;a-glimpse-of-ai-native-editing&quot;&gt;A Glimpse of AI-Native Editing&lt;/h3&gt;
&lt;p&gt;This simple feature highlights how easy it is to make your editor &lt;strong&gt;AI-native,&lt;/strong&gt; combining traditional editing tools with generative intelligence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Practical use cases include:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;E-commerce&lt;/strong&gt;&lt;br&gt;
Show products “in action” or animate between styles and configurations.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Marketing&lt;/strong&gt;&lt;br&gt;
Create quick product reveal animations from static assets.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Content creation&lt;/strong&gt;&lt;br&gt;
Generate short motion clips or “tween” between creative scenes.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And if users want longer clips, they can simply &lt;strong&gt;add another 8-second track&lt;/strong&gt; and &lt;strong&gt;transition seamlessly&lt;/strong&gt; into the next.&lt;/p&gt;
&lt;h3 id=&quot;the-future-of-creative-workflows&quot;&gt;The Future of Creative Workflows&lt;/h3&gt;
&lt;p&gt;With Veo 3 integrated, CE.SDK becomes a powerful playground for AI-driven creativity from image-to-video to scene interpolation and contextual animation.&lt;/p&gt;
&lt;p&gt;We already empowered over 600 innovative startups, government entities, and Fortune 500 companies to add powerful design, video, and photo editing workflows to their products. &lt;a href=&quot;https://img.ly/forms/contact-sales/&quot;&gt;Get in touch&lt;/a&gt;, to see how we can do the same for you.&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><dc:creator>Marcin</dc:creator><media:content url="https://blog.img.ly/2025/10/veo-3_1-creative-editor-sdk-video-editor-imgly-2.jpg" medium="image"/><category>AI</category><category>Video Editing</category><category>Image2Video</category></item><item><title>What is Visual Prompting?</title><link>https://img.ly/blog/what-is-visual-prompting/</link><guid isPermaLink="true">https://img.ly/blog/what-is-visual-prompting/</guid><description>Visual Prompting is a new way to guide AI using visual input instead of just text. By composing layouts with images, annotations, and design cues directly on the canvas, creators can prompt AI more intuitively. Learn how IMG.LY’s CE.SDK brings this paradigm to life.</description><pubDate>Tue, 29 Jul 2025 13:06:11 GMT</pubDate><content:encoded>&lt;h2 id=&quot;a-new-paradigm-for-creative-ai-built-by-imgly&quot;&gt;A New Paradigm for Creative AI, Built by IMG.LY&lt;/h2&gt;
&lt;p&gt;To say it’s trite to refer to the impact of AI in this or that domain as disruptive or groundbreaking would be an understatement. Yet, few areas have been as profoundly affected as the creative process. With just a text prompt, anyone can produce stunning images, remix visual styles, and explore design possibilities at a scale and speed never seen before. AI has inserted itself so quickly into this process that its gone from curious novelty to an essential part of the creator toolchain.&lt;/p&gt;
&lt;p&gt;The more serious adoption we see, however, the more key limitations of today’s AI tooling come into focus: &lt;strong&gt;the prompt itself&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Text alone, for all its expressive power, struggles to capture the essence of visual intent. Most creative work doesn’t begin with a sentence it begins with a sketch, a layout, a mood board, or an arrangement of elements. Visual ideas are shared by pointing, placing, showing.&lt;/p&gt;
&lt;p&gt;At &lt;strong&gt;IMG.LY&lt;/strong&gt;, we have begun to think about better ways to direct AI for visual generation, the term we use is Visual Prompting.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Visual Prompting: the practice of composing a visual scene or layout as input for a generative model.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Instead of describing what you want with paragraphs of text, you show it directly using a canvas of images, text, spatial cues, and annotations. This visual composition then becomes the prompt for the AI to generate new content in return. It’s a more natural, intuitive, and powerful way to collaborate with AI, especially when integrated directly into the creative process.&lt;/p&gt;
&lt;h2 id=&quot;problem-the-chat-disconnect&quot;&gt;Problem: the Chat Disconnect&lt;/h2&gt;
&lt;p&gt;The current generation of AI tools has largely been shaped by language-first interfaces. Whether it’s ChatGPT for writing or Midjourney for image generation, the assumption is the same: the user will type a descriptive prompt, and the AI will generate a result based on it.&lt;/p&gt;
&lt;p&gt;But when it comes to &lt;strong&gt;design&lt;/strong&gt;, this workflow quickly runs into friction. Visual ideas are inherently spatial and non-linear. Trying to express layout, balance, mood, or specific spatial relationships through text can feel like trying to describe a painting over the phone. It’s possible but unnecessarily cumbersome.&lt;/p&gt;
&lt;p&gt;A designer might want to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Indicate that a certain area in the image should be blue.&lt;/li&gt;
&lt;li&gt;Replace a background with a texture sample.&lt;/li&gt;
&lt;li&gt;Position a character precisely in a composition.&lt;/li&gt;
&lt;li&gt;Annotate which parts of a scene to preserve or modify.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;All of these are difficult to express fluently in text. But they’re &lt;strong&gt;effortless&lt;/strong&gt; in a visual interface. The truth is: &lt;strong&gt;an image is worth more than a thousand words when prompting an image.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id=&quot;what-is-visual-prompting&quot;&gt;What Is Visual Prompting?&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Visual Prompting&lt;/strong&gt; is a multimodal approach to generative AI, where the &lt;strong&gt;input to the model is not just text, but a full visual composition&lt;/strong&gt;: images, text, annotations, and layout.&lt;/p&gt;
&lt;p&gt;Rather than prompting AI in isolation, the user builds their intent on a canvas. This might include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Reference images that communicate mood or style.&lt;/li&gt;
&lt;li&gt;Text blocks indicating desired copy or instructions.&lt;/li&gt;
&lt;li&gt;Annotations pointing to specific areas with notes like “make this glow” or “replace this object.”&lt;/li&gt;
&lt;li&gt;Spatial composition: where elements are arranged meaningfully to convey intent.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The visual prompt is then interpreted by a multimodal model such as OpenAI’s &lt;code&gt;gpt-image-1&lt;/code&gt; to generate new visual content that reflects not only the textual description, but also the visual context.&lt;/p&gt;
&lt;h2 id=&quot;how-visual-prompting-works-in-cesdk&quot;&gt;How Visual Prompting Works in CE.SDK&lt;/h2&gt;
&lt;p&gt;About time for an example. As part of our recent AI released &lt;a href=&quot;https://img.ly/demos/ai-editor/web/&quot;&gt;we demoed how to use OpenAIs &lt;code&gt;gpt-image-1&lt;/code&gt; model&lt;/a&gt; to build visual prompting into &lt;strong&gt;CreativeEditor SDK (CE.SDK)&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Here’s what the process looks like inside CE.SDK:&lt;/p&gt;
&lt;p&gt;&lt;video src=&quot;https://storage.googleapis.com/imgly-static-assets/static/blog/videos/imgly-ai/visualprompt_05.mp4&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Compose Visually:&lt;/strong&gt; The user creates a layout with reference content, uploaded images, icons, color schemes, design elements, placeholder text, and annotations. This composition represents the “prompt” in visual form.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Add AI Layers:&lt;/strong&gt; With a single click, the user can trigger image generation using CE.SDK’s built-in AI plugin. The plugin sends the visual context (alongside any optional text input) to a multimodal model capable of interpreting both.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Refine and Iterate:&lt;/strong&gt; Users can adjust the layout, reposition elements, change annotations, or layer in new references, then prompt again. Because the canvas is interactive and editable, the feedback loop is tight.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Build Up Complexity:&lt;/strong&gt; Over time, users can layer generated images with manually designed components or other generated outputs, creating rich compositions that blend AI creativity with human direction.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This workflow turns the traditional prompt/response cycle into a &lt;strong&gt;conversation between the designer and the model&lt;/strong&gt;, with the canvas acting as the shared language.&lt;/p&gt;
&lt;h2 id=&quot;who-is-visual-prompting-for&quot;&gt;Who Is Visual Prompting For?&lt;/h2&gt;
&lt;p&gt;The use cases for Visual Prompting extend across industries:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Creative teams&lt;/strong&gt; can go from reference to generation in seconds, iterating visually instead of wrangling prompts.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Marketing teams&lt;/strong&gt; can generate regionalized or personalized creative variants from a shared layout.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Product designers&lt;/strong&gt; can prototype in context, turning layouts into realistic screens without leaving the editor.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Storytellers and content creators&lt;/strong&gt; can use annotated sketches to generate detailed illustrations or scene variations.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;E-commerce platforms&lt;/strong&gt; can give sellers the power to visually customize their brand materials with AI assistance.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In every case, Visual Prompting replaces friction with flow and text-based prompting with something more expressive, more reliable, and more fun.&lt;/p&gt;
&lt;h2 id=&quot;built-for-this-multimodal-models-and-cesdks-plugin-system&quot;&gt;Built for This: Multimodal Models and CE.SDK’s Plugin System&lt;/h2&gt;
&lt;p&gt;Visual Prompting is only possible because of two parallel advancements:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Multimodal AI models&lt;/strong&gt;, such as OpenAI’s &lt;code&gt;gpt-image-1&lt;/code&gt;, that can interpret both images and text, understand spatial relationships, and respond to annotated cues.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A flexible, composable editor SDK&lt;/strong&gt; like CE.SDK, which enables the construction of visual prompts on a live canvas, and makes it easy to integrate AI models directly into the design flow.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Our SDK was built from the ground up to support &lt;strong&gt;AI-first creative workflows&lt;/strong&gt;. Its plugin architecture allows you to add any model or API, image generation, video generation, captioning, text rewriting and use it natively inside the editor without the need to switch tools or copy/paste.&lt;/p&gt;
&lt;p&gt;Generative AI’s full potential is only unlocked when it is embedded directly into the tools creatives use not siloed in chatbots or separate interfaces. Visual Prompting allows that embedding to go even deeper, aligning the &lt;strong&gt;mode of input (visual)&lt;/strong&gt; with the &lt;strong&gt;desired output (visual)&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id=&quot;explore-it-yourself&quot;&gt;Explore It Yourself&lt;/h3&gt;
&lt;p&gt;🎨 Try out Visual Prompting in our &lt;a href=&quot;https://img.ly/demos/ai-editor/web/&quot;&gt;AI Editor demo&lt;/a&gt;&lt;br&gt;
📘 &lt;a href=&quot;https://img.ly/docs/cesdk/js/user-interface/ai-integration/integrate-8e906c/&quot;&gt;Learn How to Integrate AI into CE.SDK&lt;/a&gt;&lt;br&gt;
💬 &lt;a href=&quot;https://img.ly/forms/contact-sales/&quot;&gt;Contact Us&lt;/a&gt; to Bring Visual Prompting to Your Product&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="https://blog.img.ly/2025/07/cesdk-2025-07-25T12_20_58.735Z--1-.png" medium="image"/><category>AI</category><category>CE.SDK</category><category>Creativity</category><category>Creative Workflows</category></item><item><title>Vibe-Engineering: When AI Does All the Coding, What Do We Actually Do?</title><link>https://img.ly/blog/vibe-engineering-when-ai-does-all-the-coding-what-do-we-actually-do/</link><guid isPermaLink="true">https://img.ly/blog/vibe-engineering-when-ai-does-all-the-coding-what-do-we-actually-do/</guid><description>When AI handles the coding, our value shifts to orchestrating, reasoning, and adapting. Guiding the process, not writing code.</description><pubDate>Mon, 07 Jul 2025 09:51:05 GMT</pubDate><content:encoded>&lt;p&gt;This isn’t an article about AI replacing engineers, it’s about discovering what (software) engineering will become when freed from the mechanics of coding itself.&lt;/p&gt;
&lt;p&gt;Vibe-Coding, and Vibe-Engineering are omnipresent in my social feeds these days. As with every new trend, it’s hard to judge what really works and what is pure marketing hype. But the central question nagged at me: if AI really can do all the coding, what exactly are we humans supposed to do?&lt;/p&gt;
&lt;p&gt;Therefore, I wanted to find out for myself. I didn’t want to build just a &lt;em&gt;Hello World&lt;/em&gt; example, so I pulled an old idea out of the closet and dusted it off.&lt;/p&gt;
&lt;h2 id=&quot;the-experiment-ai-agents-code-humans-curate&quot;&gt;The Experiment: AI Agents Code, Humans Curate&lt;/h2&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1536px) 1536px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1536&quot; height=&quot;1024&quot; src=&quot;https://img.ly/_astro/20250705_1708_Human-Robot-Coding-Collaboration_remix_01jzdj1by9ebmaxjtmw6z04z8q--1-_Z1wpIkA.webp&quot; srcset=&quot;/_astro/20250705_1708_Human-Robot-Coding-Collaboration_remix_01jzdj1by9ebmaxjtmw6z04z8q--1-_Z1I0A6E.webp 640w, /_astro/20250705_1708_Human-Robot-Coding-Collaboration_remix_01jzdj1by9ebmaxjtmw6z04z8q--1-_29uY3R.webp 750w, /_astro/20250705_1708_Human-Robot-Coding-Collaboration_remix_01jzdj1by9ebmaxjtmw6z04z8q--1-_ZEchyF.webp 828w, /_astro/20250705_1708_Human-Robot-Coding-Collaboration_remix_01jzdj1by9ebmaxjtmw6z04z8q--1-_Z1vkIgK.webp 1080w, /_astro/20250705_1708_Human-Robot-Coding-Collaboration_remix_01jzdj1by9ebmaxjtmw6z04z8q--1-_Z17XtcX.webp 1280w, /_astro/20250705_1708_Human-Robot-Coding-Collaboration_remix_01jzdj1by9ebmaxjtmw6z04z8q--1-_Z1wpIkA.webp 1536w&quot;&gt;&lt;/p&gt;
&lt;p&gt;For some time, I had considered porting our &lt;a href=&quot;https://github.com/imgly/background-removal-js&quot;&gt;IMG.LY background removal library&lt;/a&gt; from JavaScript to other platforms using Rust. However, I had postponed this side project due to the anticipated effort involved. This seemed like the perfect project to discover what my role would be in a world where AI does the heavy lifting. To pull it off, I self-imposed only one strict rule:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Hands off the code!&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Every commit, every fix, every feature would be handled by an AI coding agent. My role? Being the curator, formulating my intent, reviewing, play-testing, and guiding the agent with feedback.&lt;/p&gt;
&lt;p&gt;The agent should be the sole coder, from start to finish.&lt;/p&gt;
&lt;p&gt;This was a true hands-off experiment: could an AI build an entire Rust library if I never touched the code?&lt;/p&gt;
&lt;p&gt;Is it really something new? After years of handing off development to teammates as a CTO or teaching students, this didn’t feel completely different.&lt;/p&gt;
&lt;p&gt;So, what happens when you &lt;strong&gt;let an AI do all the coding&lt;/strong&gt;, and you just review and give feedback? What does it actually mean to be a human in this new paradigm?&lt;/p&gt;
&lt;p&gt;The result is baffling and gives a glimpse of what we as engineers can expect for the future.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;10,000+ lines of production-ready code. Authored by AI, orchestrated by a human.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Seeing is believing? Then explore the results and judge for yourself at &lt;a href=&quot;https://github.com/imgly/background-removal-rs/tree/v0.2.0&quot;&gt;github.com/imgly/background-removal-rs&lt;/a&gt;.&lt;br&gt;
I open-sourced the full code and the agent rules used during this experience.&lt;/p&gt;
&lt;p&gt;Here’s what I learned: vibe-engineering an entire Rust library with a coding agent CLI as my only pair programmer.&lt;/p&gt;
&lt;p&gt;On a quick note, I chose to use Claude Code with Opus 4 and Sonnet 4, as I was already familiar with them and had achieved the best results so far. The costs were kept in check by using a Claude Max Account, which at the time is around ~€100.&lt;/p&gt;
&lt;h3 id=&quot;the-general-workflow-that-emerged&quot;&gt;The General Workflow That Emerged&lt;/h3&gt;
&lt;p&gt;After some initial time to get used to working with an AI agent, I realized that the workflow follows pretty closely a typical software engineering flow, but with an important twist that answers our central question: what do humans actually do when AI codes? While most people focus on the &lt;code&gt;code writing&lt;/code&gt; part, I discovered that the &lt;code&gt;quality assurance&lt;/code&gt; part becomes the human’s primary domain. In my experiments, I realized that to keep the code intact over time, this QA phase is essential, and this is where humans truly shine. Most interestingly, with the right rules and prompting, we can let the agent do the heavy lifting here, too.&lt;/p&gt;
&lt;p&gt;Here’s how the collaboration actually worked in practice:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Getting Started&lt;/strong&gt;&lt;br&gt;
The project began with me setting up the environment and configuring timeouts for long-running tasks, because AI agents need time to compile Rust code and run extensive test suites. Together, the AI and I created a &lt;code&gt;CLAUDE.md&lt;/code&gt; file with coding standards and rules (you can see the &lt;a href=&quot;https://github.com/imgly/background-removal-rs/tree/main/.claude/rules&quot;&gt;actual rules we used&lt;/a&gt;). We also wrote project requirements documents collaboratively, with me providing the vision and the AI helping structure the technical details.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Development Dance&lt;/strong&gt;&lt;br&gt;
Once we got rolling, a natural rhythm emerged. I’d describe what I wanted, something like “add support for different image formats with color profile preservation.” The AI would then create a detailed implementation plan, breaking down the work into manageable chunks. I’d review this plan, often tweaking the approach or scope, then give the green light.&lt;/p&gt;
&lt;p&gt;What happened next was fascinating: the AI would write the code, run correctness checks, update all the tests (unit tests, documentation tests, end-to-end tests), run the linter, analyze coverage, and even update documentation and examples. Sometimes it would run performance benchmarks without me asking, the overzealous builder at work.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;My New Role: The Quality Guardian&lt;/strong&gt;&lt;br&gt;
My job became verifying that the intent was met and the API actually made sense. I’d play-test the functionality, try to break it in ways a real user might, and provide feedback. When something worked well, I’d document what we learned for future development. The AI would then write tests to lock in that new functionality, ensuring we never accidentally broke what was working.&lt;/p&gt;
&lt;p&gt;This cycle repeated for every feature, with the AI handling the heavy lifting while I focused on direction, quality, and user experience.&lt;/p&gt;
&lt;p&gt;The interesting part was that the more I got used to the workflow, the more I handed off to the agent. The validation and general experience had to be largely influenced or orchestrated by me, but building, testing, and even updating documentation could be delegated. In the beginning, I started just building things, but the more we built, the more the AI agent failed to keep existing functionality intact. With those small context windows, it’s kind of understandable after all.&lt;/p&gt;
&lt;h3 id=&quot;quality-assurance--knowledge-capturing-matters-more-than-ever&quot;&gt;Quality Assurance &amp;#x26; Knowledge Capturing Matters More Than Ever&lt;/h3&gt;
&lt;p&gt;As engineers, we know that at some points we cannot grasp every influence of our code changes on the whole code base anymore. Therefore, we use tests and tooling to help us keep our sanity.&lt;/p&gt;
&lt;p&gt;The good thing is that the Rust and Cargo toolchains are exceptionally high quality when it comes to providing the right tooling also for agents.&lt;/p&gt;
&lt;p&gt;Due to its lack of long-term context and limited knowledge of all the library’s capabilities, the tests and verification mechanisms become crucial. While these &lt;em&gt;are&lt;/em&gt; part of the context, they’re not stored in the agent’s memory but baked into the codebase itself. With the tests and the tooling being available to the agent, it will still make mistakes, as I do, but they can effectively &lt;code&gt;auto-heal&lt;/code&gt; or &lt;code&gt;auto-correct&lt;/code&gt; their own mistakes with the provided feedback.&lt;/p&gt;
&lt;p&gt;I can only advise being even stricter with QA from the very beginning.&lt;/p&gt;
&lt;p&gt;Another thing that became apparent very quickly is that, unlike a human colleague with lots of context and a good memory, agents &lt;em&gt;still&lt;/em&gt; need significant help to remember things to avoid rediscovering the same information repeatedly. This revealed another crucial human role: &lt;code&gt;context engineering&lt;/code&gt; becomes critically important. As such, we engineers become the keepers of institutional knowledge, the ones that help the agent remember what worked and what didn’t. I would describe it as &lt;code&gt;knowledge capture&lt;/code&gt;, for lack of a better term.&lt;/p&gt;
&lt;h2 id=&quot;tips--tricks&quot;&gt;Tips &amp;#x26; Tricks&lt;/h2&gt;
&lt;h3 id=&quot;ai-is-eager-resourceful-and-has-memory-like-a-sieve&quot;&gt;AI Is Eager, Resourceful, and Has Memory Like a Sieve&lt;/h3&gt;
&lt;p&gt;Here are the key insights I gained from building the library with an AI coding agent:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Code that works:&lt;/strong&gt; This might seem obvious, but it wasn’t clear to me if the library would ever be usable or publishable, but it is. The AI proved it could deliver production-ready code.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Time Effectiveness:&lt;/strong&gt; While creating the library took around 3 weeks, don’t let this fool you. Most of the time, the agent worked alone, and I only had to check in once in a while to review and give feedback. Iteration cycles and rewriting still take time. So I wouldn’t say development was faster than if I did it myself, but much more scalable and effective with my own time. Here’s what I actually spent my time on: reviewing architectural decisions, ensuring the API made sense, and play-testing the final product.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stamina:&lt;/strong&gt; The agent never complained, even when fixing 400+ warnings after setting the linter to pedantic. It pushed through. The AI powered through hundreds of warnings, even clustering them and suggesting which to fix first. To be fair, most humans would have either gone crazy or aborted the effort claiming “it’s good enough.” This revealed something important: as humans, we don’t need to be the ones grinding through tedious tasks anymore.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Overzealous Builder:&lt;/strong&gt; The agent often built more than I asked for. Scoping and clear implementation plans were essential. This taught me that one of the key human roles is setting boundaries and maintaining focus.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Premature “Done”:&lt;/strong&gt; Sometimes the agent claimed it was finished but left code mocked or incomplete. I learned to always ask, “Is there anything left to implement?” Quality assurance becomes fundamentally a human responsibility.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Old Knowledge:&lt;/strong&gt; The AI sometimes relied on outdated library knowledge. I had to remind it to check the latest docs and version of dependencies, but when asked, it searched the web, GitHub, crates.io, etc., and gathered all necessary information for using a library. Humans become the validators of information freshness.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Forgetting and Compaction:&lt;/strong&gt; Larger features grow outside the context window, so the agent would “compact” context, sometimes forgetting important details. Forcing it to write implementation plans in markdown files helps maintain flow and allows better resuming and bookkeeping. We humans become the keepers of long-term project memory.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Helps Establishing Rules:&lt;/strong&gt; While working with the agent, I saw repeating patterns like it forgetting how to check, build, and test the application, or best practices like using &lt;code&gt;git worktrees&lt;/code&gt;. I started adding new rules to &lt;code&gt;Claude.md&lt;/code&gt; by myself at first, but quickly realized that the agent is far better at formulating and adding new rules if I asked it to. The human role evolved into being the pattern recognizer and rule creator.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Doesn’t Always Obey the Rules:&lt;/strong&gt; Most of the time, the rules are followed, but occasionally it doesn’t follow them and forgets to automatically run all tests. This reinforced that humans must remain the final guardians of process and quality.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Code Deletion:&lt;/strong&gt; Occasionally, the AI “fixed” issues by removing important code paths. Solution: Insist on never removing functionality without discussion, and build a robust test suite. We become the protectors of existing functionality.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Process Management:&lt;/strong&gt; The AI struggled with background processes, understanding when to run things in the background, and keeping track of what it started. For example, starting a web server via &lt;a href=&quot;https://docs.anthropic.com/en/docs/agents-and-tools/tool-use/bash-tool&quot;&gt;&lt;code&gt;Bash&lt;/code&gt;&lt;/a&gt; (see &lt;a href=&quot;https://docs.anthropic.com/en/docs/agents-and-tools/tool-use/bash-tool&quot;&gt;Claude Code Bash tool docs&lt;/a&gt;), and then trying to use a &lt;a href=&quot;https://github.com/microsoft/playwright&quot;&gt;&lt;code&gt;playwright&lt;/code&gt; tool&lt;/a&gt; (see also &lt;a href=&quot;https://github.com/microsoft/playwright-mcp&quot;&gt;playwright-mcp&lt;/a&gt;) to access this server was blocked by these issues. Complex orchestration remains a distinctly human skill.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Overly Agreeable:&lt;/strong&gt; Too often, it seems to just agree with whatever I said. For the future, I wish agents wouldn’t be so obedient. This highlighted that humans need to be the challengers and devil’s advocates in the process.&lt;/p&gt;
&lt;h3 id=&quot;further-remarks&quot;&gt;Further Remarks&lt;/h3&gt;
&lt;p&gt;After three weeks of this experiment, the patterns became clear. Here’s what I learned about the human role in AI-assisted development:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Ensure Verify&lt;/strong&gt;, &lt;strong&gt;and Self-Correct:&lt;/strong&gt; Always ask the AI to check, format, lint, test, and benchmark its work. Your job becomes quality orchestration, not quality execution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Always Plan First:&lt;/strong&gt; Insist on a clear implementation plan before coding begins. Humans excel at high-level architectural thinking and breaking down complex problems.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Write It Down:&lt;/strong&gt; Let an agent keep notes, todos, and open issues in markdown files. Ideally in the repository itself. Don’t rely on the AI’s memory. We become the institutional memory keepers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Features:&lt;/strong&gt; Break tasks into small, manageable pieces. Feature scoping and boundary setting become a core human skill.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Test Suite is King:&lt;/strong&gt; A robust test suite prevents accidental “fixes” that remove functionality. Humans become the guardians of existing behavior.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stay Involved:&lt;/strong&gt; Fast iteration and discussion with the AI is crucial. Don’t go fully hands-off. Active curation and guidance remain essential.&lt;/p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;So, what do humans actually do when AI does all the coding?&lt;/p&gt;
&lt;p&gt;To answer that, let’s look at what AI agents are today:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Agents are super-talented, overly obedient coding assistants with lots of stamina and endless potential.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;They’re fast, never complain, and iterate like champs. But they need your guidance, structure, and a healthy dose of skepticism.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;In the end, the best results come from a true partnership: you plan and guide—the AI builds, fixes, and learns.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;If you follow these concepts, then&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;“Vibe engineering is a highly engaging experience.”&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And if you’re wondering: yes, I’d do it again. But next time, I’ll make sure the AI writes everything down too, and I’ll put even more effort into validating and securing new features as soon as they’re play-tested and verified. The human role isn’t disappearing. It’s evolving into something more strategic and impactful.&lt;/p&gt;
&lt;p&gt;Through this experiment, I’ve identified three fundamental shifts that define what engineering becomes in an AI-assisted world:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;From Execution to Orchestration&lt;/strong&gt;&lt;br&gt;
We’re no longer the ones typing code. We’re the conductors. Our new skills revolve around prompting (framing tasks so AI understands our intent), directing workflows (knowing when to use AI versus human judgment), and tool fluency (combining different AI tools to create something greater than their parts).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;From Knowledge to Reasoning&lt;/strong&gt;&lt;br&gt;
AI can recall facts and syntax better than any human ever could. But what matters now is our ability to interpret, contextualize, and make sense of uncertainty, something current AI still struggles with. We become the ones who ask “why” and “what if” rather than just “how.”&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;From Repetition to Adaptation&lt;/strong&gt;&lt;br&gt;
Skills rooted in routine are increasingly automated. The enduring value lies in adaptive thinking, problem-framing, and reinventing approaches when the usual patterns don’t apply. We’re the ones who recognize when something feels off, even if we can’t immediately articulate why.&lt;/p&gt;
&lt;h2 id=&quot;afterthought&quot;&gt;Afterthought&lt;/h2&gt;
&lt;p&gt;During this experiment, I saw that the agent repeated similar tasks over and over again. One thing that struck me was looking up documentation. Whenever something with a third-party library didn’t pan out, it started web-searching for info and guides about the library. It rarely started out with reading the Rust docs first, so I had to tell it to use the Rust docs. Therefore, I assume that providing quick access via tools to guides, docs, API references, and examples will spare me some roundtrips, time, and tokens and accelerate the process. I know that there are tools like &lt;a href=&quot;https://context7.com/&quot;&gt;context7&lt;/a&gt;, but it seems to be focused on JavaScript. At least for the Rust ecosystem, the docs are centralized and standardized and can even be used locally. A tool to access those quickly and find things in them would be highly beneficial.&lt;/p&gt;
&lt;p&gt;Additionally, for coding projects, there are typical steps to follow to verify if the code is sound. It’s probably also best to provide specific tools or workflows to follow these steps exactly. Rules already help with this, though.&lt;/p&gt;
&lt;p&gt;Last but not least, a natural next step would be for an agent to analyze its history periodically and provide new memory entries (rules) based on that to reduce repetition and common mistakes. For more details, see the &lt;a href=&quot;https://img.ly/blog/vibe-engineering-when-ai-does-all-the-coding-what-do-we-actually-do/#appendix-most-used-prompts&quot;&gt;&lt;em&gt;Appendix: Most used prompts&lt;/em&gt;&lt;/a&gt; section, where I analyzed the top 10 most common prompts to help agents proactively develop useful rules.&lt;/p&gt;
&lt;p&gt;If you’re curious about improvements, I’ve put together a &lt;a href=&quot;https://img.ly/blog/vibe-engineering-when-ai-does-all-the-coding-what-do-we-actually-do/#appendix&quot;&gt;&lt;em&gt;wishlist for Claude Code&lt;/em&gt;&lt;/a&gt; that would make the agent even more useful.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;try-it-yourself&quot;&gt;Try It Yourself&lt;/h2&gt;
&lt;p&gt;If you want to experience the results of vibe-engineering firsthand:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Clone and explore the codebase&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;git&lt;/span&gt;&lt;span&gt; clone&lt;/span&gt;&lt;span&gt; https://github.com/imgly/background-removal-rs.git&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;git&lt;/span&gt;&lt;span&gt; checkout&lt;/span&gt;&lt;span&gt; --tag&lt;/span&gt;&lt;span&gt; v0.2.0&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Or install the CLI directly&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;cargo&lt;/span&gt;&lt;span&gt; install&lt;/span&gt;&lt;span&gt; --git&lt;/span&gt;&lt;span&gt; https://github.com/imgly/background-removal-rs.git&lt;/span&gt;&lt;span&gt; --tag&lt;/span&gt;&lt;span&gt; v0.2.0&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Remember: every line of code in this library was written by AI while I played the role of curator, architect, and quality guardian. Judge for yourself whether this new paradigm produces production-ready results.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;appendix&quot;&gt;Appendix&lt;/h2&gt;
&lt;h3 id=&quot;appendix-claude-code-wishlist&quot;&gt;Appendix: Claude Code Wishlist&lt;/h3&gt;
&lt;p&gt;During the experiment, I encountered some issues whose resolution would improve the agent.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Project-scoped history (not global).&lt;/li&gt;
&lt;li&gt;Project-scoped todos (not global).&lt;/li&gt;
&lt;li&gt;Project-scoped implementation plans (not global).&lt;/li&gt;
&lt;li&gt;Customizable &lt;code&gt;/compact&lt;/code&gt; prompt or choose other compact strategies.&lt;/li&gt;
&lt;li&gt;Improved handling of background processes.&lt;/li&gt;
&lt;li&gt;History analytics with memory proposal.&lt;/li&gt;
&lt;li&gt;Code-specific predefined &lt;code&gt;Code and Validate&lt;/code&gt; workflows.&lt;/li&gt;
&lt;li&gt;Allow forking of multiple agents from a single point in the session history to create multiple trials with the same intent and context.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;appendix-most-used-prompts&quot;&gt;Appendix: Most Used Prompts&lt;/h3&gt;
&lt;p&gt;Claude code stores the history under &lt;code&gt;~/.claude/history&lt;/code&gt; as json formats. I asked claude code to read them in and categorize the most used prompts.&lt;/p&gt;



































































































































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;#&lt;/th&gt;&lt;th&gt;Category&lt;/th&gt;&lt;th&gt;% Usage (Count)&lt;/th&gt;&lt;th&gt;Description / Examples&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;1&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Other/Specific Instructions&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;19.5%&lt;/td&gt;&lt;td&gt;Technical specs, detailed requirements&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;em&gt;e.g.&lt;/em&gt;: “Relax the success metrics”, “Use tensor data directly as alpha channel”&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;2&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Questions&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;12.7%&lt;/td&gt;&lt;td&gt;Status checks, clarifications&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;em&gt;e.g.&lt;/em&gt;: “What’s next?”, “How can I test it myself?”, “Do we have any mock implementations?“&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;3&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Implementation Requests&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;11.1%&lt;/td&gt;&lt;td&gt;Direct build/create requests&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;em&gt;e.g.&lt;/em&gt;: “Create a PRD for Rust port”, “Make ONNX Runtime injectable”&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Simple Responses&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;10.0%&lt;/td&gt;&lt;td&gt;Short confirmations, approvals&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;em&gt;e.g.&lt;/em&gt;: “ok”, “go for it”, specific technical choices&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;5&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Continuation Commands&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;8.4%&lt;/td&gt;&lt;td&gt;Requests to proceed&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;em&gt;e.g.&lt;/em&gt;: “go on” (59 times)&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;&lt;strong&gt;File Operations&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;8.3%&lt;/td&gt;&lt;td&gt;File management&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;em&gt;e.g.&lt;/em&gt;: “Write this into PRD.md”, “Move packages into crates directory”&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;7&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Context Summaries&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;8.3%&lt;/td&gt;&lt;td&gt;Session continuation due to context limits&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;em&gt;e.g.&lt;/em&gt;: “This session is being continued from a previous conversation…“&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;8&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Analysis/Review Requests&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;7.7%&lt;/td&gt;&lt;td&gt;Requests to analyze/review&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;em&gt;e.g.&lt;/em&gt;: “Analyze my background removal project”, “Check the preprocessing”&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;9&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Bug Fix/Issue Resolution&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;4.4%&lt;/td&gt;&lt;td&gt;Problem identification/fixes&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;em&gt;e.g.&lt;/em&gt;: “You mix up release and debug”, “This is wrong”&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;10&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Testing&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;3.8%&lt;/td&gt;&lt;td&gt;Running/validating tests&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;em&gt;e.g.&lt;/em&gt;: “Test images are incorrect”, “Rerun comparison tests”&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;td&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;h3 id=&quot;appendix-claude-code-settings&quot;&gt;Appendix: Claude Code Settings&lt;/h3&gt;
&lt;p&gt;In real-world projects, the default timeout of two minutes is not enough. I bumped the timeouts to the maximum values to allow execution of unit tests, E2E tests, benchmarks, and long-running tasks.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;BASH_DEFAULT_TIMEOUT_MS&lt;/code&gt;: Sets the default timeout (in milliseconds) for long-running bash commands.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;BASH_MAX_TIMEOUT_MS&lt;/code&gt;: Specifies the maximum timeout (in milliseconds) that can be set for bash commands.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;BASH_MAX_OUTPUT_LENGTH&lt;/code&gt;: Limits the maximum number of characters in bash outputs before they are truncated in the middle.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;CLAUDE_BASH_MAINTAIN_PROJECT_WORKING_DIR&lt;/code&gt;: Ensures the working directory is reset to the original after each Bash command.&lt;/li&gt;
&lt;/ul&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;json&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;{&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &quot;env&quot;&lt;/span&gt;&lt;span&gt;: {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    &quot;CLAUDE_BASH_MAINTAIN_PROJECT_WORKING_DIR&quot;&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;&quot;true&quot;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    &quot;BASH_MAX_TIMEOUT_MS&quot;&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;&quot;3600000&quot;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    &quot;BASH_DEFAULT_TIMEOUT_MS&quot;&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;&quot;3600000&quot;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    &quot;MCP_TIMEOUT&quot;&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;&quot;3600000&quot;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    &quot;MCP_TOOL_TIMEOUT&quot;&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;&quot;3600000&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  }&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;}&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;library-capabilities&quot;&gt;Library Capabilities&lt;/h3&gt;
&lt;p&gt;High-performance Rust library for AI-powered background removal with hardware acceleration, built for&lt;br&gt;
production scale and developer productivity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Performance Highlights&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Hardware acceleration: CUDA (NVIDIA), CoreML (Apple Silicon), CPU fallback&lt;/li&gt;
&lt;li&gt;Sub-second processing on modern hardware (100-1200ms depending on image size)&lt;/li&gt;
&lt;li&gt;Memory efficient with optimized threading and session reuse&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Key Features&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;AI Models &amp;#x26; Quality&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Multiple state-of-the-art models (ISNet, BiRefNet)&lt;/li&gt;
&lt;li&gt;FP16/FP32 precision variants for performance vs. quality&lt;/li&gt;
&lt;li&gt;Portrait-optimized and general-purpose models&lt;/li&gt;
&lt;li&gt;Custom model support via ONNX models from &lt;a href=&quot;https://huggingface.co/&quot;&gt;Hugging Face&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Format and Color Profile Support&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Input: JPEG, PNG, WebP, TIFF, BMP with ICC color profile preservation&lt;/li&gt;
&lt;li&gt;Output: PNG (transparency), JPEG, WebP, TIFF, raw RGBA8&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Integration Patterns&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;One-liner API for simple use cases&lt;/li&gt;
&lt;li&gt;Session-based processing for batch operations&lt;/li&gt;
&lt;li&gt;Stream processing from any AsyncRead source&lt;/li&gt;
&lt;li&gt;CLI tool for standalone usage and pipeline integration&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Architecture &amp;#x26; Platforms&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Dual Backend System&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;ONNX Runtime: Maximum performance, GPU acceleration&lt;/li&gt;
&lt;li&gt;Tract: Pure Rust, zero external dependencies&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Platform Support&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;macOS: Apple Silicon + Intel with CoreML acceleration&lt;/li&gt;
&lt;li&gt;Linux/Windows: NVIDIA CUDA + CPU fallback&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Developer Experience&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Modern Rust Ecosystem&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Async/await native support&lt;/li&gt;
&lt;li&gt;Comprehensive documentation and examples&lt;/li&gt;
&lt;li&gt;Zero-warning policy with extensive testing&lt;/li&gt;
&lt;li&gt;Structured tracing for production observability&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;CLI Capabilities&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Batch processing with recursive directory support&lt;/li&gt;
&lt;li&gt;Model management (download, cache, clear)&lt;/li&gt;
&lt;li&gt;Provider diagnostics and performance monitoring&lt;/li&gt;
&lt;li&gt;Pipeline-friendly with stdin/stdout support&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;3,000+ creative professionals get early access to new features and updates. Don’t miss out, and&lt;/strong&gt; &lt;a href=&quot;https://share.hsforms.com/1IgAOV1wASXGPnFG4ZPLejg1hk3i?ref=img.ly&quot;&gt;&lt;strong&gt;subscribe&lt;/strong&gt;&lt;/a&gt; &lt;strong&gt;to our newsletter.&lt;/strong&gt;&lt;/p&gt;</content:encoded><dc:creator>Daniel</dc:creator><media:content url="https://blog.img.ly/2025/07/20250705_1708_Human-Robot-Collaboration_remix_01jzdj0cdxeacs9y9rbhsn04yy.png" medium="image"/><category>AI</category><category>Vibe Coding</category></item><item><title>IMG.LY x AI</title><link>https://img.ly/blog/img-ly-x-ai-future-of-ai-powered-creative-workflows/</link><guid isPermaLink="true">https://img.ly/blog/img-ly-x-ai-future-of-ai-powered-creative-workflows/</guid><description>Shaping the Future of Creative Workflows</description><pubDate>Mon, 30 Jun 2025 13:18:14 GMT</pubDate><content:encoded>&lt;p&gt;Over the past two years, one question has consistently emerged in conversations with customers: “What about AI?”&lt;br&gt;
While AI promises to disrupt many industries, it remains difficult to grasp how this technology will reshape creative workflows.&lt;/p&gt;
&lt;p&gt;Our customers look to us for guidance: What is IMG.LY’s vision? How will our SDK help them harness this wave of innovation?&lt;br&gt;
Last month, we took our first significant step by launching an &lt;a href=&quot;https://img.ly/use-cases/ai-editor/&quot;&gt;initial suite of generative AI-powered features for our web SDK&lt;/a&gt;. We’ve embedded these tools deeply into photo, video, and design editing workflows and the response from customers and prospects was overwhelmingly positive.&lt;/p&gt;
&lt;p&gt;And this is just the beginning. An immense opportunity lies ahead for both IMG.LY and our customers to drive transformation in the creative domain through our SDK. In this post, I want to share our vision for the future of creative tools powered by AI and our SDK.&lt;/p&gt;
&lt;p&gt;Going forward, we’re focusing on three central goals:&lt;/p&gt;
&lt;h3 id=&quot;1-deep-integration-of-ai-capabilities-into-editing-workflows&quot;&gt;1. Deep Integration of AI Capabilities into Editing Workflows&lt;/h3&gt;
&lt;p&gt;The pace of AI innovation continues to be remarkably fast. New models and configurations emerge almost daily, some as APIs, others open-sourced on platforms like Hugging Face. While some offer generalist features, others provide specialized, industry-specific capabilities, such as automatically obscuring license plates, blurring faces or replacing skies for property exteriors.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The true value of these AI tools emerges not in isolation, but when they’re seamlessly combined within existing workflows&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;That’s why we built a &lt;strong&gt;plugin system&lt;/strong&gt; for CE.SDK.&lt;br&gt;
The CE.SDK plugin system has fundamentally changed how quickly we, and our customers, can integrate new capabilities. We’ve created a way to bring AI models and agents directly into connected workflows within the editor, making AI feel native rather than bolted on.&lt;/p&gt;
&lt;p&gt;A recent example is &lt;a href=&quot;https://img.ly/blog/ai-first-visual-editor-for-gpt-4o-image-gen/&quot;&gt;our integration of OpenAI’s gpt-image-1 API&lt;/a&gt;, which we implemented in just a few days after its release. We used it to build a visual prompting workflow that takes into account all elements of a page, text, images, and annotations, to generate results based on the complete layout context.&lt;/p&gt;
&lt;p&gt;&lt;video src=&quot;https://storage.googleapis.com/imgly-static-assets/static/blog/videos/imgly-ai/visualprompt_05.mp4&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;p&gt;These early successes are encouraging us to ship even faster and bring new features to our customers.&lt;/p&gt;
&lt;p&gt;We’re now partnering with &lt;a href=&quot;https://fal.ai/&quot;&gt;fal.ai&lt;/a&gt; to provide a nearly out-of-the-box experience to use any new generative AI model through their platform (more on this soon). Fal.ai excels at speed: when exciting new models emerge, they offer API access within days. This partnership ensures our users can quickly access new AI capabilities with minimal effort. More partnerships and integrations are on the horizon, bringing world-class AI APIs with intuitive interfaces directly into the editor.&lt;/p&gt;
&lt;h3 id=&quot;2-enabling-ai-agents-as-creative-collaborators&quot;&gt;2. Enabling AI Agents as Creative Collaborators&lt;/h3&gt;
&lt;p&gt;AI agents, like humans, leverage tools to accomplish tasks efficiently. CE.SDK occupies a unique position as a highly adaptable, multi-platform technology for editing various media types, including a fully documented headless version that’s perfect for programmatic control.&lt;/p&gt;
&lt;p&gt;Thanks to CE.SDK’s fully documented API and headless architecture, AI agents can be created to navigate and operate the editor programmatically. This opens the door to agents that act as powerful scaffolders:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Generating initial designs and layouts based on brief descriptions&lt;/li&gt;
&lt;li&gt;Automatically aligning content with brand guidelines&lt;/li&gt;
&lt;li&gt;Transforming static designs into dynamic videos with a single command&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://img.ly/blog/how-to-build-a-short-video-generator-using-ce-sdk-2/&quot;&gt;Creating engaging short video content from prompts&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Adapting existing designs to new formats and dimensions&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Crucially, everything an AI agent creates remains &lt;strong&gt;fully editable&lt;/strong&gt; by human users. You can refine, add nuance, and perfect the results through collaboration. AI provides the scaffolding, humans remain the tastemakers, adding the creative spark that makes designs truly exceptional.&lt;/p&gt;
&lt;p&gt;&lt;video src=&quot;https://storage.googleapis.com/imgly-static-assets/static/blog/videos/imgly-ai/shortyroll.mp4&quot; controls loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;h3 id=&quot;3-ai-powered-sdk-configuration-and-customization&quot;&gt;3. AI-Powered SDK Configuration and Customization&lt;/h3&gt;
&lt;p&gt;AI agents are already actively involved in building, refining, and optimizing software tools, a trend that’s accelerating rapidly. This shift directly impacts developer experience and points to an exciting future where our SDK serves not just human developers, but AI agents as well.&lt;/p&gt;
&lt;p&gt;To facilitate this interaction, we’ve already taken steps like making our documentation available in LLM-friendly formats. But this is just the beginning of our journey toward radically improving the experience for both developers and AI agents.&lt;/p&gt;
&lt;p&gt;Our ultimate goal is conversational configuration. Imagine describing your requirements in plain language, and having an AI agent handle the rest:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Configuring the SDK’s visual aesthetics to match your brand&lt;/li&gt;
&lt;li&gt;Setting up custom functionality and workflows&lt;/li&gt;
&lt;li&gt;Integrating media libraries and plugins&lt;/li&gt;
&lt;li&gt;Optimizing performance for specific use cases&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This isn’t just about making development faster. It’s about democratizing access to powerful creative tools, allowing anyone to build sophisticated editing experiences, regardless of technical expertise.&lt;/p&gt;
&lt;p&gt;&lt;video src=&quot;https://storage.googleapis.com/imgly-static-assets/static/blog/videos/imgly-ai/vibe-coding-design-ads-visuals-editor.mp4&quot; controls loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;h2 id=&quot;looking-forward&quot;&gt;Looking Forward&lt;/h2&gt;
&lt;p&gt;We believe that AI collaboration represents a transformative shift in creative technology, empowering users and developers to achieve extraordinary outcomes. At IMG.LY, we’re committed to being at the forefront of this exciting journey. Our vision extends beyond simply adding AI features, we’re reimagining how creative tools are built, configured, and used in an AI-augmented world.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Stay ahead with us: &lt;a href=&quot;https://share.hsforms.com/1IgAOV1wASXGPnFG4ZPLejg1hk3i?ref=img.ly&quot;&gt;subscribe&lt;/a&gt; to our newsletter for exclusive updates.&lt;/p&gt;</content:encoded><dc:creator>Eray</dc:creator><media:content url="https://blog.img.ly/2025/06/creative-editor-sdk-1-53--1-.png" medium="image"/><category>AI</category><category>Company</category></item><item><title>AI-first Visual Editor using GPT-4o’s gpt-image-1 Model</title><link>https://img.ly/blog/ai-first-visual-editor-for-gpt-4o-image-gen/</link><guid isPermaLink="true">https://img.ly/blog/ai-first-visual-editor-for-gpt-4o-image-gen/</guid><description>We embedded OpenAI’s gpt-image-1 API into our CreativeEditor SDK, enabling users to generate and refine images directly in their creative workflow, no tool switching. This brings multimodal AI into real-world design tasks with seamless prompting, editing, and visual iteration.</description><pubDate>Mon, 05 May 2025 20:58:07 GMT</pubDate><content:encoded>&lt;h2 id=&quot;what-we-built&quot;&gt;What We Built&lt;/h2&gt;
&lt;p&gt;We integrated OpenAI’s new &lt;code&gt;gpt-image-1&lt;/code&gt; API (from GPT-4o) directly into our fully functional visual editor, &lt;strong&gt;CreativeEditor SDK (CE.SDK)&lt;/strong&gt;, enabling generation, editing, and refinement of images without ever leaving your creative workflow.&lt;/p&gt;
&lt;div class=&quot;cta-button-wrapper&quot;&gt;&lt;a href=&quot;https://cdn.img.ly/demo/gpt-image-1/v1/&quot; target=&quot;_blank&quot; class=&quot;cta-button&quot;&gt;Open AI Editor Demo Page&lt;/a&gt;&lt;/div&gt;
&lt;h3 id=&quot;from-simple-image-generation-to-visual-prompting-on-a-canvas&quot;&gt;From Simple Image Generation to Visual Prompting on a Canvas&lt;/h3&gt;
&lt;p&gt;Inside the editor, users can now:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Generate Images&lt;/strong&gt;&lt;br&gt;
Use prompts to generate images from scratch.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Generate Images from Visual Prompts&lt;/strong&gt;&lt;br&gt;
Turn full compositions (images, text, and annotations) into fresh visual content. Just select your page and let AI handle the rest, as shown in the video.&lt;/p&gt;
&lt;p&gt;&lt;video src=&quot;https://storage.googleapis.com/imgly-static-assets/static/blog/videos/imgly-ai/visualprompt_05.mp4&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Reimagine Images &amp;#x26; Text&lt;/strong&gt;&lt;br&gt;
Edit existing images and text with prompts to iterate faster and create &lt;strong&gt;variants&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;video src=&quot;https://storage.googleapis.com/imgly-static-assets/static/blog/videos/imgly-ai/variants_02.mp4&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Create Incredible Compositions&lt;/strong&gt;&lt;br&gt;
Combine generated and uploaded images into complex compositions.&lt;/p&gt;
&lt;p&gt;&lt;video src=&quot;https://storage.googleapis.com/imgly-static-assets/static/blog/videos/imgly-ai/combine.mp4&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;p&gt;Each step builds on the last, evolving from basic generation into true visual prompting powered by multiple input modes, all within one canvas. Check out the live demo &lt;a href=&quot;https://cdn.img.ly/demo/gpt-image-1/v1/index.rc6.html&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;how-we-built-it&quot;&gt;How We Built It&lt;/h2&gt;
&lt;p&gt;We built this integration using our &lt;strong&gt;CE.SDK&lt;/strong&gt; and its flexible &lt;strong&gt;plugin&lt;/strong&gt; system, designed from the ground up to support AI-first creative workflows.&lt;/p&gt;
&lt;p&gt;This approach lets developers plug in any model or API (text, image, video, or audio), and run them all in one seamless editing flow. Whether you’re using OpenAI, Stability, or an in-house model, CE.SDK gives you the tools to bring it into the visual workflow natively.&lt;/p&gt;
&lt;p&gt;🔗 Check out our &lt;a href=&quot;https://img.ly/use-cases/ai-editor/&quot;&gt;AI Editor&lt;/a&gt;.&lt;br&gt;
📘 Learn how to &lt;a href=&quot;https://img.ly/docs/cesdk/js/user-interface/ai-integration-5aa356/&quot;&gt;integrate AI into CE.SDK&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;why-this-matters&quot;&gt;Why This Matters&lt;/h2&gt;
&lt;p&gt;Generative AI’s full potential isn’t unlocked by prompting alone, it’s unlocked when embedded into real-world creative workflows.&lt;/p&gt;
&lt;p&gt;Designers, marketers, and content teams don’t just need outputs; they need &lt;strong&gt;control&lt;/strong&gt;, &lt;strong&gt;iteration&lt;/strong&gt;, and &lt;strong&gt;context&lt;/strong&gt;. By bringing AI directly into the canvas where assets are created and edited, we turn generative models into tools for actual production, not just ideation.&lt;/p&gt;
&lt;p&gt;This shift enables:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Creative work in context&lt;/strong&gt;: No switching between ChatGPT and design tools.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Real-time augmentation&lt;/strong&gt;: Prompt, edit, refine &lt;em&gt;in place&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scalable content generation&lt;/strong&gt;: Automate localization, personalization, and variants.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Multimodal orchestration&lt;/strong&gt;: Use visuals, layouts, and annotations as inputs.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It’s a step toward making multimodal AI usable for &lt;strong&gt;real design workflows&lt;/strong&gt;, not just concept generation.&lt;/p&gt;
&lt;h2 id=&quot;integration--feedback&quot;&gt;Integration &amp;#x26; Feedback&lt;/h2&gt;
&lt;p&gt;This linked demo is &lt;strong&gt;rate-limited&lt;/strong&gt;, if you would like to test more extensively or if you are interested in giving the AI editor a spin inside your own app, you can get started with our &lt;a href=&quot;https://img.ly/docs/cesdk/js/user-interface/ai-integration-5aa356/&quot;&gt;documentation&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;We’d love your feedback, any thoughts, questions, and ideas are welcome!&lt;br&gt;
&lt;a href=&quot;mailto:ai@img.ly&quot;&gt;Reach out&lt;/a&gt; to us.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;3,000+ creative professionals gain early access to new features and updates. Don’t miss out, and&lt;/strong&gt; &lt;a href=&quot;https://share.hsforms.com/1IgAOV1wASXGPnFG4ZPLejg1hk3i?ref=img.ly&quot;&gt;&lt;strong&gt;subscribe&lt;/strong&gt;&lt;/a&gt; &lt;strong&gt;to our newsletter.&lt;/strong&gt;&lt;/p&gt;</content:encoded><dc:creator>Eray</dc:creator><media:content url="https://blog.img.ly/2025/05/cesdk-2025-05-12T08_58_58.284Z.png" medium="image"/><category>AI</category><category>gpt-4o</category></item><item><title>OpenAI GPT-4o Image Generation (gpt-image-1) API: A Complete Guide for Creative Workflows for 2025</title><link>https://img.ly/blog/openai-gpt-4o-image-generation-api-gpt-image-1-a-complete-guide-for-creative-workflows-for-2025/</link><guid isPermaLink="true">https://img.ly/blog/openai-gpt-4o-image-generation-api-gpt-image-1-a-complete-guide-for-creative-workflows-for-2025/</guid><description>Learn how to integrate OpenAI’s gpt-image-1 API into modern creative applications. This complete 2025 guide covers technical setup, CE.SDK integration, prompt engineering, and tips for building real multimodal creative workflows.</description><pubDate>Mon, 28 Apr 2025 07:55:48 GMT</pubDate><content:encoded>&lt;h2 id=&quot;update-ai-first-visual-editing&quot;&gt;Update: AI-first Visual Editing&lt;/h2&gt;
&lt;p&gt;A day after the release of the &lt;code&gt;gpt-image-1&lt;/code&gt; API, we took it for a spin and integrated it into CreativeEditor SDK. Users can now generate images, create variants and use the canvas to compose visual prompts with our design editor. See it in action:&lt;/p&gt;
&lt;div class=&quot;cta-button-wrapper&quot;&gt;&lt;a href=&quot;https://cdn.img.ly/demo/gpt-image-1/v1/&quot; target=&quot;_blank&quot; class=&quot;cta-button&quot;&gt;Open AI Editor Demo Page&lt;/a&gt;&lt;/div&gt;
&lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;The release of OpenAI’s &lt;code&gt;gpt-image-1&lt;/code&gt; model signals a pivotal shift in the creative developer landscape, one that moves beyond static, one-shot image generation and toward a more dynamic, multimodal interaction model. Until recently, most image APIs followed a predictable pattern: submit a prompt, receive a finished image. The process was useful, but flat. What’s changing now is not just image quality or style fidelity, but the shape of the workflow itself. With &lt;code&gt;gpt-image-1&lt;/code&gt;, built on the GPT-4o foundation, developers can start designing creative tools that feel conversational and iterative. This evolution invites a new kind of interface where prompting, tweaking, and refining happen inside the canvas, not outside of it.&lt;/p&gt;
&lt;p&gt;For teams building creative editing experience into their app, this moment coincides with the release of &lt;a href=&quot;https://img.ly/demos/ai-editor/web/&quot;&gt;IMG.LY’s AI Editor SDK&lt;/a&gt;, a powerful, fully integrated toolkit designed for generative workflows. The SDK is already equipped to support interactive image generation, contextual editing, and multimodal inputs, and you can try it today through &lt;a href=&quot;https://img.ly/demos/&quot;&gt;this live demo&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;This guide is a comprehensive introduction to the &lt;code&gt;gpt-image-1&lt;/code&gt; API, but it also goes further. It’s not just about wiring up an endpoint, it’s about rethinking what image generation means in a user-centric product.&lt;/p&gt;
&lt;p&gt;From prompt handling to interactive iteration, we’ll walk through how to design creative cycles, not just outputs. This guide explores how to make that shift, how to go from generating images to integrating &lt;code&gt;gpt-image-1&lt;/code&gt; into real creative cycles, where AI becomes a tool that bends to user intent, not the other way around.&lt;/p&gt;
&lt;h2 id=&quot;overview-of-gpt-image-1&quot;&gt;Overview of &lt;code&gt;gpt-image-1&lt;/code&gt;&lt;/h2&gt;
&lt;p&gt;OpenAI’s &lt;code&gt;gpt-image-1&lt;/code&gt; model, released in April 2025, is the latest evolution in the company’s generative image lineup and marks a turning point in how developers approach visual creation inside applications. Built on the same multimodal foundation as GPT-4o, this model allows applications to move beyond one-shot static generation and instead build toward more conversational, iterative image workflows.&lt;/p&gt;
&lt;h3 id=&quot;model-architecture-and-capabilities&quot;&gt;Model Architecture and Capabilities&lt;/h3&gt;
&lt;p&gt;&lt;code&gt;gpt-image-1&lt;/code&gt; is rooted in GPT-4o’s ability to understand and generate across modalities. It is designed to produce high-resolution images (up to 4096×4096 pixels) based on natural language prompts. The model handles complex scenes with more fidelity than previous iterations and provides improved consistency in how it interprets detailed descriptions. This is particularly relevant for tools that need reliability when turning prompt inputs into design elements.&lt;/p&gt;
&lt;h3 id=&quot;parameter-control&quot;&gt;Parameter Control&lt;/h3&gt;
&lt;p&gt;Developers working with &lt;code&gt;gpt-image-1&lt;/code&gt; have access to a streamlined set of parameters, here is a subset of the most important ones:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;prompt&lt;/code&gt;: The primary text input describing the desired image.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;size&lt;/code&gt;: Choose between “1024x1024”, “1024x1536” (portrait), “1536x1024” (landscape), or “auto” (default, based on prompt).&lt;/li&gt;
&lt;li&gt;&lt;code&gt;n&lt;/code&gt;: Number of images to generate (default is 1).&lt;/li&gt;
&lt;li&gt;&lt;code&gt;response_format&lt;/code&gt;: Always returns &lt;code&gt;b64_json&lt;/code&gt;. URL outputs are not supported.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Unlike DALL·E 3, &lt;code&gt;gpt-image-1&lt;/code&gt; does not accept &lt;code&gt;style&lt;/code&gt; modifiers or &lt;code&gt;quality&lt;/code&gt; settings. It is designed for straightforward, high-fidelity image creation driven purely by the text prompt and size selection.&lt;/p&gt;
&lt;p&gt;Full documentation of these options is available via &lt;a href=&quot;https://platform.openai.com/docs/guides/images/usage&quot;&gt;OpenAI’s official guide&lt;/a&gt;.&lt;/p&gt;
&lt;h3 id=&quot;style-and-use-case-alignment&quot;&gt;Style and Use Case Alignment&lt;/h3&gt;
&lt;p&gt;By supporting a wide range of stylistic templates, &lt;code&gt;gpt-image-1&lt;/code&gt; positions itself as a flexible backend for everything from marketing collateral to storyboarding tools. The output can be tailored to suit technical illustrations, concept art, or even photorealistic renderings, allowing developers to map visual outputs more directly to brand or product requirements.&lt;/p&gt;
&lt;h3 id=&quot;limitations-and-future-direction&quot;&gt;Limitations and Future Direction&lt;/h3&gt;
&lt;p&gt;As of April 2025, &lt;code&gt;gpt-image-1&lt;/code&gt; supports only one image per request and does not offer fine-grained image editing or inpainting. However, its tight coupling with GPT-4o suggests that future iterations may embrace persistent context, conversational refinement, or even integrated image-plus-text exchanges within the same session. For developers building editors or multimodal workflows, the current model lays a strong foundation for these future capabilities.&lt;/p&gt;
&lt;h2 id=&quot;api-setup-and-usage&quot;&gt;API Setup and Usage&lt;/h2&gt;
&lt;h3 id=&quot;21-get-access&quot;&gt;2.1 Get Access&lt;/h3&gt;
&lt;p&gt;To start using &lt;code&gt;gpt-image-1&lt;/code&gt;, developers must first register for access via the OpenAI platform at &lt;a href=&quot;https://platform.openai.com/&quot;&gt;platform.openai.com&lt;/a&gt;. Access requires an API key, which is tied to your OpenAI account and associated usage limits based on your billing tier. Be sure to confirm that your account is approved for image generation, as availability may differ by region and subscription level. Once authenticated, keys can be created in your dashboard and stored securely in your server or development environment.&lt;/p&gt;
&lt;h3 id=&quot;22-first-image-generation-nodejs-example&quot;&gt;2.2 First Image Generation (Node.js Example)&lt;/h3&gt;
&lt;p&gt;The image generation API for &lt;code&gt;gpt-image-1&lt;/code&gt; can be used directly via OpenAI’s official Node.js client. Below is a complete example showing how to send a prompt and receive an image URL in response:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; OpenAI &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;openai&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; fs &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;fs&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; openai&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; new&lt;/span&gt;&lt;span&gt; OpenAI&lt;/span&gt;&lt;span&gt;({&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  apiKey: process.env.&lt;/span&gt;&lt;span&gt;OPENAI_API_KEY&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;// make sure this is securely set&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;});&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;async&lt;/span&gt;&lt;span&gt; function&lt;/span&gt;&lt;span&gt; generateImage&lt;/span&gt;&lt;span&gt;() {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  try&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    const&lt;/span&gt;&lt;span&gt; prompt&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; `&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    A studio ghibli style illustration of a cyberpunk girl holding a butterfly on her finger.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    `&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    const&lt;/span&gt;&lt;span&gt; result&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; await&lt;/span&gt;&lt;span&gt; openai.images.&lt;/span&gt;&lt;span&gt;generate&lt;/span&gt;&lt;span&gt;({&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      model: &lt;/span&gt;&lt;span&gt;&apos;gpt-image-1&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      prompt,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      size: &lt;/span&gt;&lt;span&gt;&apos;1024x1024&apos;&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;// or &quot;1024x1536&quot;, &quot;1536x1024&quot;, or &quot;auto&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    });&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    const&lt;/span&gt;&lt;span&gt; image_base64&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; result.data[&lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;].b64_json;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    const&lt;/span&gt;&lt;span&gt; image_bytes&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; Buffer.&lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt;(image_base64, &lt;/span&gt;&lt;span&gt;&apos;base64&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    fs.&lt;/span&gt;&lt;span&gt;writeFileSync&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;butterfly.png&apos;&lt;/span&gt;&lt;span&gt;, image_bytes);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    console.&lt;/span&gt;&lt;span&gt;log&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;Image saved as butterfly.png&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  } &lt;/span&gt;&lt;span&gt;catch&lt;/span&gt;&lt;span&gt; (err) {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    console.&lt;/span&gt;&lt;span&gt;error&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;Error generating image:&apos;&lt;/span&gt;&lt;span&gt;, err);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  }&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;}&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;generateImage&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Remember that all outputs from &lt;code&gt;gpt-image-1&lt;/code&gt; are delivered as base64-encoded JSON. Developers should decode this data for display, storage, or further processing within their applications. For complete parameter options and examples, consult the &lt;a href=&quot;https://platform.openai.com/docs/guides/images&quot;&gt;OpenAI Images API guide&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;integrating-with-cesdk&quot;&gt;Integrating with CE.SDK&lt;/h2&gt;
&lt;p&gt;Embedding &lt;code&gt;gpt-image-1&lt;/code&gt; into a creative editor like CE.SDK is about more than just piping an image into a canvas. It reshapes how users interact with content creation, bridging manual design work and AI-driven generation within the same editing environment. Rather than operating as a standalone prompt generator, &lt;code&gt;gpt-image-1&lt;/code&gt; becomes a continuous creative partner inside your editor. For in in-depth technical guide on how to integrate &lt;code&gt;gpt-image-1&lt;/code&gt; stay tuned for our upcoming tutorial, &lt;a href=&quot;https://share.hsforms.com/1IgAOV1wASXGPnFG4ZPLejg1hk3i?ref=img.ly&quot;&gt;sign up to our newsletter&lt;/a&gt; to be notified when it goes live.&lt;/p&gt;
&lt;h3 id=&quot;embedding-image-generation-in-a-creative-editing-workflow&quot;&gt;Embedding Image Generation in a Creative Editing Workflow&lt;/h3&gt;
&lt;p&gt;The natural entry point for &lt;code&gt;gpt-image-1&lt;/code&gt; inside CE.SDK is through a dual-mode experience: offering users the option to start either from scratch or from existing context. In “from scratch” mode, a user might open a blank scene and initiate an image generation by writing a prompt for example, “Create a vibrant festival scene at sunset.” The result appears directly on the canvas, immediately editable like any other design element.&lt;/p&gt;
&lt;p&gt;Where &lt;code&gt;gpt-image-1&lt;/code&gt; shows its real potential is in “in-context editing.” Here, users interact with existing content (a background, a product shot, or a decorative element) and trigger AI enhancements based on that visual context. A user might select an image of a bird, as in the example below and ask for variants, initiate a background swap, or request a change like adding more birds in a conversational interface embedded in the editor. Because CE.SDK treats generated images as first-class canvas elements, context such as positioning, layering, and cropping is preserved throughout the process.&lt;/p&gt;
&lt;p&gt;Let’s see what this might look like in practice. We positioned an image of a single bird on our canvas, opening the AI context menu we can now manipulate that image in place using the OpenAI API:&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 940px) 940px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;940&quot; height=&quot;560&quot; src=&quot;https://img.ly/_astro/Screenshot-2025-04-25-at-16.23.20_2408iY.webp&quot; srcset=&quot;/_astro/Screenshot-2025-04-25-at-16.23.20_hoIP0.webp 640w, /_astro/Screenshot-2025-04-25-at-16.23.20_Z1lGLpv.webp 750w, /_astro/Screenshot-2025-04-25-at-16.23.20_SQHtM.webp 828w, /_astro/Screenshot-2025-04-25-at-16.23.20_2408iY.webp 940w&quot;&gt;&lt;/p&gt;
&lt;p&gt;We edit the image and prompt the API to add more birds:&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1104px) 1104px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1104&quot; height=&quot;638&quot; src=&quot;https://img.ly/_astro/Screenshot-2025-04-25-at-11.19.07_Z1nRit4.webp&quot; srcset=&quot;/_astro/Screenshot-2025-04-25-at-11.19.07_Z2uDN7i.webp 640w, /_astro/Screenshot-2025-04-25-at-11.19.07_Z1XMUJ0.webp 750w, /_astro/Screenshot-2025-04-25-at-11.19.07_9QPa7.webp 828w, /_astro/Screenshot-2025-04-25-at-11.19.07_rW40P.webp 1080w, /_astro/Screenshot-2025-04-25-at-11.19.07_Z1nRit4.webp 1104w&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1056px) 1056px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1056&quot; height=&quot;548&quot; src=&quot;https://img.ly/_astro/Screenshot-2025-04-25-at-11.19.31_Z2pGvMv.webp&quot; srcset=&quot;/_astro/Screenshot-2025-04-25-at-11.19.31_Abfg0.webp 640w, /_astro/Screenshot-2025-04-25-at-11.19.31_1OeNO8.webp 750w, /_astro/Screenshot-2025-04-25-at-11.19.31_ZFflGu.webp 828w, /_astro/Screenshot-2025-04-25-at-11.19.31_Z2pGvMv.webp 1056w&quot;&gt;&lt;/p&gt;
&lt;p&gt;We see that the model correctly identified the type of bird in the picture (seagull) and filled it in with a swarm of flying seagulls.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1130px) 1130px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1130&quot; height=&quot;622&quot; src=&quot;https://img.ly/_astro/Screenshot-2025-04-25-at-11.35.50_Z1v4WfN.webp&quot; srcset=&quot;/_astro/Screenshot-2025-04-25-at-11.35.50_Z1yln80.webp 640w, /_astro/Screenshot-2025-04-25-at-11.35.50_Z1hUCg9.webp 750w, /_astro/Screenshot-2025-04-25-at-11.35.50_Z24uBxa.webp 828w, /_astro/Screenshot-2025-04-25-at-11.35.50_Z2ufBFR.webp 1080w, /_astro/Screenshot-2025-04-25-at-11.35.50_Z1v4WfN.webp 1130w&quot;&gt;&lt;/p&gt;
&lt;p&gt;We can now continue to work with the image, overlaying filters, changing the texture, cropping etc.&lt;/p&gt;
&lt;h3 id=&quot;switching-between-manual-edits-and-ai-powered-enhancements&quot;&gt;Switching Between Manual Edits and AI-Powered Enhancements&lt;/h3&gt;
&lt;p&gt;A critical design principle when integrating &lt;code&gt;gpt-image-1&lt;/code&gt; is giving users freedom to toggle between manual edits and AI suggestions. Manual edits should always remain possible after generation, e.g. cropping, masking, compositing while users can also seamlessly prompt &lt;code&gt;gpt-image-1&lt;/code&gt; for additional changes without losing prior work. Think of variant generation as a branch: a user picks a generated image and creates “forks” by asking for alternate styles, different lighting, or new thematic elements.&lt;/p&gt;
&lt;p&gt;In this setup, the generated image serves as a stable node in the creative graph, while edits and regenerations can attach contextually. This workflow minimizes user frustration by avoiding the “start over” penalty typical of isolated generation APIs. It also opens up more complex creative behaviors, like blending user-drawn sketches with AI-augmented refinements, or iteratively developing an asset library around a consistent visual theme.&lt;/p&gt;
&lt;p&gt;An upcoming in-depth tutorial will walk through implementing this multimodal workflow step-by-step, but the key takeaway is that &lt;code&gt;gpt-image-1&lt;/code&gt; shines brightest when it is embedded into a creative loop, not treated as a black-box generator, but as an interactive, iterative design companion.&lt;/p&gt;
&lt;h2 id=&quot;prompt-engineering-tips&quot;&gt;Prompt Engineering Tips&lt;/h2&gt;
&lt;p&gt;One of the most overlooked but critical factors in successful image generation is prompt design. With &lt;code&gt;gpt-image-1&lt;/code&gt;, prompt engineering isn’t just about describing an image. It’s about steering the model toward intent, tone, composition, and usability. Because the model is capable of rendering complex scenes and a wide range of styles, thoughtful phrasing and contextual hints can dramatically affect the outcome.&lt;/p&gt;
&lt;h3 id=&quot;writing-for-visual-intent&quot;&gt;Writing for Visual Intent&lt;/h3&gt;
&lt;p&gt;Start by clarifying what the image is supposed to communicate. Are you looking for atmosphere, action, product detail, or narrative clarity? A prompt like “a city skyline at night” is a starting point, but it leaves too much to chance. Adding elements like “view from a rooftop bar, with glowing signage and overcast haze” gives the model anchors for both composition and mood.&lt;/p&gt;
&lt;h3 id=&quot;leveraging-artistic-language&quot;&gt;Leveraging Artistic Language&lt;/h3&gt;
&lt;p&gt;You can further refine outputs by referencing mediums or artistic schools. Prompts that include terms like “in watercolor style,” “oil painting,” ”80s anime aesthetic,” or “studio photography” help the model lock onto a particular visual identity. These cues not only improve stylistic fidelity but also align the output with specific brand or genre expectations, which is especially important for products with a defined look and feel.&lt;/p&gt;
&lt;h3 id=&quot;creating-consistency-in-branded-outputs&quot;&gt;Creating Consistency in Branded Outputs&lt;/h3&gt;
&lt;p&gt;When generating a set of related images, such as social media creatives, campaign assets, or UI visuals, consistency becomes more important than variety. To achieve this, structure prompts with repeatable patterns and include brand elements such as color palettes, motifs, or reference characters. While &lt;code&gt;gpt-image-1&lt;/code&gt; doesn’t yet support persistent memory across requests, consistency can be enforced by prompting with the same style terms, layout descriptions, and constraints. Teams working within CE.SDK can even pair prompt templates with locked canvas layers to preserve composition between generations.&lt;/p&gt;
&lt;p&gt;Ultimately, good prompt engineering is not about verbosity but about clarity and constraint. It’s less like writing poetry and more like drafting a product spec. The best prompts are focused, directive, and give the model just enough creative freedom within clear boundaries. However, effective prompting should not burden the user. In practice, the interface should abstract most of the complexity away. Users can be guided toward better outputs through simple UI choices (selecting predefined styles, choosing themes, or adjusting mood settings) while the system dynamically enhances and augments their input behind the scenes. By managing the technical depth invisibly, you enable a creative process that feels intuitive and powerful without ever making prompt engineering the center of the user experience.&lt;/p&gt;
&lt;h2 id=&quot;real-world-use-cases&quot;&gt;Real-World Use Cases&lt;/h2&gt;
&lt;p&gt;The versatility of &lt;code&gt;gpt-image-1&lt;/code&gt; makes it especially impactful across a variety of industries where visual content creation is either a core product feature or a major operational need. Beyond isolated image generation, the model supports workflows that demand contextual awareness, brand consistency, and iterative refinement, key ingredients for modern digital products.&lt;/p&gt;
&lt;h3 id=&quot;web-to-print&quot;&gt;Web-to-Print&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/use-cases/web-to-print-design-tool/&quot;&gt;In web-to-print applications&lt;/a&gt;, customers expect to customize marketing materials, event invitations, signage, or packaging with minimal friction. By integrating &lt;code&gt;gpt-image-1&lt;/code&gt;, developers can offer template-driven personalization where users simply select a theme or enter a few keywords, and receive ready-to-edit visual assets. Combined with CE.SDK’s layout and editing capabilities, this enables a highly interactive experience where generated backgrounds, graphical elements, or themed illustrations can be dynamically placed into editable templates.&lt;/p&gt;
&lt;h3 id=&quot;social-media-marketing-and-martech&quot;&gt;Social Media Marketing and MarTech&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/industries/marketing-tech/&quot;&gt;Marketing teams rely on high-frequency content creation&lt;/a&gt;, often needing visually consistent, campaign-specific assets. &lt;code&gt;gpt-image-1&lt;/code&gt; can assist by automating the generation of background scenes, promotional visuals, and thematic graphics based on campaign briefs. Brands can define style presets aligned with their visual identity, making it easy for marketing teams to produce “on-brand” assets without heavy design overhead. Integrating image generation directly into campaign builders or social scheduling tools amplifies speed without sacrificing quality.&lt;/p&gt;
&lt;h3 id=&quot;digital-asset-management-dam&quot;&gt;Digital Asset Management (DAM)&lt;/h3&gt;
&lt;p&gt;Asset libraries often suffer from gaps: missing variants, seasonal versions, or content tailored to different demographics. &lt;a href=&quot;https://img.ly/industries/digital-asset-management/&quot;&gt;DAM systems&lt;/a&gt; can integrate &lt;code&gt;gpt-image-1&lt;/code&gt; to extend asset catalogs dynamically. Instead of manually commissioning variations, users can generate alternative backgrounds, localize visuals with region-specific elements, or adjust brand visuals for different markets, all from a single master file. With CE.SDK handling structured editing, teams maintain asset consistency while boosting creative flexibility.&lt;/p&gt;
&lt;h3 id=&quot;e-commerce&quot;&gt;E-Commerce&lt;/h3&gt;
&lt;p&gt;Product visualization remains a huge &lt;a href=&quot;https://img.ly/industries/e-commerce/&quot;&gt;challenge in e-commerce&lt;/a&gt;, especially for smaller retailers. &lt;code&gt;gpt-image-1&lt;/code&gt; can be used to automatically create product lifestyle imagery, context backgrounds, or thematic campaigns without expensive photo shoots. For example, a single shoe photograph can be placed into a generated “urban,” “sporty,” or “luxury” background, customized according to target audiences. When tightly integrated into e-commerce platforms, this enables faster product launches, A/B tested visuals, and localized campaigns at scale.&lt;/p&gt;
&lt;h3 id=&quot;e-learning&quot;&gt;E-Learning&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/industries/e-learning/&quot;&gt;Educational platforms&lt;/a&gt; can harness &lt;code&gt;gpt-image-1&lt;/code&gt; to generate explanatory diagrams, thematic illustrations, or scene-based visual storytelling assets. Instead of relying solely on static stock imagery, teachers, course designers, or even learners themselves can prompt the generation of custom visuals aligned with the curriculum. When embedded into authoring tools, this approach accelerates content creation and enables more engaging, visually enriched learning experiences tailored to specific topics and age groups.&lt;/p&gt;
&lt;h2 id=&quot;cost-optimization&quot;&gt;Cost Optimization&lt;/h2&gt;
&lt;p&gt;While &lt;code&gt;gpt-image-1&lt;/code&gt; opens up impressive creative possibilities, it also introduces new cost considerations that developers and product teams must plan for carefully. Since image generation typically incurs higher API costs than text-based operations, structuring workflows efficiently becomes critical, especially at scale.&lt;/p&gt;
&lt;h3 id=&quot;balancing-price-quality-and-resolution&quot;&gt;Balancing Price, Quality, and Resolution&lt;/h3&gt;
&lt;p&gt;The cost of generating an image with &lt;code&gt;gpt-image-1&lt;/code&gt; depends significantly on both the requested resolution and the selected quality setting. Higher resolutions like 4096×4096 produce sharper, more detailed results, but they also consume more compute resources-and therefore cost more. For many use cases, especially for previews, lower resolutions such as 1024×1024 or 2048×2048 strike an excellent balance between visual fidelity and API efficiency. Reserving the highest quality settings for final exports or premium workflows can help manage overall spend without compromising user experience.&lt;/p&gt;
&lt;h3 id=&quot;image-reuse-and-smart-upscaling&quot;&gt;Image Reuse and Smart Upscaling&lt;/h3&gt;
&lt;p&gt;One practical cost-saving approach is to design workflows that encourage image reuse. Instead of regenerating similar images for every small variation, applications can create high-quality master images and allow users to crop, edit, or layer additional design elements dynamically. Integrating smart upscaling techniques-for instance, using specialized image enhancement libraries after initial generation-also allows teams to work with smaller base images without sacrificing end-user quality.&lt;/p&gt;
&lt;h3 id=&quot;rate-limits-and-batching-strategies&quot;&gt;Rate Limits and Batching Strategies&lt;/h3&gt;
&lt;p&gt;Every call to &lt;code&gt;gpt-image-1&lt;/code&gt; counts toward your usage quota, and OpenAI imposes rate limits depending on account tier. To optimize performance and cost, it’s helpful to batch generation requests thoughtfully where possible-for instance, combining multiple prompts into structured queues or allowing users to preview low-res draft versions before finalizing a high-res render. Building this logic into your app’s generation flow not only controls expenses but also improves perceived responsiveness, an important UX factor for creative applications.&lt;/p&gt;
&lt;p&gt;By considering cost optimization as an early design constraint rather than a late-stage patch, developers can build scalable, sustainable creative tools powered by &lt;code&gt;gpt-image-1&lt;/code&gt;.&lt;/p&gt;
&lt;h2 id=&quot;bonus-starter-kit-repo&quot;&gt;Bonus: Starter Kit Repo&lt;/h2&gt;
&lt;p&gt;We are currently in the process of integrating the new GPT-4o-powered &lt;code&gt;gpt-image-1&lt;/code&gt; model into CE.SDK. As part of this effort, we are preparing a comprehensive Starter Kit will showcase a complete with CE.SDK integration, real-time prompt input, image generation workflows, and best practices for building an AI-powered creative editor.&lt;/p&gt;
&lt;p&gt;Both a public GitHub repository and a live demo will be made available soon. If you want to be notified when the Starter Kit launches, you can subscribe to updates &lt;a href=&quot;https://share.hsforms.com/1IgAOV1wASXGPnFG4ZPLejg1hk3i&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;This Starter Kit is designed to help developers move beyond simple image generation into building full creative cycles, where users can generate, edit, refine, and remix visuals seamlessly inside the editor.&lt;/p&gt;
&lt;h2 id=&quot;faqs&quot;&gt;FAQs&lt;/h2&gt;

&lt;p&gt;Choosing to work with &lt;code&gt;gpt-image-1&lt;/code&gt; raises a number of practical and strategic questions. Below, we address the most common topics for teams evaluating the model for integration into creative workflows.&lt;/p&gt;
&lt;h3 id=&quot;how-is-gpt-image-1-different-from-dalle-3&quot;&gt;How is &lt;code&gt;gpt-image-1&lt;/code&gt; different from DALL·E 3?&lt;/h3&gt;
&lt;p&gt;While DALL·E 3 and &lt;code&gt;gpt-image-1&lt;/code&gt; both translate text prompts into images, the underlying architecture and integration paths are different. &lt;code&gt;gpt-image-1&lt;/code&gt; is built on GPT-4o’s multimodal framework, making it better suited for future conversational and iterative workflows. It also offers support for a wider range of styles, higher resolutions up to 4096×4096 pixels, and is positioned for deeper integration into dynamic user experiences rather than one-off generation tasks.&lt;/p&gt;
&lt;h3 id=&quot;can-you-fine-tune-or-train-gpt-image-1&quot;&gt;Can you fine-tune or train &lt;code&gt;gpt-image-1&lt;/code&gt;?&lt;/h3&gt;
&lt;p&gt;As of April 2025, OpenAI does not allow fine-tuning of &lt;code&gt;gpt-image-1&lt;/code&gt;. The model is optimized for broad creative use cases out of the box. Developers seeking more control typically customize the user-facing prompt engineering or combine outputs with structured editing tools like CE.SDK to achieve brand or project-specific consistency.&lt;/p&gt;
&lt;h3 id=&quot;is-offline-support-available&quot;&gt;Is offline support available?&lt;/h3&gt;
&lt;p&gt;Currently, &lt;code&gt;gpt-image-1&lt;/code&gt; requires access to OpenAI’s cloud APIs. There is no offline inference mode or local deployment option. Teams requiring strict data residency, offline workflows, or private model hosting should consider hybrid architectures where images are generated securely via backend services and then edited locally using embedded tools like CE.SDK.&lt;/p&gt;
&lt;h3 id=&quot;what-about-copyright-and-licensing&quot;&gt;What about copyright and licensing?&lt;/h3&gt;
&lt;p&gt;Images generated by &lt;code&gt;gpt-image-1&lt;/code&gt; can be used commercially according to OpenAI’s &lt;a href=&quot;https://openai.com/en-GB/policies/usage-policies/&quot;&gt;usage policies&lt;/a&gt;, but developers are encouraged to review the latest terms. Outputs are not directly copyrighted by OpenAI or the user, and responsibility for ensuring compliance with branding, likeness, or content standards typically falls on the developer or platform operator. When deploying generation features to end-users, it is good practice to provide clear terms of use and, if needed, additional moderation or review layers.&lt;/p&gt;
&lt;p&gt;By addressing these considerations early, teams can integrate &lt;code&gt;gpt-image-1&lt;/code&gt; more effectively and responsibly into creative products and workflows.&lt;/p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;&lt;code&gt;gpt-image-1&lt;/code&gt; offers developers a significant opportunity to rethink what image generation can mean inside creative applications. It is not simply a tool for producing pictures on command, but a foundation for building interactive, iterative design workflows where users stay in control of the creative process. When combined with CE.SDK, it becomes even easier to move from static outputs to living, editable canvases that support real-world design needs. As we continue to integrate GPT-4o capabilities, the next wave of creative tooling will be about more than prompting images-it will be about shaping truly collaborative creative environments. Now is the time to start experimenting, iterating, and reimagining the user experience around this new generation of multimodal AI.&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="https://blog.img.ly/2025/04/GPT-Image-1-ultimate-guide.jpg" medium="image"/><category>AI</category><category>Creative Workflows</category><category>gpt-4o</category></item><item><title>How OpenAI&apos;s Upcoming GPT-4o Image Generation API Will Change Creative Workflows</title><link>https://img.ly/blog/open-ai-gpt-4o-image-generation-api-will-change-creative-workflows/</link><guid isPermaLink="true">https://img.ly/blog/open-ai-gpt-4o-image-generation-api-will-change-creative-workflows/</guid><description>OpenAI’s GPT-4o enables real-time, interactive image generation. Instead of one-off prompts, users can refine visuals through conversation. This unlocks new UX patterns like editable outputs and character persistence. IMG.LY’s CE.SDK makes GPT-4o easy to integrate into your editor.</description><pubDate>Mon, 14 Apr 2025 10:51:53 GMT</pubDate><content:encoded>&lt;p&gt;If you’ve been working with image-generation APIs over the past year, you’ve probably gotten used to a certain flow: send a prompt, wait a few seconds, and get a flat image back. It’s a one-shot deal. Useful? Definitely. But not exactly interactive. That’s what will change with OpenAI’s upcoming GPT-4o image-generation capabilities.&lt;br&gt;
IMG.LY, which recently &lt;a href=&quot;https://img.ly/use-cases/ai-editor/&quot;&gt;released a suite of AI features for its design editor&lt;/a&gt;, is eagerly awaiting the release to expand how users can interact with AI-driven creativity even further.&lt;/p&gt;
&lt;h2 id=&quot;update-ai-first-visual-editing&quot;&gt;Update: AI-first Visual Editing&lt;/h2&gt;
&lt;p&gt;A day after the release of the &lt;code&gt;gpt-image-1&lt;/code&gt; API, we put the UX principles outlined in this post into practice and integrated it into CreativeEditor SDK. Users can now generate images, create variants and use the canvas to compose visual prompts with our design editor. See it in action:&lt;/p&gt;
&lt;div class=&quot;cta-button-wrapper&quot;&gt;&lt;a href=&quot;https://cdn.img.ly/demo/gpt-image-1/v1/&quot; target=&quot;_blank&quot; class=&quot;cta-button&quot;&gt;Open AI Editor Demo Page&lt;/a&gt;&lt;/div&gt;
&lt;h2 id=&quot;gpt-4o-beyond-the-prompt-to-image-pipeline&quot;&gt;GPT-4o: Beyond the Prompt-to-Image Pipeline&lt;/h2&gt;
&lt;p&gt;GPT-4o isn’t just another version of DALL·E. It represents a shift in how developers will integrate AI into creative applications. While DALL·E 3 is powerful it is also somewhat siloed (you send a prompt, you get an image), GPT-4o looks like it will be part of a much more dynamic, conversational model one that accepts both text and image inputs, and could soon generate visual content in context, on the fly, and as part of a back-and-forth user interaction.&lt;/p&gt;
&lt;p&gt;If you’ve used ChatGPT recently, you’ve already seen glimpses of this. You can drop an image into the chat, ask GPT to describe or edit it, and get a response that feels fluid and visual. Developers should expect the API version to follow a similar pattern. It likely won’t just be a &lt;code&gt;/generate-image&lt;/code&gt; endpoint. Instead, we may be looking at an extension of the &lt;code&gt;chat/completions&lt;/code&gt; endpoint that handles multimodal messages. That changes the way you integrate this capability into your application. Rather than simply placing an image generation step in your pipeline, you will have to build your app’s UX around this new user flow. This comes with its own set of unique challenges.&lt;/p&gt;
&lt;h3 id=&quot;rethinking-the-interface-prompting-as-a-conversation&quot;&gt;Rethinking the Interface: Prompting as a Conversation&lt;/h3&gt;
&lt;p&gt;So what does this mean if you’re planning to integrate multi-modal image generation into your own product? For starters, you’ll probably need to rethink how users initiate and refine prompts. In the DALL·E flow, you might offer a text box with a few style dropdowns and call it a day. But in a GPT-4o world, your UI needs to support image inputs, persistent context, and dynamic editing, image gen becomes more like a conversation than a command.&lt;/p&gt;
&lt;p&gt;This is where the rubber meets the road. The tools that will benefit most from GPT-4o aren’t static generators but interactive editors. Think collaborative design apps, video editors with generative overlays, or product customizers that let users sketch or upload a photo and then iterate with AI. Put differently, the model output isn’t the endpoint but rather a checkpoint in the creation process.&lt;/p&gt;
&lt;h3 id=&quot;a-typical-iteration-cycle-in-a-multimodal-workflow&quot;&gt;A Typical Iteration Cycle in a Multimodal Workflow&lt;/h3&gt;
&lt;p&gt;Here’s a rough sketch of a workflow we might be seeing more of: The user starts with a prompt and an image, maybe a rough sketch or collage created inside an editor, a product photo, or a UI frame. GPT-4o returns a generated image based on that input. The user then edits or annotates the result, maybe adds new prompt text for refinement, and resubmits that combination to further develop the output. &lt;strong&gt;This cycle might loop several times&lt;/strong&gt;: generate, tweak, refine, regenerate.&lt;/p&gt;
&lt;p&gt;That’s a fundamentally different interaction model from past AI tooling. It’s less about one-off generation and more about a guided creative journey, where the user is in dialogue with the model. The result: better alignment with the original intent, more control, and more usable creative outputs.&lt;/p&gt;
&lt;p&gt;There is an additional, more subjective benefit to this kind of workflow: it gives the user a sense of autonomy again; they are back in the driver’s seat and less at the whim of an inscrutable machine. In many contexts, that makes a difference. Most notably, as we discussed in our white paper on &lt;a href=&quot;https://img.ly/white-papers/&quot;&gt;print personalization&lt;/a&gt;, the psychological benefit of personalization lies to a large extent in the investment, the sense of ownership that comes about when you create something. “Make it yours” is the common tagline attached to personalization campaigns in e-commerce. That only works if the user exerts more control over the output than iterating over a set of prompts.&lt;/p&gt;
&lt;p&gt;The most pithy encapsulation of this paradigm that I have heard is &lt;strong&gt;Humans on top, AI on tap&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id=&quot;persistent-elements-and-visual-consistency&quot;&gt;Persistent Elements and Visual Consistency&lt;/h3&gt;
&lt;p&gt;One particularly interesting frontier here is character and object persistence. If a user defines a character early in the workflow, either via prompt, image, or a combination, they’ll increasingly expect that character to appear consistently across assets. Think of it as visual continuity, whether you’re generating scenes in a story, slides in a deck, or frames in a video.&lt;/p&gt;
&lt;p&gt;If the user of a creative marketing cloud creates a campaign avatar or mascot, that &lt;strong&gt;character needs to be&lt;/strong&gt; &lt;strong&gt;consistent&lt;/strong&gt; within and across campaigns.&lt;/p&gt;
&lt;p&gt;Being able to reference earlier outputs, prompts, or style cues gives the user control over not just individual assets but the &lt;strong&gt;whole arc of the design narrative&lt;/strong&gt;. GPT-4o’s ability to maintain that continuity is a game-changer for workflows that involve storytelling, brand identity, or serialized design work.&lt;/p&gt;
&lt;h2 id=&quot;what-to-expect-from-the-api&quot;&gt;What to Expect from the API&lt;/h2&gt;
&lt;p&gt;Technically, if GPT-4o follows OpenAI’s recent design philosophy, you can expect a JSON-based API with a &lt;code&gt;messages&lt;/code&gt; array, where content can include both &lt;code&gt;text&lt;/code&gt; and &lt;code&gt;image_url&lt;/code&gt; types. The output will likely be returned either as an image URL hosted by OpenAI or as base64-encoded image data, depending on the format you request.&lt;/p&gt;
&lt;p&gt;That structure plays nicely with modern JavaScript front-end frameworks. React, Svelte, and Vue are all well-suited to async generation flows with visual previews. If you’re already using tools like &lt;a href=&quot;https://zustand.docs.pmnd.rs/&quot;&gt;Zustand&lt;/a&gt; or &lt;a href=&quot;https://jotai.org/&quot;&gt;Jotai&lt;/a&gt; for local state or something like &lt;a href=&quot;https://trpc.io/&quot;&gt;tRPC&lt;/a&gt; or &lt;a href=&quot;https://graphql.org/&quot;&gt;GraphQL&lt;/a&gt; for structured calls, you’re in a good position to layer GPT-4o in without breaking the flow.&lt;/p&gt;
&lt;h3 id=&quot;trade-offs-and-technical-considerations&quot;&gt;Trade-offs and Technical Considerations&lt;/h3&gt;
&lt;p&gt;There are trade-offs, of course. GPT-4o will probably cost more per call than a standard DALL·E 2 or 3 generation. Its latency is still an open question, and the multimodal input support will likely require more thoughtful UX decisions. What happens when a user drops an image and wants to undo just part of the generation? Where do you store prompt context for edits? How do you communicate what’s editable and what’s not?&lt;/p&gt;
&lt;p&gt;This is where design and engineering need to work together. You’ll want to build an interface that makes AI feel like a creative partner, not just a backend service. That might mean giving users a visual prompt history or allowing partial re-generations of specific canvas elements. You’ll need sensible fallback states. What happens when generation fails or the result isn’t what the user wanted?&lt;/p&gt;
&lt;h2 id=&quot;where-imglys-cesdk-fits-in&quot;&gt;Where IMG.LY’s CE.SDK Fits In&lt;/h2&gt;
&lt;p&gt;We have already given the questions raised above some serious thought, and most of the complexities introduced by this new workflow are the table stakes for the Creative Editor. So, if you’ve already integrated IMG.LY’s CE.SDK, we have taken care of most of these problems, and you can seamlessly integrate with any AI model. We are actively working on an off-the-shelf integration of the GPT-4o image model once its public API launches.&lt;/p&gt;
&lt;p&gt;In general, you can treat GPT-4o’s image outputs as just another layer in the editing canvas, positioned, styled, cropped, and ultimately editable in the same environment as everything else. That’s the real power of multimodal workflows: not just generating but integrating. And once GPT-4o’s API goes live, you’ll want your infrastructure ready to slot it in with minimal friction.&lt;/p&gt;
&lt;h3 id=&quot;the-loop-prompt-generate-refine&quot;&gt;The Loop: Prompt, Generate, Refine&lt;/h3&gt;
&lt;p&gt;The era of single-shot generation is winding down. What’s coming next is a loop: edit, prompt, generate, refine, repeat. And this loop doesn’t just belong in the backend, it needs to live in the UI, in a way that invites user input, creativity, and correction.&lt;/p&gt;
&lt;p&gt;We’ll be publishing more on how this integrates into IMG.LY’s upcoming AI workflows soon. Expect tools that don’t just generate visuals but help teams and individuals work through ideas in real time. Because especially as AI gets more potent, it needs humans on top.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;3,000+ creative professionals gain early access to new features and updates. Don’t miss out, and&lt;/strong&gt; &lt;a href=&quot;https://share.hsforms.com/1IgAOV1wASXGPnFG4ZPLejg1hk3i?ref=img.ly&quot;&gt;&lt;strong&gt;subscribe&lt;/strong&gt;&lt;/a&gt; &lt;strong&gt;to our newsletter.&lt;/strong&gt;&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="https://blog.img.ly/2025/04/GPT-4o-API-Changes-1.jpg" medium="image"/><category>AI</category><category>gpt-4o</category><category>Creative Workflows</category></item></channel></rss>