<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>Jan – IMG.LY Blog</title><description>Jan has orbited IMG.LY for over ten years, first as a developer, then as the company&apos;s first salesperson, and since 2022 as Head of Growth. He works at the seam where design meets engineering, and his mission is simple: get IMG.LY&apos;s first-class editing tools into the hands of ever more product builders, engineers, and tinkerers. He likes his articles long enough to proof-read over a Montecristo No. 5 Perla... yes, on a printout.</description><link>https://img.ly/blog/author/jan/</link><language>en-us</language><image><url>https://img.ly/apple-touch-icon.png</url><title>Jan – IMG.LY Blog</title><link>https://img.ly/blog/author/jan/</link></image><atom:link href="https://img.ly/blog/author/jan/rss.xml" rel="self" type="application/rss+xml"/><generator>Astro</generator><lastBuildDate>Tue, 01 Sep 2026 09:42:33 GMT</lastBuildDate><ttl>60</ttl><item><title>AI Design Agents in 2026: Which One Fits Your Work</title><link>https://img.ly/blog/ai-design-agents/</link><guid isPermaLink="true">https://img.ly/blog/ai-design-agents/</guid><description>Six AI design agents grouped by the work they suit: app screens, campaign assets, or design inside a larger task. Each entry says what you can still edit once the agent stops.</description><pubDate>Mon, 17 Aug 2026 09:00:00 GMT</pubDate><content:encoded>&lt;figure class=&quot;kg-embed-card&quot;&gt;
  &lt;iframe src=&quot;https://www.youtube.com/embed/T6jdkJMsi4Q?feature=oembed&quot; title=&quot;AI Design Agents Tested: Figma, Stitch, Lovart, Canva, CoDesign&quot; loading=&quot;lazy&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; referrerpolicy=&quot;strict-origin-when-cross-origin&quot; allowfullscreen&gt;&lt;/iframe&gt;
&lt;/figure&gt;
&lt;p&gt;Most roundups of AI design agents rank the tools one to six. That only works if every reader wants the same thing. A developer wiring up app screens and a brand manager producing a campaign are not shopping in the same market, even though the same six products keep showing up in both their search results.&lt;/p&gt;
&lt;p&gt;This survey sorts by the work instead. Two questions do most of the sorting: what you are making, and what you are left holding when the agent stops. If you want the wider tool category rather than agents specifically, we have &lt;a href=&quot;https://img.ly/blog/vibe-design-tools-compared/&quot;&gt;compared the AI design tools&lt;/a&gt; built around prompt-first workflows.&lt;/p&gt;
&lt;p&gt;We ran five of the six ourselves, on the same two prompts, and the findings below are what we saw rather than what the marketing pages say.&lt;/p&gt;
&lt;p&gt;Disclosure up front. IMG.LY publishes this survey, and two of the six tools here have a commercial relationship with us: CoDesign is our own product, and Manus is an IMG.LY customer. Both entries say so where they appear, we assess both against the same criteria as everything else, and anything we did not test ourselves rests on public sources.&lt;/p&gt;
&lt;h2 id=&quot;what-counts-as-a-design-agent&quot;&gt;What counts as a design agent&lt;/h2&gt;
&lt;p&gt;Three tests separate agents from the wider AI design tool space. The tool must plan multi-step work from a single goal. It must execute design operations rather than only advising. Its output must be usable design work: files, screens, or code.&lt;/p&gt;
&lt;p&gt;These tests exclude two familiar groups. Image generators such as Midjourney and Ideogram produce strong single images without planning or executing a task around them. Chat assistants configured with design prompts, such as Taskade’s design agent templates, return recommendations in text. Both are useful, but neither does design work on your behalf.&lt;/p&gt;
&lt;p&gt;Six products meet that definition in August 2026. We work through the definition in more detail in &lt;a href=&quot;https://img.ly/blog/what-is-a-design-agent/&quot;&gt;What Is an AI Design Agent?&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;how-we-tested&quot;&gt;How we tested&lt;/h2&gt;
&lt;p&gt;One prompt per category, given identically to every tool in that category.&lt;/p&gt;
&lt;p&gt;For the interface tools: design a three-screen onboarding flow for a habit-tracking app, the screens in order being a welcome screen, a goal-setting screen, and a first habit formation screen, styled clean, high contrast, large type, one accent color.&lt;/p&gt;
&lt;p&gt;For the graphic design tools: a launch campaign for a cold brew called Northwind, in three formats, an Instagram post, a story, and a DIN A5 flyer, warm and minimal, cream background, one product photo, headline supplied.&lt;/p&gt;
&lt;p&gt;Two things decide it, and neither is how good the first draft looks.&lt;/p&gt;
&lt;p&gt;Is the agent context aware? Can you feed it data, make it familiar with your brand, and does it hold that across every variant? We do not much care how fancy the design is. The table stakes have to be right before anything else counts.&lt;/p&gt;
&lt;p&gt;And does it fit the habits you carried over from the before-AI age? Making a small revision by going back and prompting an AI again is the wrong shape. You want to be the human in the loop who makes that edit manually, because by the time you are prompting for the third time your frustration is already high enough that you have stopped wanting to.&lt;/p&gt;
&lt;p&gt;One caveat on the Figma run, since it cuts against us: it executed the prompt inside IMG.LY’s own design system, so its icons, fonts, and patterns inherit work our designers had already done. That makes the head-to-head somewhat unfair in Figma’s favor.&lt;/p&gt;
&lt;p&gt;Manus is the exception. We have not run it hands-on, and its entry below rests on public sources.&lt;/p&gt;
&lt;h2 id=&quot;what-you-hold-afterward&quot;&gt;What you hold afterward&lt;/h2&gt;
&lt;p&gt;Generation quality converges fast. Every tool here produces a credible first draft, and the differences narrow every quarter. If you want to see where the underlying models differ, we &lt;a href=&quot;https://img.ly/ai-benchmarks/&quot;&gt;benchmark 15 of them&lt;/a&gt; against the same design prompts.&lt;/p&gt;
&lt;p&gt;What you keep does not converge. Some tools hand you a file inside an editor where you can select any element and change it. Others give you an export to finish in a different program. The rest keep the work inside their own platform, editable there and nowhere else.&lt;/p&gt;
&lt;h2 id=&quot;ai-design-agents-for-app-and-web-interfaces&quot;&gt;AI design agents for app and web interfaces&lt;/h2&gt;
&lt;p&gt;Screen design has the most mature agent tooling, partly because UI has structure a model can reason about and partly because the output can be code.&lt;/p&gt;
&lt;h3 id=&quot;google-stitch&quot;&gt;Google Stitch&lt;/h3&gt;
&lt;p&gt;Stitch is a free experimental tool from Google Labs that generates multi-screen mobile and web interfaces, plus frontend code, from text or image prompts. It launched at I/O 2025 and gained Gemini 3 and interactive prototype flows in December 2025.&lt;/p&gt;
&lt;p&gt;You describe an app and Stitch plans the screens, which you refine through further prompts. Screens render in the browser, with code export for the web stack and a paste-to-Figma path that preserves layers. Stitch has no source format of its own, so you can only keep editing in the exported code or the Figma file it hands off to. Reviewers recommend exploring in Stitch and refining elsewhere.&lt;/p&gt;
&lt;p&gt;On our onboarding prompt it did not adhere to the brief. Habit formation came before goal setting, so the screen order was wrong. Button styles were somewhat inconsistent, the look and feel differed between screens, visible in the progress indicator on the welcome screen against the one on goal setting, and the aspect ratios were weird.&lt;/p&gt;
&lt;p&gt;Editing is where it separates from Figma. Double-clicking does let you change the text, but you are apparently supposed to edit with AI, which is a nuisance when you only want to make a few design fixes. The brand kit is the part we liked: colors, accent colors, fonts, and spacing can all be changed globally. Though it looked like Stitch did not adhere to its own system that well.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;Google Stitch showing the generated onboarding screens, with habit formation appearing before goal setting and the progress indicators differing between screens&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1920px) 1920px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1920&quot; height=&quot;1158&quot; src=&quot;https://img.ly/_astro/stitch-ui.BpsuLGO1_21CgRX.webp&quot; srcset=&quot;/_astro/stitch-ui.BpsuLGO1_2mzJWp.webp 640w, /_astro/stitch-ui.BpsuLGO1_2c3zyR.webp 750w, /_astro/stitch-ui.BpsuLGO1_Z21Ahyw.webp 828w, /_astro/stitch-ui.BpsuLGO1_1RtJSh.webp 1080w, /_astro/stitch-ui.BpsuLGO1_Z1l1Hy3.webp 1280w, /_astro/stitch-ui.BpsuLGO1_2aSrbg.webp 1668w, /_astro/stitch-ui.BpsuLGO1_21CgRX.webp 1920w&quot;&gt;&lt;/p&gt;
&lt;p&gt;It is free while in Google Labs. Google does not publish the generation caps, and third-party sources describe the limits inconsistently. Best fit is getting a first version on screen fast, especially for developers who want screens and starter code in one pass. Its Labs status makes it a tool to explore with rather than build a production process on.&lt;/p&gt;
&lt;h3 id=&quot;figma-agent&quot;&gt;Figma Agent&lt;/h3&gt;
&lt;p&gt;Figma’s agent works directly on the design canvas. It entered beta in May 2026 and gained custom skills, web search, and external MCP connections at Config in June 2026.&lt;/p&gt;
&lt;p&gt;You prompt from any layer. The agent executes multi-step tasks: bulk edits across screens, dark-mode conversion, populating designs with real content, turning feedback threads into revisions. Custom skills written as markdown steer how it works.&lt;/p&gt;
&lt;p&gt;Output is native Figma layers with components, variables, and design tokens preserved. This is the strongest design-system integration of the six, because the agent works inside the system your team already maintains rather than approximating it. Code generation routes through Figma Make, a separate product.&lt;/p&gt;
&lt;p&gt;On our onboarding prompt it executed almost perfectly. The design was very clean, it adhered to the standards of modern apps, and it almost looked like something you would find browsing the App Store. Even the copy was decent, and it did not read like AI slop. The screens looked thought through: asked for a primary goal, it offered routines, breaking a bad habit, and staying consistent, which are the broad categories without too many assumptions on top. It stuck to the screen order, it baked in some interactivity, and it caught small details like “remind me every day at 8 a.m.” on the habit formation screen. As a novice you could take this and start prototyping.&lt;/p&gt;
&lt;p&gt;Remember the caveat above, though: this run inherited IMG.LY’s design system, so the icons, fonts, and patterns were already good before the agent started.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;The Figma Agent&amp;#39;s three-screen onboarding flow on the Figma canvas: welcome, goal setting and first habit, in order, with consistent buttons and type&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1920px) 1920px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1920&quot; height=&quot;1386&quot; src=&quot;https://img.ly/_astro/figma-ui.Y16RdKY8_Z268T6m.webp&quot; srcset=&quot;/_astro/figma-ui.Y16RdKY8_2v5Q6d.webp 640w, /_astro/figma-ui.Y16RdKY8_Z1qzg7w.webp 750w, /_astro/figma-ui.Y16RdKY8_ZXCx2r.webp 828w, /_astro/figma-ui.Y16RdKY8_MFTUh.webp 1080w, /_astro/figma-ui.Y16RdKY8_Z2rSWLd.webp 1280w, /_astro/figma-ui.Y16RdKY8_1Dj2Pf.webp 1668w, /_astro/figma-ui.Y16RdKY8_Z268T6m.webp 1920w&quot;&gt;&lt;/p&gt;
&lt;p&gt;The agent is free during beta, and Figma has said AI credits will apply at general availability without publishing amounts. It requires a full seat on Professional, Organization, or Enterprise; full seats start at $16 per person per month on the annual Professional plan. Early reviews report rough edges, including responsive layouts that broke on mobile.&lt;/p&gt;
&lt;p&gt;In the head-to-head, Figma beat Stitch by a wide margin, which was to be expected.&lt;/p&gt;
&lt;p&gt;Best fit is product and UI teams already working in Figma with an established design system.&lt;/p&gt;
&lt;h2 id=&quot;ai-design-agents-for-marketing-and-brand-assets&quot;&gt;AI design agents for marketing and brand assets&lt;/h2&gt;
&lt;p&gt;In campaign work you are producing dozens of assets at once, they all have to stay on brand, and some of them have to survive contact with a printer.&lt;/p&gt;
&lt;h3 id=&quot;lovart&quot;&gt;Lovart&lt;/h3&gt;
&lt;p&gt;Lovart is a dedicated design agent that turns a brief into a coordinated set of deliverables: logos, posters, packaging, social assets. It launched publicly in July 2025 after a closed beta that drew several hundred thousand users.&lt;/p&gt;
&lt;p&gt;Give it a campaign-level brief and it analyzes intent, researches references, then generates dozens of assets sharing one visual system. You and the agent iterate on ChatCanvas, an infinite canvas where you both edit the same design through conversation. Results land as layers with typography kept separately, and Lovart offers targeted element and text edits in place.&lt;/p&gt;
&lt;p&gt;On the Northwind campaign it was fairly quick, and the result was very aesthetically pleasing. We had very little to quibble with. The design subtleties were there, the product placement, the coffee beans, some cloth, and it stuck very well to the specification. Exactly what we had in mind.&lt;/p&gt;
&lt;p&gt;The one quibble is brand drift. Look at the logo, look at the bottle, and there is significant drift in brand identity across the three formats. That is fixable in a real scenario, because you can link a brand, so you would have the product image and the logo ready and we would not expect much trouble then, whether you use Lovart to create variations, produce marketing material to A/B test, or adjust for different formats.&lt;/p&gt;
&lt;p&gt;Editing is the problem. Ask to change the headline text and you cannot tell whether you are even looking at the same font. Quick edit means writing another prompt. There are no layers, and there is nothing you can do directly. So while the output looks nice, for revision work it is effectively useless.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;Lovart&amp;#39;s canvas with the Northwind campaign assets, where editing the headline means writing another prompt rather than selecting a text layer&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1920px) 1920px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1920&quot; height=&quot;1027&quot; src=&quot;https://img.ly/_astro/lovart-ui.DIzU1wlA_Zh3UPf.webp&quot; srcset=&quot;/_astro/lovart-ui.DIzU1wlA_248rGS.webp 640w, /_astro/lovart-ui.DIzU1wlA_3cGnN.webp 750w, /_astro/lovart-ui.DIzU1wlA_ZQVJHY.webp 828w, /_astro/lovart-ui.DIzU1wlA_2hTJac.webp 1080w, /_astro/lovart-ui.DIzU1wlA_Z18qvYg.webp 1280w, /_astro/lovart-ui.DIzU1wlA_UlQg1.webp 1668w, /_astro/lovart-ui.DIzU1wlA_Zh3UPf.webp 1920w&quot;&gt;&lt;/p&gt;
&lt;p&gt;Reviewers report that precise pixel-level work tends to move into Photoshop or Figma anyway, which matches what we found.&lt;/p&gt;
&lt;p&gt;Pricing runs on credit tiers from Starter to Ultimate, 2,000 to 27,000 credits per month. Lovart renders prices dynamically and third-party reports of the dollar amounts vary between roughly $15 and $90 monthly, so check the pricing page directly. Credit consumption is hard to predict on agent-driven tasks.&lt;/p&gt;
&lt;p&gt;Best fit is marketing teams whose unit of work is the campaign rather than the asset, and who accept a finishing pass elsewhere.&lt;/p&gt;
&lt;h3 id=&quot;canva-ai-assistant&quot;&gt;Canva AI Assistant&lt;/h3&gt;
&lt;p&gt;Canva’s assistant plans and executes design tasks by calling Canva’s own tools. It launched in April 2025 and received a tool-calling rebuild, announced as a research preview, in April 2026.&lt;/p&gt;
&lt;p&gt;It runs multi-step jobs such as producing a multi-channel campaign, and pulls context from connected sources including Slack, Gmail, Google Drive, and Notion. Scheduled tasks run in the background, with finished work arriving as drafts for review.&lt;/p&gt;
&lt;p&gt;Output is layered, fully editable Canva designs across presentations, social posts, documents, and spreadsheets. You can change any element without regenerating. The constraint is Canva itself. The work stays editable inside Canva, and getting it out means exporting.&lt;/p&gt;
&lt;p&gt;On the same Northwind prompt it was fairly quick, and it appears to match the brief against the vast template library Canva already has. From the get-go, though, it did not really stick to the specification. We asked for a cream background and one product photo. There is no cream background on the first asset, and we do not know why the product would be displayed on a laptop, which is plain weird. Only the last of the three is anywhere near acceptable.&lt;/p&gt;
&lt;p&gt;We also had no indication of whether it generated variations or ignored that part of the brief. Opening one of the designs gives you more variations, including the one you did not select. What you do not get is a multi-page layout you can change, or any way to specify changes precisely. At that point you are simply inside Canva, editing.&lt;/p&gt;
&lt;p&gt;As a starting point we would have been just as well off picking one of the Canva templates, uploading the product image, and adjusting the text. Compared with Lovart on the identical brief, Canva really does fall short.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;Canva&amp;#39;s assistant returning the Northwind assets, without the cream background the brief asked for and with the product shown on a laptop&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1920px) 1920px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1920&quot; height=&quot;1092&quot; src=&quot;https://img.ly/_astro/canva-ui.DpSjPHH5_Z1vYvTM.webp&quot; srcset=&quot;/_astro/canva-ui.DpSjPHH5_jD6q9.webp 640w, /_astro/canva-ui.DpSjPHH5_2cxk3p.webp 750w, /_astro/canva-ui.DpSjPHH5_Z1NMplW.webp 828w, /_astro/canva-ui.DpSjPHH5_Z1fW75G.webp 1080w, /_astro/canva-ui.DpSjPHH5_Z1aCzQe.webp 1280w, /_astro/canva-ui.DpSjPHH5_Z1vGnHn.webp 1668w, /_astro/canva-ui.DpSjPHH5_Z1vYvTM.webp 1920w&quot;&gt;&lt;/p&gt;
&lt;p&gt;The assistant is available on the free tier with rate limits and monthly credits. Pro is $18 monthly. Business is $25 per user monthly with larger allowances. Reviewers warn that letting the assistant run a job end to end uses premium credits faster than editing by hand.&lt;/p&gt;
&lt;p&gt;Best fit is teams already on Canva who want campaign production handled conversationally with brand kits applied automatically.&lt;/p&gt;
&lt;h3 id=&quot;imgly-codesign&quot;&gt;IMG.LY CoDesign&lt;/h3&gt;
&lt;p&gt;Disclosure: CoDesign is made by IMG.LY, which publishes this survey.&lt;/p&gt;
&lt;p&gt;CoDesign is a free local MCP server that gives an agent you already use a full design engine. It entered technical preview in June 2026.&lt;/p&gt;
&lt;p&gt;Instead of going to CoDesign, you install it into your own agent with one command, and the agent then performs design operations in conversation with you. IMG.LY documents the install for &lt;a href=&quot;https://img.ly/codesign/&quot;&gt;Claude Code and Codex&lt;/a&gt;, and any client that can run an MCP server locally works the same way. CoDesign can generate, rebrand, resize, edit, localize, import, and judge, the last checking a design against your brand rules. A representative chain: import a PSD, rebrand it, resize it to a dozen formats, check it against the rules, export.&lt;/p&gt;
&lt;p&gt;Output is a structured scene rather than a flat image. It opens in an editor where you select any element and change it without regenerating anything. Underneath is CE.SDK, IMG.LY’s production design SDK. It imports PSD, IDML, PDF, and PPTX, and exports PDF with CMYK, bleed, and ICC profiles when a printer needs them. The agent, your code, and a human editor all work on the same file.&lt;/p&gt;
&lt;p&gt;We ran the Northwind prompt through Claude Code. Before touching the canvas it asked where the product photo should come from, what accent color to use, what kind of cream we meant, and it drafted the copy for approval. In normal use you would run this inside your marketing documentation and brand assets, so your coding agent already carries a pile of context and can answer most of that itself. We gave it no brand context at all, to keep the comparison fair with the others.&lt;/p&gt;
&lt;p&gt;It left us waiting a bit longer, though that is not quite fair to say, since it works in tandem with the coding agent. It was thorough and diligent. It ran a self-check against a set of axes before handing the design back, and because we supplied no brand context those came back blank. It passed, and returned one editable master file.&lt;/p&gt;
&lt;p&gt;When it needed the product image it opened the IMG.LY dashboard, where the AI Gateway picked the model for it and credited the generation, which was frictionless.&lt;/p&gt;
&lt;p&gt;The design that came back is fairly conservative and does not assume too much: some copy, “now pouring, limited first batch”, a CTA. Fairly bare bones, but it works. What matters is that it stuck perfectly to the specification. The square asset, the story asset, and a PDF carrying its color space and a resolution we could hand to a printer. Edit one element, tell it to regenerate, and it propagates: same colors, same logo, same product image across every variant.&lt;/p&gt;
&lt;p&gt;All three formats came out of a single master file, generated in one pass from the same brief:&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;The three Northwind formats CoDesign returned from one brief: the DIN A5 flyer on the left, then the square Instagram post and the vertical story side by side, all carrying the same headline, colors and product photo&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 2168px) 2168px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;2168&quot; height=&quot;900&quot; src=&quot;https://img.ly/_astro/codesign-outputs.ByTlv-EZ_1l0LHA.webp&quot; srcset=&quot;/_astro/codesign-outputs.ByTlv-EZ_1WiJfB.webp 640w, /_astro/codesign-outputs.ByTlv-EZ_ZMeRqI.webp 750w, /_astro/codesign-outputs.ByTlv-EZ_Z2c0iCf.webp 828w, /_astro/codesign-outputs.ByTlv-EZ_Z2hT6nK.webp 1080w, /_astro/codesign-outputs.ByTlv-EZ_Z1t3588.webp 1280w, /_astro/codesign-outputs.ByTlv-EZ_1T2Tz8.webp 1668w, /_astro/codesign-outputs.ByTlv-EZ_y88Cz.webp 2048w, /_astro/codesign-outputs.ByTlv-EZ_1l0LHA.webp 2168w&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;The CoDesign editor with the Northwind master file open: the headline text layer selected for a direct edit, a shapes library on the left, and a Download PDF button in the toolbar&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1497px) 1497px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1497&quot; height=&quot;827&quot; src=&quot;https://img.ly/_astro/codesign-ui.Cr8GPPF2_Kvoq4.webp&quot; srcset=&quot;/_astro/codesign-ui.Cr8GPPF2_1S0RsV.webp 640w, /_astro/codesign-ui.Cr8GPPF2_Z1IJheT.webp 750w, /_astro/codesign-ui.Cr8GPPF2_1RszbR.webp 828w, /_astro/codesign-ui.Cr8GPPF2_ZhArp0.webp 1080w, /_astro/codesign-ui.Cr8GPPF2_1FlAe7.webp 1280w, /_astro/codesign-ui.Cr8GPPF2_Kvoq4.webp 1497w&quot;&gt;&lt;/p&gt;
&lt;p&gt;The local server is free to install and run with no account needed to start. An optional free account adds AI image generation, which routes out through IMG.LY’s AI Gateway. Paid licensing applies when you host the server for other people or embed it in a product.&lt;/p&gt;
&lt;p&gt;Best fit is anyone running an agent-first local workflow who needs production-grade editing, print-ready files, and brand consistency across every variant. As a technical preview it is early, and behavior can change between releases.&lt;/p&gt;
&lt;h2 id=&quot;delegating-design-inside-a-larger-task&quot;&gt;Delegating design inside a larger task&lt;/h2&gt;
&lt;h3 id=&quot;manus&quot;&gt;Manus&lt;/h3&gt;
&lt;p&gt;Disclosure: Manus is an IMG.LY customer. This is also the one entry we did not run ourselves, so unlike the five above it rests on public sources rather than hands-on testing.&lt;/p&gt;
&lt;p&gt;Manus is a general-purpose autonomous agent with a design workspace inside it. Design is one capability among research, app building, and document production.&lt;/p&gt;
&lt;p&gt;Manus decomposes a goal into subtasks, browses for context, runs for long stretches without supervision, and continues while you are offline. For design work it researches before designing, then produces assets you refine through object-level edits: click an element to change colors, swap a background, or reshape it, without regenerating the whole asset.&lt;/p&gt;
&lt;p&gt;Output covers images, video, 3D assets, and presentations. Its public documentation does not describe a layered source file of the kind a dedicated design tool exposes, so you edit objects rather than a design file. General agents are closing the gap on design output faster than anything else in this survey.&lt;/p&gt;
&lt;p&gt;The free tier includes 300 daily refresh credits. Paid plans run $20, $40, and $200 monthly for 4,000, 8,000, and 40,000 credits. Reviews describe unpredictable credit burn on complex tasks.&lt;/p&gt;
&lt;p&gt;Best fit is founders and operators who want one agent handling research, documents, and adequate design output, and who value autonomy over design-tool depth.&lt;/p&gt;
&lt;h2 id=&quot;which-tools-leave-you-a-real-editor&quot;&gt;Which tools leave you a real editor&lt;/h2&gt;
&lt;p&gt;Three tools in this survey give you a real editor after generation. Figma Agent puts native layers on the Figma canvas. Canva’s assistant produces fully editable Canva designs. CoDesign returns a structured scene that opens in a full editor. In all three you select an element and change it, instead of rewriting the prompt and hoping the next version keeps the parts you liked.&lt;/p&gt;
&lt;p&gt;Lovart sits between. It has a canvas and in-place editing, and precision work still tends to finish in Photoshop. Stitch hands off to Figma or to code. Manus edits at the object level, which suits an agent whose remit is much wider than design.&lt;/p&gt;
&lt;p&gt;Revision takes longer than the first draft, and tools that keep you in control of the file absorb those cycles.&lt;/p&gt;
&lt;h2 id=&quot;destination-app-or-capability-in-your-stack&quot;&gt;Destination app or capability in your stack&lt;/h2&gt;
&lt;p&gt;Five of the six tools here are destinations. You open Figma, Canva, Lovart, Stitch, or Manus, do the work there, and take the result away. That model is familiar and it works.&lt;/p&gt;
&lt;p&gt;The sixth, CoDesign, is a design engine that installs into the agent you already use, so the design step happens inside your existing workflow.&lt;/p&gt;
&lt;p&gt;That split is narrower than it first appears. Figma opened an MCP server in March 2026, so external coding agents can drive its canvas. Canva shipped an MCP server in February 2026 that exposes design generation inside ChatGPT, Claude, and Microsoft Copilot, and says more than 12 million designs have been created that way. Being reachable from your agent is not unique to CoDesign.&lt;/p&gt;
&lt;p&gt;The tools differ in what they hand back. Drive Figma’s MCP and you get a Figma file, which is useful if your team lives in Figma. Drive Canva’s and you get a Canva design. Drive CoDesign and you get a portable scene file, PDF with CMYK and bleed, PSD and IDML import, and no platform you have to keep an account with. For teams whose output ends at a printer or inside their own product, that portability decides it. It matters much less if your work already lives in Figma or Canva.&lt;/p&gt;
&lt;h2 id=&quot;how-to-choose&quot;&gt;How to choose&lt;/h2&gt;
&lt;p&gt;Start with what you are making. UI and app screens point to Stitch for speed or the Figma Agent for depth. Campaign and brand assets point to Lovart for volume from one brief, Canva for teams already there, or &lt;a href=&quot;https://img.ly/codesign/&quot;&gt;CoDesign&lt;/a&gt; for output that has to stay editable and reach print. If design is a side task inside something larger, Manus covers it.&lt;/p&gt;
&lt;p&gt;Then check the exit path before you commit, because it is the part you cannot change later. Ask where the file lives, whether you can edit it without regenerating, and what happens when it needs to leave the tool.&lt;/p&gt;
&lt;h2 id=&quot;frequently-asked-questions&quot;&gt;Frequently asked questions&lt;/h2&gt;
&lt;h3 id=&quot;what-is-an-ai-design-agent&quot;&gt;What is an AI design agent?&lt;/h3&gt;
&lt;p&gt;An AI design agent plans and executes multi-step design work from a goal, producing usable design output such as files, screens, or code. &lt;a href=&quot;https://img.ly/blog/what-is-a-design-agent/&quot;&gt;What Is an AI Design Agent?&lt;/a&gt; goes into where the line falls.&lt;/p&gt;
&lt;h3 id=&quot;how-is-an-ai-design-agent-different-from-an-image-generator&quot;&gt;How is an AI design agent different from an image generator?&lt;/h3&gt;
&lt;p&gt;An image generator returns a picture for a prompt. An agent plans a sequence of steps, performs design operations, and produces work you can carry forward.&lt;/p&gt;
&lt;h3 id=&quot;are-ai-design-agents-free&quot;&gt;Are AI design agents free?&lt;/h3&gt;
&lt;p&gt;Three of the six have a free way in: Stitch while it is in Labs, Manus on its free tier, and CoDesign’s local server. Figma’s agent is free during beta but needs a paid seat. Lovart and full Canva use are subscriptions.&lt;/p&gt;
&lt;h3 id=&quot;do-ai-design-agents-replace-designers&quot;&gt;Do AI design agents replace designers?&lt;/h3&gt;
&lt;p&gt;No. Every tool in this survey keeps a person reviewing the output, and all six advertise editing after generation.&lt;/p&gt;
&lt;h3 id=&quot;which-ai-design-agents-produce-editable-files&quot;&gt;Which AI design agents produce editable files?&lt;/h3&gt;
&lt;p&gt;Figma Agent, Canva’s assistant, and CoDesign keep element-level editability in a design file. Lovart offers in-canvas editing, though external finishing is common. Manus edits objects in place. Stitch relies on its Figma and code exports.&lt;/p&gt;
&lt;h3 id=&quot;can-i-use-a-design-agent-inside-claude-or-chatgpt&quot;&gt;Can I use a design agent inside Claude or ChatGPT?&lt;/h3&gt;
&lt;p&gt;Yes, through MCP. Canva exposes design generation in ChatGPT, Claude, and Copilot. Figma opened its MCP server to external coding agents. CoDesign runs as a local MCP server in coding agents such as Claude Code and Codex.&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="/_astro/hero.DBk-VO6j.webp" medium="image"/><category>AI</category><category>Insights</category></item><item><title>Introducing IMG.LY AI Benchmarks</title><link>https://img.ly/blog/introducing-imgly-ai-benchmarks/</link><guid isPermaLink="true">https://img.ly/blog/introducing-imgly-ai-benchmarks/</guid><description>We ran 15 image models through 37 production design prompts and published every score, every prompt and every generated image. Why we built it, why the scores are weighted per job, and the result that surprised us most.</description><pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/&quot;&gt;IMG.LY GenAI Benchmarks&lt;/a&gt; is live. 15 text-to-image models, 37 prompts taken from real design jobs, three fixed seeds each, 1,662 generated images, roughly 7,400 scores. Every prompt, every score, every generated image and the full &lt;a href=&quot;https://img.ly/ai-benchmarks/methodology/&quot;&gt;methodology&lt;/a&gt; are public.&lt;/p&gt;
&lt;h2 id=&quot;why-we-ran-it&quot;&gt;Why we ran it&lt;/h2&gt;
&lt;p&gt;Our customers embed image generation in products that ship physical, commercial output. Web-to-print storefronts, merch platforms, campaign tooling, template systems. When they asked us which model to use, the honest answer was a shrug and a link to a leaderboard that measures something else.&lt;/p&gt;
&lt;p&gt;The existing arenas rank aesthetic preference by crowd vote. That’s a signal, sure, but not the one most relevant to these use cases. They want to know whether the image has a real alpha channel, or whether users got the hex they asked for and whether a word is spelled right. A 10,000-piece personalization run needs to hold the character across every variant. Those are the failures that cost you a reprint or a support ticket, and none of them show up in a preference ranking.&lt;/p&gt;
&lt;p&gt;So we built the one that answers our customers’ question. Which model should you build on for production design work.&lt;/p&gt;
&lt;h2 id=&quot;why-some-of-it-is-judged-not-measured&quot;&gt;Why some of it is judged, not measured&lt;/h2&gt;
&lt;p&gt;Plenty of what matters here is measurable, and we measure it. Native output size, wall-clock latency and price per image come straight off the run. Transparency is a file-level check for a real alpha channel rather than a painted-on white rectangle. Color accuracy is the CIEDE2000 distance between the hex the prompt required and the color that came back.&lt;/p&gt;
&lt;p&gt;The rest needs a judgment call. Whether an image satisfies “three blue cubes stacked on the left and one red sphere on the right” is a checklist, and something has to look at the picture and tick it off. Whether a banner leaves a clean area for a headline has a right answer, but not a computable one.&lt;/p&gt;
&lt;p&gt;For those criteria a vision model does the judging, against a checklist written per prompt before the run. We added two rules to make the VLM’s judgment more transparent and control for arbitrariness. Every judgment records a written rationale beside the score, so you can read why an image lost a point and tell us we’re wrong. The criterion that really does need a person, design-readiness, the “would a professional ship this with under five minutes of cleanup” question, is surfaced to a human in the loop (that would be yours truly, the author).&lt;/p&gt;
&lt;p&gt;The rest of the protocol is the boring part that makes the numbers worth anything. Identical prompts for every model, default parameters, fixed seeds, no per-model prompt tuning ever, a versioned suite, and published scores that are never quietly restated. Partner models get no special treatment, and the Gateway links stay well away from the rankings.&lt;/p&gt;
&lt;h2 id=&quot;why-we-split-it-by-use-case&quot;&gt;Why we split it by use case&lt;/h2&gt;
&lt;p&gt;Average our eight criteria with equal weight and &lt;a href=&quot;https://img.ly/ai-benchmarks/models/nano-banana-2-lite/&quot;&gt;Nano Banana 2 Lite&lt;/a&gt; leads at 3.87 out of 5, while last place, at 2.55, goes to &lt;a href=&quot;https://img.ly/ai-benchmarks/models/nano-banana-pro/&quot;&gt;Nano Banana Pro&lt;/a&gt;, the most expensive flagship in the set. That’s not a data error. Nano Banana Pro has the best brand-color fidelity in the field and renders every required string in the suite exactly. It also costs 28 times more per image than the cheapest model and takes 23 seconds. An unweighted average treats all of that as equally important, so it buries the model with the best output.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;Cost per image plotted against the unweighted overall score for all 15 models. The cheap, fast models sit high on the left; the most expensive flagships sit low on the right. Stable Diffusion 1.5, FLUX.2 and Nano Banana 2 Lite form the Pareto frontier.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1200px) 1200px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1200&quot; height=&quot;660&quot; src=&quot;https://img.ly/_astro/cost-vs-quality.B8vY-NEd_2iGGAu.svg&quot; srcset=&quot;/_astro/cost-vs-quality.B8vY-NEd_4yLYC.svg 640w, /_astro/cost-vs-quality.B8vY-NEd_ZmGGij.svg 750w, /_astro/cost-vs-quality.B8vY-NEd_Z2txY4l.svg 828w, /_astro/cost-vs-quality.B8vY-NEd_KI4Im.svg 1080w, /_astro/cost-vs-quality.B8vY-NEd_2iGGAu.svg 1200w&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;This is the chart everyone asks for, and its vertical axis is the number you should not rank on. Explore the &lt;a href=&quot;https://img.ly/ai-benchmarks/models/&quot;&gt;live, interactive version&lt;/a&gt;, where every point links to that model’s scores.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Which is what leaderboards do. So we publish the same measurements &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/print-files/&quot;&gt;per use case&lt;/a&gt; as well, with a weight matrix per job and the reasoning behind each matrix written out. Print files weight resolution and color heavily, because a 1024px generation is a 3.4-inch print and an out-of-gamut red is a reprint. Personalization at scale weights cost and adherence, because unit economics decide the run. Merch weights the alpha channel above everything else, because a sticker without a clean cutout isn’t a sticker.&lt;/p&gt;
&lt;p&gt;Same numbers, different question, different answer. That’s why the section opens by asking what you’re building instead of handing you a ranked list.&lt;/p&gt;
&lt;h2 id=&quot;the-result-that-surprised-us-most&quot;&gt;The result that surprised us most&lt;/h2&gt;
&lt;p&gt;Reweighting moves models the length of the table.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/gpt-image-1-5/&quot;&gt;GPT Image 1.5&lt;/a&gt; sits twelfth of fifteen on the unweighted blend, dragged down by cost and latency. Reweight for &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/merch-and-stickers/&quot;&gt;merch and stickers&lt;/a&gt;, where transparency carries 35 percent of the matrix, and it wins outright at 3.72, nearly a full point clear of second place. It’s the only model in the run that reliably returns a genuine alpha channel. Thirteen of the fifteen return none at all, and a few paint a fake checkerboard into the pixels, which looks like transparency right up until it reaches a printer.&lt;/p&gt;
&lt;p&gt;On a general leaderboard, the best model for one of our customers’ most common jobs reads as a mid-table also-ran. That’s the argument for the whole project.&lt;/p&gt;
&lt;p&gt;The typography column surprised us a second time. Ideogram has the strongest public reputation for rendering text, and in our run it ranks twelfth of fifteen on measured text accuracy, behind several general-purpose flagships that now render every required string in the suite exactly. Route typographic work to a specialist on reputation alone and you can end up with worse type than the model you already call by default.&lt;/p&gt;
&lt;p&gt;Two more capability gaps got written up as standalone &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/&quot;&gt;findings&lt;/a&gt;. &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/brand-color/&quot;&gt;Brand color&lt;/a&gt;, where the best model manages 3.67 out of 5 and most of the field sits below 3. And &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/consistency/&quot;&gt;consistency&lt;/a&gt;, where no model holds a described character across three scenes and the ceiling is 3.89.&lt;/p&gt;
&lt;h2 id=&quot;what-this-means-for-an-editor&quot;&gt;What this means for an editor&lt;/h2&gt;
&lt;p&gt;We build editors, so our conclusion isn’t neutral. The data still points where it points.&lt;/p&gt;
&lt;p&gt;Read the four findings together and they describe the same problem four times. The cut-out is missing, the color is close but not the color in the brand book, the headline is right until it isn’t, and the character drifts between scenes. None of these are bugs waiting on the next model release. They’re what generation is, a probabilistic first draft. Getting from that draft to something a customer can order takes a handful of deterministic corrections, and somebody has to be able to make them.&lt;/p&gt;
&lt;p&gt;That’s an editor’s job. Background removal and edge cleanup close the transparency gap. Brand kits and exact color values close the color gap. Editable text layers turn a wrong character into a two-second fix instead of a regeneration. Reusable assets and templates carry identity that regeneration won’t. Our &lt;a href=&quot;https://img.ly/use-cases/ai-editor/&quot;&gt;AI Editor&lt;/a&gt; gives your users those controls over whatever model produced the image, and the &lt;a href=&quot;https://img.ly/ai-gateway/&quot;&gt;AI Gateway&lt;/a&gt; makes routing per job, which the data says you should be doing, a config change instead of another integration.&lt;/p&gt;
&lt;p&gt;We’d rather make that case with numbers anyone can check, including the ones that make models we partner with look bad.&lt;/p&gt;
&lt;h2 id=&quot;caveats-and-whats-next&quot;&gt;Caveats, and what’s next&lt;/h2&gt;
&lt;p&gt;This is a pilot dataset. Quality criteria are judged by a vision model, not yet by the blind expert panel, and design-readiness isn’t scored at all. The weight matrices are provisional. When the frozen suite lands, scores restate once and the suite version changes; after that, nothing gets restated without a clearly labeled new methodology version.&lt;/p&gt;
&lt;p&gt;The roster will grow. Image editing and video are the obvious next modalities, and we plan to re-run the suite whenever a notable model ships.&lt;/p&gt;
&lt;p&gt;Go &lt;a href=&quot;https://img.ly/ai-benchmarks/&quot;&gt;pick a use case&lt;/a&gt; and see which model wins your job. If the numbers disagree with your own experience, the prompts and the raw images are all sitting there, and we’d like to hear about it.&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="/_astro/hero.CyGIWvSf.webp" medium="image"/><category>AI</category><category>Image Gen</category><category>Insights</category></item><item><title>The Best GenAI Models for Web-to-Print: A Buyer&apos;s Guide</title><link>https://img.ly/blog/best-generative-ai-models-for-web-to-print/</link><guid isPermaLink="true">https://img.ly/blog/best-generative-ai-models-for-web-to-print/</guid><description>The best general image model can still fail at print. What decides it (vector output, print resolution, CMYK behavior, legible text) is exactly what general model rankings ignore. This guide gives you eleven print-specific evaluation criteria, a scoring rubric, and a model-by-model read on where each one wins.</description><pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;No single generative model is “best” for web-to-print. The right pick depends on the specific job, and print has jobs that screen-first leaderboards don’t measure. A model that tops a general image-quality benchmark can still be useless for print if it can’t render legible text, can’t output vector, or shifts hard out of CMYK gamut.&lt;/p&gt;
&lt;p&gt;This guide gives you the criteria that actually matter for print, a scoring rubric to apply them, and a model-by-model read on where each one wins. The goal is to route each task to the model that fits it, which is why a &lt;strong&gt;model-agnostic&lt;/strong&gt; (“bring your own model”) integration matters more than any one provider’s marketing.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How to read this guide.&lt;/strong&gt; The criteria and rubric are the durable part. The model roster moves monthly: treat the specific model notes as a snapshot, and re-run the rubric against current versions before you commit. The per-model benchmark scores come from IMG.LY’s &lt;a href=&quot;https://img.ly/ai-benchmarks/&quot;&gt;GenAI benchmark suite&lt;/a&gt;, which is now live: a pilot dataset you can explore today, re-run whenever new models drop. The gated PDF report that packages the weighted per-use-case totals and the full test set is the piece we’re opening to early-access subscribers first (sign-up at the end of this guide).&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;why-print-needs-its-own-criteria&quot;&gt;Why print needs its own criteria&lt;/h2&gt;
&lt;p&gt;A web-to-print product &lt;strong&gt;sells a physical object&lt;/strong&gt;. That changes what “good” means:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The output is measured at 300 DPI on a substrate, not at 72 PPI on a retina screen.&lt;/li&gt;
&lt;li&gt;Color is CMYK and spot, not sRGB.&lt;/li&gt;
&lt;li&gt;Logos and type must scale and stay crisp: vector, not pixels.&lt;/li&gt;
&lt;li&gt;The customer is a non-designer, so the model has to be controllable enough to stay inside a template.&lt;/li&gt;
&lt;li&gt;And because it’s a commercial product sold to a customer, the &lt;strong&gt;licensing&lt;/strong&gt; of the output carries real legal exposure.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;General benchmarks (LMArena-style human preference, aesthetic scores) miss most of this. The eleven criteria below are built around it.&lt;/p&gt;
&lt;h2 id=&quot;the-eleven-evaluation-criteria&quot;&gt;The eleven evaluation criteria&lt;/h2&gt;
&lt;h3 id=&quot;1-output-type-and-format-raster-vs-vector&quot;&gt;1. Output type and format: raster vs. vector&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters for print:&lt;/strong&gt; Logos, icons, line art, and type need to scale to any size and print crisply. Raster can’t; vector can. A model that outputs (or can be cleanly traced to) &lt;strong&gt;native SVG / vector&lt;/strong&gt; is uniquely valuable for the print-specific elements that screen apps don’t care about.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Can it produce vector natively? If not, how cleanly does its output vectorize? Does the vector survive into PDF/X as real paths?&lt;/p&gt;
&lt;h3 id=&quot;2-resolution-and-maximum-dimensions&quot;&gt;2. Resolution and maximum dimensions&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; A ~1024px generation is ~3.4” at 300 DPI. Large-format and even A4 need far more. Native max resolution, and how well the model holds detail when upscaled, determines whether you can print the output at size without a quality penalty.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Native max output dimensions; detail retention at 4×/8× upscale; effective printable size at 300 DPI.&lt;/p&gt;
&lt;h3 id=&quot;3-visual-quality-and-fidelity&quot;&gt;3. Visual quality and fidelity&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; The baseline. Realism, coherence, absence of artifacts, anatomical and structural correctness. Crucially, how it holds up under print’s unforgiving close inspection: banding, mushy detail, and plasticky textures show more on paper than on screen.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Side-by-side preference on print-representative prompts; artifact rate; detail at print scale.&lt;/p&gt;
&lt;h3 id=&quot;4-text-and-typography-rendering&quot;&gt;4. Text and typography rendering&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Print is full of words, and most models garble text inside images. For anything where text-in-image is unavoidable (a generated poster, a label motif), legibility and kerning are make-or-break. Best practice is still to set real type as a separate layer, but some jobs need it baked in. Our &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/text-reliability/&quot;&gt;text-reliability finding&lt;/a&gt; measures exactly this. The top of the field has largely solved Latin headlines: three models rendered every required string in our suite exactly. Small type and non-Latin scripts have not been solved, and the model most famous for typography ranks twelfth of fifteen. One wrong character still means regenerating the whole image, which is the real argument for keeping type on its own layer. Compare every model on the &lt;a href=&quot;https://img.ly/ai-benchmarks/prompts/t04-wordmark/&quot;&gt;same wordmark prompt&lt;/a&gt; to see the spread.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Legibility of short and long strings; correct spelling; kerning; multi-line; non-Latin scripts.&lt;/p&gt;
&lt;h3 id=&quot;5-prompt-adherence-and-controllability&quot;&gt;5. Prompt adherence and controllability&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Non-designers need predictable results, and templates need the output to land where it’s told. Controllability spans prompt adherence, &lt;strong&gt;image-to-image / inpainting / outpainting&lt;/strong&gt;, style and reference-image conditioning, and seed stability for repeatable results.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Does it honor constraints? Quality of edits (generative fill, expand, object removal)? Reference-image fidelity? Same-seed reproducibility?&lt;/p&gt;
&lt;h3 id=&quot;6-color-accuracy-and-cmyk-readiness&quot;&gt;6. Color accuracy and CMYK-readiness&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Vivid RGB outputs clip hard in CMYK and shift on press. Models that stay closer to printable gamut, and outputs that convert predictably, mean fewer surprises and less soft-proofing friction. Even before the CMYK conversion, hitting an exact brand hex in RGB is hard: our &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/brand-color/&quot;&gt;brand-color finding&lt;/a&gt; measures color accuracy with CIEDE2000 against the required hex, and no model reliably lands the color in your brand book.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Gamut coverage vs. CMYK; shift magnitude after ICC conversion; behavior on brand-critical colors and near-whites/blacks.&lt;/p&gt;
&lt;h3 id=&quot;7-transparency-and-background-handling&quot;&gt;7. Transparency and background handling&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Merch, stickers, die-cuts, and product compositing need clean subjects on transparent backgrounds and accurate edges (hair, glass, fine detail). Native transparent output and clean cutout edges save a manual masking step. In the models themselves this is close to unsolved: our &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/transparency/&quot;&gt;transparency finding&lt;/a&gt; shows 13 of 15 tested models emit no real alpha channel at all, which is exactly why a background-removal step belongs in the pipeline rather than a prayer that the model returns a clean cutout.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Native alpha output; edge quality on hard cases; cutout-path readiness.&lt;/p&gt;
&lt;h3 id=&quot;8-responsiveness--latency&quot;&gt;8. Responsiveness / latency&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Two very different bars. In an &lt;strong&gt;interactive editor&lt;/strong&gt;, anything over a few seconds breaks flow. In a &lt;strong&gt;&lt;a href=&quot;https://img.ly/use-cases/variable-data-printing/&quot;&gt;batch/VDP run&lt;/a&gt;&lt;/strong&gt;, throughput and concurrency matter more than per-image speed. The right model differs by context.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; P50/P95 latency at your resolution; throughput under concurrency; cold-start behavior.&lt;/p&gt;
&lt;h3 id=&quot;9-cost-per-generation&quot;&gt;9. Cost per generation&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Unit economics. A few cents per image is invisible for interactive one-offs and brutal across a 100k-recipient VDP run. Price has to be weighed against quality &lt;em&gt;for the specific job&lt;/em&gt;. You don’t pay flagship prices for a draft thumbnail.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Cost per image at your resolution/step settings; cost of edit vs. generate; volume pricing.&lt;/p&gt;
&lt;h3 id=&quot;10-consistency-and-repeatability&quot;&gt;10. Consistency and repeatability&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; Brand work needs the same input to yield the same (or controllably similar) output, and product likenesses must not drift. Seed control, character and style consistency, and low variance separate “brand-safe” models from “slot-machine” ones. Our &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/consistency/&quot;&gt;consistency finding&lt;/a&gt; puts numbers on the drift: no model holds a described character across scenes, so identity has to come from templates and editing, not from a regeneration lottery.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Variance across runs at fixed seed; character/style consistency across a set; drift on iterative edits.&lt;/p&gt;
&lt;h3 id=&quot;11-commercial-licensing-and-ip-safety&quot;&gt;11. Commercial licensing and IP safety&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Why it matters:&lt;/strong&gt; You’re selling the printed output. Commercial-use rights, training-data provenance, indemnification, and content/safety filtering are legal exposure, not fine print. A model that’s brilliant but legally murky is a non-starter for a product customers resell.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What to test:&lt;/strong&gt; Commercial-use terms; indemnity; provenance/transparency; enterprise and data-handling terms; content-filter behavior.&lt;/p&gt;
&lt;h2 id=&quot;the-scoring-rubric&quot;&gt;The scoring rubric&lt;/h2&gt;
&lt;p&gt;Score each model &lt;strong&gt;1–5&lt;/strong&gt; on each criterion, then weight by your use case. Suggested weights for the three most common web-to-print contexts:&lt;/p&gt;













































































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Criterion&lt;/th&gt;&lt;th&gt;Interactive design tool&lt;/th&gt;&lt;th&gt;High-volume VDP / automation&lt;/th&gt;&lt;th&gt;Merch / POD&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;1. Vector / SVG output&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;Medium&lt;/td&gt;&lt;td&gt;Medium&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;2. Resolution / max size&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;3. Visual quality&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;Medium&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;4. Text rendering&lt;/td&gt;&lt;td&gt;Medium&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;Medium&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;5. Controllability / editing&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;6. CMYK-readiness&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;7. Transparency / cutout&lt;/td&gt;&lt;td&gt;Medium&lt;/td&gt;&lt;td&gt;Low&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;8. Responsiveness&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;Medium (throughput)&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;9. Cost per generation&lt;/td&gt;&lt;td&gt;Medium&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;Medium&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;10. Consistency&lt;/td&gt;&lt;td&gt;Medium&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;11. Licensing / IP&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;td&gt;High&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Scoring scale:&lt;/strong&gt; 1 = unusable for print · 2 = weak · 3 = workable with mitigation · 4 = strong · 5 = best-in-class.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Drop your own numbers into this rubric. The measured scores in the &lt;a href=&quot;https://img.ly/blog/best-generative-ai-models-for-web-to-print//#the-imgly-genai-benchmark-measured-results&quot;&gt;results table later in this guide&lt;/a&gt; give you a starting hypothesis to test; the qualitative notes are not a substitute for your own runs on your prompts.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;models-by-job&quot;&gt;Models by job&lt;/h2&gt;
&lt;p&gt;New models ship faster than any ranking can keep up with, so this section is organized &lt;strong&gt;by job&lt;/strong&gt;, the way you actually route requests. Models are drawn from the roster currently wired into CE.SDK’s AI plugins.&lt;/p&gt;
&lt;h3 id=&quot;job-a-text-to-image-generate-a-new-asset&quot;&gt;Job A: Text-to-image (generate a new asset)&lt;/h3&gt;
&lt;p&gt;Candidates: &lt;strong&gt;Recraft V3&lt;/strong&gt;, &lt;strong&gt;Recraft 20B&lt;/strong&gt;, &lt;strong&gt;Seedream V4&lt;/strong&gt;, &lt;strong&gt;Nano Banana&lt;/strong&gt;, &lt;strong&gt;Nano Banana Pro&lt;/strong&gt;, &lt;strong&gt;GPT Image 1&lt;/strong&gt;, &lt;strong&gt;Ideogram V3&lt;/strong&gt;.&lt;/p&gt;















































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Stand-out for print&lt;/th&gt;&lt;th&gt;Watch-outs&lt;/th&gt;&lt;th&gt;Benchmark&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Recraft V3&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;The print specialist: &lt;strong&gt;native vector/SVG generation&lt;/strong&gt;, strong text rendering, brand-style controls. Often the single most print-relevant text-to-image model.&lt;/td&gt;&lt;td&gt;Verify max raster resolution for large-format jobs.&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/recraft-v3/&quot;&gt;3.29 / 5&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Recraft 20B&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Faster, cheaper Recraft tier. Good for interactive iteration before a final Recraft V3 render.&lt;/td&gt;&lt;td&gt;Quality step-down vs. V3.&lt;/td&gt;&lt;td&gt;not yet in the suite&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Seedream V4&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;High visual fidelity and resolution; strong general-purpose hero imagery.&lt;/td&gt;&lt;td&gt;Text rendering and vector are not its strength.&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/seedream-4-5/&quot;&gt;3.80 / 5&lt;/a&gt; (measured on Seedream 4.5)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Nano Banana / Pro&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Fast, controllable, strong prompt adherence; Pro for higher quality. Good interactive default.&lt;/td&gt;&lt;td&gt;Confirm commercial-licensing terms for resale.&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/nano-banana-2/&quot;&gt;3.24 / 5&lt;/a&gt; · Pro &lt;a href=&quot;https://img.ly/ai-benchmarks/models/nano-banana-pro/&quot;&gt;2.55 / 5&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;GPT Image 1&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Strong instruction-following and &lt;strong&gt;comparatively reliable in-image text&lt;/strong&gt;; good for layout-aware generation.&lt;/td&gt;&lt;td&gt;Latency and cost on the higher side for interactive use.&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/gpt-image-1-5/&quot;&gt;3.07 / 5&lt;/a&gt; (measured on GPT Image 1.5)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Ideogram V3&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Built around typography and still a credible pick for poster and typographic art, but our measured text column no longer puts it on top.&lt;/td&gt;&lt;td&gt;Ranks twelfth of fifteen on text accuracy in our run.&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/ideogram-v3/&quot;&gt;2.97 / 5&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Routing heuristic:&lt;/strong&gt; logos, icons, and typographic art go to &lt;strong&gt;Recraft V3&lt;/strong&gt; (vector; the &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/print-files/&quot;&gt;print files use case&lt;/a&gt;). Words-in-image goes to whichever flagship currently tops the measured text column, which in our July 2026 run means &lt;strong&gt;GPT Image 1.5&lt;/strong&gt;, &lt;strong&gt;FLUX.2 [pro]&lt;/strong&gt; or &lt;strong&gt;Nano Banana Pro&lt;/strong&gt; rather than a typography specialist (the &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/text-heavy-designs/&quot;&gt;text-heavy designs use case&lt;/a&gt;). Photoreal hero imagery goes to &lt;strong&gt;Seedream V4&lt;/strong&gt; or &lt;strong&gt;Nano Banana Pro&lt;/strong&gt; (the &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/product-ads/&quot;&gt;product ads use case&lt;/a&gt;). Fast interactive drafts go to &lt;strong&gt;Recraft 20B / Nano Banana&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id=&quot;job-b-image-editing-adapt-a-customers-asset&quot;&gt;Job B: Image editing (adapt a customer’s asset)&lt;/h3&gt;
&lt;p&gt;Candidates: &lt;strong&gt;Flux Pro Kontext&lt;/strong&gt;, &lt;strong&gt;Flux Pro Kontext Max&lt;/strong&gt;, &lt;strong&gt;Nano Banana Edit&lt;/strong&gt;, &lt;strong&gt;Nano Banana Pro Edit&lt;/strong&gt;, &lt;strong&gt;Qwen Image Edit&lt;/strong&gt;, &lt;strong&gt;Gemini Flash Edit&lt;/strong&gt;, &lt;strong&gt;Seedream V4 Edit&lt;/strong&gt;, &lt;strong&gt;GPT Image 1&lt;/strong&gt;, &lt;strong&gt;Ideogram V3 Remix&lt;/strong&gt;.&lt;/p&gt;









































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Stand-out for print&lt;/th&gt;&lt;th&gt;Watch-outs&lt;/th&gt;&lt;th&gt;Benchmark&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Flux Pro Kontext / Max&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Strong &lt;strong&gt;instruction-based editing&lt;/strong&gt; with high subject preservation: change one thing without wrecking the rest. Good for generative fill/expand and targeted edits.&lt;/td&gt;&lt;td&gt;Max tier costs more; budget per edit.&lt;/td&gt;&lt;td&gt;editing suite planned&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Nano Banana Edit / Pro Edit&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Fast, controllable edits; good interactive default for object removal and swaps.&lt;/td&gt;&lt;td&gt;Verify edge quality on fine detail.&lt;/td&gt;&lt;td&gt;editing suite planned&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Qwen Image Edit&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Strong editing with good text handling during edits.&lt;/td&gt;&lt;td&gt;Validate commercial terms.&lt;/td&gt;&lt;td&gt;editing suite planned&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Gemini Flash Edit&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Low latency&lt;/strong&gt;: the responsiveness pick for interactive editing.&lt;/td&gt;&lt;td&gt;Quality trade-off vs. heavier models.&lt;/td&gt;&lt;td&gt;editing suite planned&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Seedream V4 Edit&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;High-fidelity edits matching its generation quality.&lt;/td&gt;&lt;td&gt;Heavier; weigh latency.&lt;/td&gt;&lt;td&gt;editing suite planned&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The benchmark currently measures text-to-image only; the editing models above join the leaderboard when the frozen suite adds an editing track, so their cells read “editing suite planned” rather than a score.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Routing heuristic:&lt;/strong&gt; precise “change X, keep everything else” goes to &lt;strong&gt;Flux Kontext&lt;/strong&gt;. Fast interactive edits go to &lt;strong&gt;Gemini Flash Edit / Nano Banana Edit&lt;/strong&gt;. Edits that must preserve in-image text go to &lt;strong&gt;Qwen Image Edit&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id=&quot;job-c-specialized-print-operations&quot;&gt;Job C: Specialized print operations&lt;/h3&gt;
&lt;p&gt;These aren’t single “models” so much as task pipelines, but they belong in the routing table:&lt;/p&gt;



































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Task&lt;/th&gt;&lt;th&gt;Approach / model&lt;/th&gt;&lt;th&gt;Note&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Vectorize (raster → SVG)&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Vectorize plugin / Recraft vector&lt;/td&gt;&lt;td&gt;The print-critical one: protects logos and PDF/X vector output.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Background removal&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;@imgly/background-removal&lt;/code&gt; (in-browser)&lt;/td&gt;&lt;td&gt;On-device, instant, privacy-preserving. Feeds cutout paths.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Image correction &amp;#x26; DPI&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Perfectly Clear plugin + DPI validation&lt;/td&gt;&lt;td&gt;Auto-corrects exposure, color, sharpness, and noise; validation flags undersized images. Prefer high-res native generation over upscaling.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Text generation / copy&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Claude Sonnet 4.6&lt;/strong&gt;, &lt;strong&gt;GPT-4.1 Nano&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Headlines, rewrites, per-recipient VDP copy.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Text-to-speech / sound&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;ElevenLabs (where relevant)&lt;/td&gt;&lt;td&gt;Less common in print; relevant for multi-channel campaigns.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;h2 id=&quot;the-quick-reference-picks&quot;&gt;The quick-reference picks&lt;/h2&gt;
&lt;p&gt;If you skip the rubric and just want today’s defaults, this is the snapshot we’d start from. Treat it as the hypothesis your own benchmarks confirm or overturn, not a verdict.&lt;/p&gt;


















































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Print job&lt;/th&gt;&lt;th&gt;Pick&lt;/th&gt;&lt;th&gt;Why&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Logos, icons, line art, anything that must scale&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Recraft V3&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;The only major model with native SVG/vector output. Everything else hands you pixels to trace. Vector survives into PDF/X as real paths and prints sharp at any size.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Legible text inside the image (posters, labels, packaging motifs)&lt;/td&gt;&lt;td&gt;&lt;strong&gt;GPT Image 1.5&lt;/strong&gt;, &lt;strong&gt;FLUX.2 [pro]&lt;/strong&gt;, &lt;strong&gt;Nano Banana Pro&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;These three rendered every required string in our suite exactly. The typography specialists did not: Ideogram ranks twelfth of fifteen on the measured column. Still, bake text in only when you must: real type as a separate editor layer beats all of them.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Photoreal hero imagery at print size&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Seedream 4.5&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Resolution is the print bottleneck, and it is the one criterion where the field genuinely splits. At default parameters Seedream returns 2048px (about 6.8” at 300 DPI); most of the field returns 1024px, which is 3.4” and a bet on the upscaler.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Transparent cut-outs for stickers, die-cuts, merch&lt;/td&gt;&lt;td&gt;&lt;strong&gt;GPT Image 1.5&lt;/strong&gt;, and a background-removal step regardless&lt;/td&gt;&lt;td&gt;It is the only model in our run that reliably emits a real alpha channel. Thirteen of fifteen emit none at all, so the cut-out has to be a pipeline step, not a prompt.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;”Change X, keep everything else” edits&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Flux Kontext Pro / Max&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Best-in-class instruction-following edits with subject preservation, which is critical when the asset is the customer’s product photo and “the AI improved it” means a reprint. The Nano Banana family is the consistency pick for variants of one subject.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Interactive drafts (speed and cost)&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Recraft 20B&lt;/strong&gt;, &lt;strong&gt;Gemini Flash Edit&lt;/strong&gt;, base &lt;strong&gt;Nano Banana&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;In an editor, a 20-second generation is a bounce. Draft cheap and fast, then re-render the final with the heavyweight. The customer never needs to know two models were involved.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Data sovereignty / self-hosting&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Qwen-Image / Qwen Image Edit&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Apache 2.0 open weights. For European operators and government buyers who won’t send customer uploads to a US API, “we can run it ourselves” beats a quality delta.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;VDP copy and headlines&lt;/td&gt;&lt;td&gt;&lt;strong&gt;Claude Sonnet 4.6&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Cheap enough per recipient at 100k-mailer volume, and strong at length-constrained rewriting: the “fit this text frame” problem.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Three of those rationales carry more weight than the rest.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vector is the criterion that decides printability.&lt;/strong&gt; Screen products never need vector, so general models never optimized for it, and Recraft sits almost alone on the one measure that determines whether a logo prints cleanly at size. If you adopt a single routing rule from this guide, make it this one.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Resolution and color are pipeline problems the model can only start.&lt;/strong&gt; No model outputs CMYK; every one generates sRGB. “CMYK-readiness” is really about how hard a model’s palette clips on conversion, and your ICC and soft-proofing pipeline owns the rest. Same with DPI: pair every generation path with validation and a dedicated upscaler rather than trusting native output. You can swap the model whenever you want; the print-correctness layer around it has to stay put.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Transparency is not a model feature you can shop for.&lt;/strong&gt; When we asked every model for a transparent background and measured the alpha channel that came back, thirteen of fifteen returned none, and a few painted a checkerboard into the pixels instead. If your product sells stickers, die-cuts, or anything composited onto a garment, budget for background removal and edge cleanup as a permanent pipeline stage. Choosing a different model does not remove that stage.&lt;/p&gt;
&lt;p&gt;One roster note: &lt;strong&gt;Adobe Firefly&lt;/strong&gt; isn’t wired into CE.SDK’s plugin roster today, but for licensing-sensitive enterprises its trained-on-licensed-content and indemnification story is the strongest answer to criterion 11. Include it in your own evaluation if criterion 11 dominates your weighting.&lt;/p&gt;
&lt;h2 id=&quot;a-worked-decision-which-model-wins-for-your-product&quot;&gt;A worked decision: which model wins for &lt;em&gt;your&lt;/em&gt; product?&lt;/h2&gt;
&lt;p&gt;The rubric resolves to a different answer per context. Three quick reads:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://img.ly/products/creative-sdk/&quot;&gt;Interactive design tool&lt;/a&gt; (non-designers personalizing templates).&lt;/strong&gt; Weight responsiveness, controllability, and vector. Likely stack: &lt;strong&gt;Recraft 20B / Nano Banana&lt;/strong&gt; for fast generation, &lt;strong&gt;Recraft V3&lt;/strong&gt; for final vector, &lt;strong&gt;Gemini Flash Edit&lt;/strong&gt; for interactive edits. The customer iterates fast and the final asset is print-clean. See the &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/in-app-ugc/&quot;&gt;in-app UGC use case&lt;/a&gt; for how the benchmark weights this context.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://img.ly/use-cases/variable-data-printing/&quot;&gt;High-volume VDP / automation&lt;/a&gt; (100k personalized mailers).&lt;/strong&gt; Weight cost, throughput, consistency, and licensing. Likely stack: a cost-efficient generation model at high concurrency, &lt;strong&gt;Claude Sonnet 4.6&lt;/strong&gt; for per-recipient copy, seed-locked for repeatability, all run headless. Per-unit cost dominates the decision. See the &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/personalization-at-scale/&quot;&gt;personalization-at-scale use case&lt;/a&gt; for the weighted read.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://img.ly/use-cases/ai-editor/print-on-demand/&quot;&gt;Merch / POD&lt;/a&gt; (user art on physical products).&lt;/strong&gt; Weight transparency/cutout, resolution, and quality. Likely stack: &lt;strong&gt;background removal&lt;/strong&gt; plus cutout for subjects, &lt;strong&gt;upscaling&lt;/strong&gt; for resolution, &lt;strong&gt;Seedream V4 / Nano Banana Pro&lt;/strong&gt; for generated designs, with hard licensing checks because the output is resold. See the &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/merch-and-stickers/&quot;&gt;merch and stickers use case&lt;/a&gt;, where the transparency gap decides the stack.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In every case, the conclusion is the same: &lt;strong&gt;no one model wins all three.&lt;/strong&gt; The advantage is in being able to route each task to the model that wins it, and to swap models as better ones appear, without re-architecting your editor.&lt;/p&gt;
&lt;h2 id=&quot;what-makes-per-job-routing-practical-native-plugins--the-ai-gateway&quot;&gt;What makes per-job routing practical: native plugins + the AI Gateway&lt;/h2&gt;
&lt;p&gt;A criteria-driven, route-per-job strategy only works if switching models is cheap. If wiring in a new model means a new API integration, new auth, new key management, new error handling, and new billing plumbing every time, you’ll quietly standardize on whatever you integrated first, and inherit its weaknesses on every job it loses.&lt;/p&gt;
&lt;p&gt;CE.SDK is built to avoid that:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Every image model in this guide is natively supported as an AI plugin.&lt;/strong&gt; Text-to-image and image editing run as drop-in providers inside the editor, not bespoke integrations you build and maintain. Background removal, vectorization, and image correction (&lt;a href=&quot;https://img.ly/demos/perfectlyclear-plugin/web/&quot;&gt;Perfectly Clear&lt;/a&gt;) ship as their own plugins, and text generation routes through built-in Anthropic and OpenAI providers. Adopting Recraft for vector, Ideogram for text, and Flux Kontext for edits is &lt;a href=&quot;https://img.ly/demos/ai-editor/&quot;&gt;configuration, not three separate engineering projects&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The &lt;a href=&quot;https://img.ly/ai-gateway/&quot;&gt;AI Gateway&lt;/a&gt; gives you one unified surface to every model.&lt;/strong&gt; Instead of stitching together a dozen provider APIs, key schemes, and rate-limit behaviors, you call models through a single managed layer. That’s what turns “route each task to the best model” from an architecture diagram into a one-line config change, and what lets you swap a model the day a better one ships, without touching your editor or your print pipeline.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Outputs land inside the print pipeline, not beside it.&lt;/strong&gt; Whatever model produces the asset, it flows into locked templates with bleed, safe-area, and DPI constraints, then through CMYK / spot-color / PDF/X-3 export. The model is interchangeable; print correctness is not. Print-on-demand company &lt;a href=&quot;https://img.ly/case-studies/print-bar/&quot;&gt;The Print Bar&lt;/a&gt; and direct-mail platform &lt;a href=&quot;https://img.ly/case-studies/postbuddy/&quot;&gt;PostBuddy&lt;/a&gt; both built on this pipeline, embedding CE.SDK rather than maintaining their own print editor.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;So the “best model” question is no longer a one-time, high-stakes bet on a single provider; it’s an ongoing optimization. The benchmark scores below tell you &lt;em&gt;which&lt;/em&gt; model to route each job to today, and the native-plugin and Gateway architecture is what lets you act on that answer and revisit it when the scores change next quarter.&lt;/p&gt;
&lt;h2 id=&quot;the-imgly-genai-benchmark-measured-results&quot;&gt;The IMG.LY GenAI benchmark: measured results&lt;/h2&gt;
&lt;p&gt;IMG.LY’s GenAI benchmark now scores these models on print-representative prompts, and the results are live and explorable at &lt;a href=&quot;https://img.ly/ai-benchmarks/&quot;&gt;img.ly/ai-benchmarks&lt;/a&gt;. The table below is the top-line read for the text-to-image models in this guide: the overall score out of 5, the four print-critical criteria (each out of 5), list price per image, and median (p50) latency. All numbers are from the pilot-0 dataset (July 2026); scores are sorted by overall.&lt;/p&gt;







































































































































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Overall&lt;/th&gt;&lt;th&gt;Text&lt;/th&gt;&lt;th&gt;Color&lt;/th&gt;&lt;th&gt;Transparency&lt;/th&gt;&lt;th&gt;Resolution&lt;/th&gt;&lt;th&gt;$ / image&lt;/th&gt;&lt;th&gt;p50&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/nano-banana-2-lite/&quot;&gt;Nano Banana 2 Lite&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.87 / 5&lt;/td&gt;&lt;td&gt;4.9&lt;/td&gt;&lt;td&gt;2.9&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;3.0&lt;/td&gt;&lt;td&gt;$0.02&lt;/td&gt;&lt;td&gt;4.0s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/seedream-4-5/&quot;&gt;Seedream 4.5&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.80 / 5&lt;/td&gt;&lt;td&gt;4.9&lt;/td&gt;&lt;td&gt;2.9&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;5.0&lt;/td&gt;&lt;td&gt;$0.048&lt;/td&gt;&lt;td&gt;12.4s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/flux-2/&quot;&gt;FLUX.2&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.64 / 5&lt;/td&gt;&lt;td&gt;4.9&lt;/td&gt;&lt;td&gt;2.8&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;2.0&lt;/td&gt;&lt;td&gt;$0.013&lt;/td&gt;&lt;td&gt;2.0s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/flux-2-turbo/&quot;&gt;FLUX.2 Turbo&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.61 / 5&lt;/td&gt;&lt;td&gt;4.9&lt;/td&gt;&lt;td&gt;2.7&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;2.0&lt;/td&gt;&lt;td&gt;$0.015&lt;/td&gt;&lt;td&gt;2.0s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/gemini-25-flash-image/&quot;&gt;Gemini 2.5 Flash Image&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.60 / 5&lt;/td&gt;&lt;td&gt;5.0&lt;/td&gt;&lt;td&gt;1.7&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;3.0&lt;/td&gt;&lt;td&gt;$0.039&lt;/td&gt;&lt;td&gt;7.4s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/recraft-v3/&quot;&gt;Recraft V3&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.29 / 5&lt;/td&gt;&lt;td&gt;4.0&lt;/td&gt;&lt;td&gt;1.9&lt;/td&gt;&lt;td&gt;2.0&lt;/td&gt;&lt;td&gt;3.0&lt;/td&gt;&lt;td&gt;$0.04&lt;/td&gt;&lt;td&gt;7.4s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/nano-banana-2/&quot;&gt;Nano Banana 2&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.24 / 5&lt;/td&gt;&lt;td&gt;5.0&lt;/td&gt;&lt;td&gt;3.2&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;3.0&lt;/td&gt;&lt;td&gt;$0.10&lt;/td&gt;&lt;td&gt;13.2s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/flux-2-pro/&quot;&gt;FLUX.2 Pro&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.23 / 5&lt;/td&gt;&lt;td&gt;5.0&lt;/td&gt;&lt;td&gt;2.3&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;2.0&lt;/td&gt;&lt;td&gt;$0.04&lt;/td&gt;&lt;td&gt;11.5s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/qwen-image/&quot;&gt;Qwen-Image&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.21 / 5&lt;/td&gt;&lt;td&gt;4.8&lt;/td&gt;&lt;td&gt;2.0&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;2.0&lt;/td&gt;&lt;td&gt;$0.03&lt;/td&gt;&lt;td&gt;6.9s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/gpt-image-1-5/&quot;&gt;GPT Image 1.5&lt;/a&gt;&lt;/td&gt;&lt;td&gt;3.07 / 5&lt;/td&gt;&lt;td&gt;5.0&lt;/td&gt;&lt;td&gt;3.4&lt;/td&gt;&lt;td&gt;4.2&lt;/td&gt;&lt;td&gt;3.0&lt;/td&gt;&lt;td&gt;$0.12&lt;/td&gt;&lt;td&gt;34.0s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/ideogram-v3/&quot;&gt;Ideogram 3.0&lt;/a&gt;&lt;/td&gt;&lt;td&gt;2.97 / 5&lt;/td&gt;&lt;td&gt;4.7&lt;/td&gt;&lt;td&gt;2.8&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;3.0&lt;/td&gt;&lt;td&gt;$0.06&lt;/td&gt;&lt;td&gt;17.9s&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;a href=&quot;https://img.ly/ai-benchmarks/models/nano-banana-pro/&quot;&gt;Nano Banana Pro&lt;/a&gt;&lt;/td&gt;&lt;td&gt;2.55 / 5&lt;/td&gt;&lt;td&gt;5.0&lt;/td&gt;&lt;td&gt;3.7&lt;/td&gt;&lt;td&gt;0.0&lt;/td&gt;&lt;td&gt;3.0&lt;/td&gt;&lt;td&gt;$0.36&lt;/td&gt;&lt;td&gt;23.3s&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Measured on the pilot-0 suite (July 2026): 37 prompts times 3 seeds per model, every image scored 0 to 5 per criterion by an automated vision judge. Two models carry newer versions than the guide names them by: Seedream V4 is measured as Seedream 4.5, and GPT Image 1 as GPT Image 1.5. The suite also covers Seedream 5.0 Lite, Luma Photon, and Stable Diffusion 1.5, which this guide’s roster does not name; their scores are on the &lt;a href=&quot;https://img.ly/ai-benchmarks/models/&quot;&gt;rankings page&lt;/a&gt;. The overall column is an unweighted blend, so read it alongside the per-use-case weighting, not as a single leaderboard: the cheap fast model tops the raw average while the priciest lands last. Full methodology is at &lt;a href=&quot;https://img.ly/ai-benchmarks/methodology/&quot;&gt;img.ly/ai-benchmarks/methodology&lt;/a&gt;, and every model, prompt, and generated image is explorable at &lt;a href=&quot;https://img.ly/ai-benchmarks/&quot;&gt;img.ly/ai-benchmarks&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;A few things the numbers make plain, and that the criteria above predicted. Transparency is close to unsolved: 13 of 15 tested models emit no real alpha channel at all (&lt;a href=&quot;https://img.ly/ai-benchmarks/findings/transparency/&quot;&gt;the transparency finding&lt;/a&gt;), and GPT Image 1.5 is the only one that scores meaningfully on it. Brand color is broadly weak: measured with CIEDE2000, no model reliably hits an exact hex (&lt;a href=&quot;https://img.ly/ai-benchmarks/findings/brand-color/&quot;&gt;the brand-color finding&lt;/a&gt;). Text is the one column where the field mostly clears the bar: on our per-required-string measure several general-purpose flagships render every string in the suite exactly, while the model with the biggest reputation for text, Ideogram, actually lands near the bottom of the column, not the top (&lt;a href=&quot;https://img.ly/ai-benchmarks/findings/text-reliability/&quot;&gt;the text-reliability finding&lt;/a&gt;). The gaps open exactly where print work gets hard. Averaged across the field, the neon sign in Japanese scores 3.91 and the dense packaging label 4.40, against 4.91 for a greeting-card cover. Latin headlines are close to solved; small type and non-Latin scripts are not.&lt;/p&gt;
&lt;h3 id=&quot;the-same-numbers-weighted-for-print&quot;&gt;The same numbers, weighted for print&lt;/h3&gt;
&lt;p&gt;The overall column is the wrong lens for a print product, which is the whole reason the benchmark carries use-case weights. Reweighted for print files (resolution 35%, color 30%, text 15%, transparency 10%), &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/print-files/&quot;&gt;Seedream 4.5 leads at 3.64&lt;/a&gt;, ahead of Seedream 5.0 Lite at 3.50 and GPT Image 1.5 at 3.45: native output size at default parameters does most of the work, and the two models that return 2048px start the job at twice the printable width of the 1024px field.&lt;/p&gt;
&lt;p&gt;Reweight again for &lt;a href=&quot;https://img.ly/ai-benchmarks/use-cases/merch-and-stickers/&quot;&gt;merch and stickers&lt;/a&gt;, where alpha carries 35%, and the order breaks completely: GPT Image 1.5 wins at 3.72, nearly a full point ahead of Recraft V3 at 2.78, because it is the only model that reliably returns a real alpha channel. Same measurements, different job, different answer. That inversion is the argument for routing per job rather than standardizing on one provider.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The gated report is coming to subscribers first.&lt;/strong&gt; &lt;em&gt;AI Image Models for Print Production: The Benchmark Report&lt;/em&gt; packages the weighted per-use-case totals, the full test set, and the print-specific analysis into one PDF. &lt;a href=&quot;https://share.hsforms.com/1IgAOV1wASXGPnFG4ZPLejg1hk3i&quot;&gt;Subscribe to get it before it is public&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The criteria and the routing logic above stand on their own. The scored numbers are what the suite adds: a measured, repeatable answer for each cell, refreshed as models change. For why the benchmark is built this way, and what else the run turned up, see &lt;a href=&quot;https://img.ly/blog/introducing-imgly-ai-benchmarks/&quot;&gt;&lt;strong&gt;Introducing IMG.LY AI Benchmarks&lt;/strong&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;bottom-line&quot;&gt;Bottom line&lt;/h2&gt;
&lt;p&gt;Pick criteria before you pick models. For &lt;a href=&quot;https://img.ly/use-cases/web-to-print-design-tool/&quot;&gt;web-to-print&lt;/a&gt;, the criteria that screen benchmarks ignore (vector output, print resolution, CMYK behavior, text legibility, licensing) are exactly the ones that decide whether a beautiful generation is a sellable print. Score the current models against those criteria for &lt;em&gt;your&lt;/em&gt; context, route each job to the model that wins it, and run it all through CE.SDK’s &lt;a href=&quot;https://img.ly/use-cases/ai-editor/&quot;&gt;native AI plugins&lt;/a&gt; and &lt;a href=&quot;https://img.ly/ai-gateway/&quot;&gt;AI Gateway&lt;/a&gt; so re-routing tomorrow, after the next leaderboard reshuffle, is a config change rather than a rebuild.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Companion piece: &lt;a href=&quot;https://img.ly/blog/how-to-leverage-generative-ai-in-web-to-print/&quot;&gt;&lt;strong&gt;How to Leverage Generative AI in Web-to-Print&lt;/strong&gt;&lt;/a&gt; covers the use cases, integration patterns, and print-specific pitfalls behind these model choices.&lt;/em&gt;&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="/_astro/cover.ZLW7wLRy.webp" medium="image"/><category>AI</category><category>Image Gen</category><category>Web-to-Print</category><category>Print</category></item><item><title>How to Leverage Generative AI in Web-to-Print</title><link>https://img.ly/blog/how-to-leverage-generative-ai-in-web-to-print/</link><guid isPermaLink="true">https://img.ly/blog/how-to-leverage-generative-ai-in-web-to-print/</guid><description>Generative AI changes the economics of web-to-print by doing the part customers were never qualified to do: generating images, fixing resolution, removing backgrounds, vectorizing logos. Here are the use cases worth building, and the print-specific realities to design around.</description><pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Web-to-print sells self-service: your customer designs a postcard, a flyer, a t-shirt, a packaging label in the browser and orders it without a designer in the loop. For years, though, the self-service stopped at the hard part. Someone still had to supply a usable image. Someone still had to know what 300 DPI meant. When they didn’t, the file landed on a prepress desk to be fixed by hand.&lt;/p&gt;
&lt;p&gt;One operator we spoke to put it bluntly: his designers spend roughly 40% of their time fixing customer artwork rather than creating templates. Tom Rowe of &lt;a href=&quot;https://img.ly/case-studies/print-bar/&quot;&gt;The Print Bar&lt;/a&gt; described what that looks like at the design step: &lt;em&gt;“Our brand didn’t match our experience. You could see it in our bounce rates.”&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Generative AI changes the economics of that hard part. It can turn a prompt into a usable asset, lift a 72-DPI logo to something printable, and strip a background in the browser before a human ever sees the file, removing the specific friction that kept self-service from being self-service in the first place.&lt;/p&gt;
&lt;p&gt;It can also go wrong: aimed carelessly, the same models produce beautiful screen images that fall apart on press. This post is about the difference between the two: the use cases worth building, and the print-specific realities you have to design around.&lt;/p&gt;
&lt;h2 id=&quot;the-shift-from-upload-a-print-ready-file-to-describe-what-you-want&quot;&gt;The shift: from “upload a print-ready file” to “describe what you want”&lt;/h2&gt;
&lt;p&gt;The old &lt;a href=&quot;https://img.ly/use-cases/web-to-print-design-tool/&quot;&gt;web-to-print&lt;/a&gt; contract asked the customer to arrive with a finished, technically correct asset. That filter excluded most of the market: the CRM manager, the franchise owner, the small-business buyer who &lt;em&gt;“knows Canva because they make their wedding invitations on it”&lt;/em&gt; but has never opened InDesign and never will.&lt;/p&gt;
&lt;p&gt;Generative AI moves the burden of production off the customer. Instead of &lt;em&gt;“upload a 300 DPI CMYK image with 3mm bleed,”&lt;/em&gt; the ask becomes &lt;em&gt;“type what you want on the card.”&lt;/em&gt; The editor, not the customer and not your prepress team, becomes responsible for turning intent into a production-safe file.&lt;/p&gt;
&lt;p&gt;That principle runs through everything below: &lt;strong&gt;AI is most valuable in web-to-print where it removes a step the customer was never qualified to do.&lt;/strong&gt; The closer a use case sits to it, the more it returns.&lt;/p&gt;
&lt;h2 id=&quot;the-use-cases-worth-building&quot;&gt;The use cases worth building&lt;/h2&gt;
&lt;h3 id=&quot;1-generate-on-brand-imagery-from-a-prompt-text-to-image&quot;&gt;1. Generate on-brand imagery from a prompt (text-to-image)&lt;/h3&gt;
&lt;p&gt;This is the most obvious win. A customer needs a hero image, a background, a seasonal motif, or a product scene, and doesn’t have one. &lt;a href=&quot;https://img.ly/demos/ai-editor/&quot;&gt;Text-to-image generation&lt;/a&gt; lets them describe it and get options in seconds, without a stock-photo license hunt or a design request ticket.&lt;/p&gt;
&lt;p&gt;In a print context the point is &lt;strong&gt;unblocking the order&lt;/strong&gt;, not novelty for its own sake. The customer who would have abandoned at “I don’t have a good image” now keeps going. Models like Recraft V3, Seedream V4, Ideogram V3, and the Nano Banana family each have different strengths here (covered in depth in the companion guide, &lt;em&gt;The Best GenAI Models for Web-to-Print&lt;/em&gt;), but the integration pattern is the same: generation feeds a placeholder inside a locked template, so the output lands inside your bleed, safe-area, and brand constraints automatically.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; generated images are usually delivered at screen resolution, often around 1024px. That’s fine for a business card photo, marginal for an A4 flyer, and unusable for large-format. Pair generation with upscaling (below) and size validation before you let it reach export. Native output sizes vary widely by model: the &lt;a href=&quot;https://img.ly/ai-benchmarks/models/&quot;&gt;measured native output in the model rankings&lt;/a&gt; (pilot-0 dataset, July 2026) shows some models topping out near 1024px while others generate at 2K.&lt;/p&gt;
&lt;h3 id=&quot;2-adapt-and-extend-customer-supplied-images-image-to-image-generative-fillexpand&quot;&gt;2. Adapt and extend customer-supplied images (image-to-image, generative fill/expand)&lt;/h3&gt;
&lt;p&gt;Customers rarely arrive with an asset that fits your canvas. It’s the wrong aspect ratio, it has the wrong background, or it’s a portrait crop where you need a landscape banner. Image-to-image editing and &lt;strong&gt;generative expand&lt;/strong&gt; (outpainting) solve the most common case: extending an image to fill bleed or a different SKU’s dimensions instead of stretching or letterboxing it.&lt;/p&gt;
&lt;p&gt;This is quietly one of the highest-value print use cases, because aspect-ratio and bleed mismatches are a top source of prepress rework. Generative fill also handles object removal (“take the coffee cup off the desk”), background swaps, and clean-up: the edits a non-designer can’t do in any tool they own.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; outpainting invents pixels. On a brand asset or a product photo, “invented” can mean “wrong.” Keep generative expand to backgrounds and ambient areas. Never let it reconstruct a logo, a face, or a product the customer is actually selling.&lt;/p&gt;
&lt;h3 id=&quot;3-background-removal-and-cutouts-for-mockups-and-merch&quot;&gt;3. Background removal and cutouts for mockups and merch&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/demos/background-removal/&quot;&gt;Background removal&lt;/a&gt; is the workhorse, and it has to be a pipeline step because the models won’t do it for you: our &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/transparency/&quot;&gt;transparency finding&lt;/a&gt; shows 13 of 15 tested models emit no real alpha channel at all. For merch and print-on-demand it’s the difference between a customer’s snapshot and a clean subject that drops onto a t-shirt, mug, or sticker. IMG.LY runs this &lt;strong&gt;in the browser&lt;/strong&gt; (via the open-source &lt;code&gt;@imgly/background-removal&lt;/code&gt; package), which matters for two reasons: it’s instant enough for an interactive editor, and the customer’s image never has to leave the device.&lt;/p&gt;
&lt;p&gt;Paired with the &lt;strong&gt;&lt;a href=&quot;https://img.ly/demos/cutout-lines/&quot;&gt;cutout plugin&lt;/a&gt;&lt;/strong&gt;, background removal also feeds the literal cut lines for die-cut stickers, labels, and packaging, turning a removed background into a production path.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; browser-based removal is excellent on clear subjects and struggles on hair, glass, and fine fringes. For a print product that will be inspected up close, expose a manual refine step rather than trusting the mask blindly.&lt;/p&gt;
&lt;h3 id=&quot;4-vectorize-raster-art-into-print-scalable-graphics&quot;&gt;4. Vectorize raster art into print-scalable graphics&lt;/h3&gt;
&lt;p&gt;This is the use case print people care about most and screen-first builders forget. A customer uploads a logo as a small, jagged PNG. Printed at size, it’s mush. &lt;strong&gt;&lt;a href=&quot;https://img.ly/demos/vectorizer-plugin/&quot;&gt;Vectorization&lt;/a&gt;&lt;/strong&gt; (AI raster-to-SVG) traces it into clean, resolution-independent paths that scale to any output size and reproduce crisply on press.&lt;/p&gt;
&lt;p&gt;Vectorize is also how you get logos and simple graphics into a state your &lt;strong&gt;&lt;a href=&quot;https://img.ly/demos/export-print-ready-pdf/&quot;&gt;PDF/X pipeline&lt;/a&gt;&lt;/strong&gt; can preserve as vector, instead of rasterizing everything and throwing away the scalability you vectorized for. Recraft V3 is notable here because it can generate &lt;strong&gt;natively vector&lt;/strong&gt; output rather than only tracing after the fact.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; vectorization is great for logos, icons, and flat art. It’s the wrong tool for a photograph. Detect the asset type and route photos to upscaling, not tracing.&lt;/p&gt;
&lt;h3 id=&quot;5-resolution-and-image-quality-upscaling-correction-and-dpi-repair&quot;&gt;5. Resolution and image quality: upscaling, correction, and DPI repair&lt;/h3&gt;
&lt;p&gt;The single most common web-to-print failure is a low-resolution image: the 72-DPI photo that looks fine on screen and prints as a blurry mess. Two different AI operations fix it, and it pays to keep them straight, because they solve different problems.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Upscaling (super-resolution)&lt;/strong&gt; raises the actual pixel count, reconstructing detail to lift a small image toward a printable size. It’s a generative step you route to a dedicated model, the same way you route a text-to-image prompt; the strong upscalers can take a 1024px generation to 4K. Reach for it when the source is simply too small for its placed size.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Image correction&lt;/strong&gt; is the other half, and it ships in CE.SDK today as the &lt;a href=&quot;https://img.ly/demos/perfectlyclear-plugin/web/&quot;&gt;Perfectly Clear plugin&lt;/a&gt;. It auto-corrects exposure, contrast, color, tint, sharpness, and noise in a single pass. It won’t add pixels, but it gets the most printable result out of the pixels you have, which is what a dim, soft phone photo usually needs more than raw resolution.&lt;/p&gt;
&lt;p&gt;Tie both to your &lt;strong&gt;DPI validation&lt;/strong&gt;. When the editor detects an image below your minimum threshold for its placed size, offer to correct or upscale it instead of just blocking export. You convert a dead end into a one-click fix, the kind of &lt;em&gt;“they don’t need to know what DPI is”&lt;/em&gt; experience operators are chasing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; upscaling fabricates detail. It rescues marginal images; it can’t conjure a sharp 4-megapixel product shot from a thumbnail, and correction can’t add resolution at all. Set honest thresholds, and still warn when the source is hopeless.&lt;/p&gt;
&lt;h3 id=&quot;6-copy-and-variable-text-generation-text--vdp-at-scale&quot;&gt;6. Copy and variable text generation (text + VDP at scale)&lt;/h3&gt;
&lt;p&gt;Generative text earns its place in two spots. First, as a writing aid inside the editor: headline options, a tagline, a “make it shorter / more formal / fit this space” rewrite for the non-writer staring at an empty text box. Second, and more powerfully for print, in &lt;strong&gt;&lt;a href=&quot;https://img.ly/use-cases/variable-data-printing/&quot;&gt;Variable Data Printing&lt;/a&gt;&lt;/strong&gt;: generating or localizing per-recipient copy across thousands of personalized postcards, mailers, or labels from a single template.&lt;/p&gt;
&lt;p&gt;VDP is the highest-leverage place to use generative text: generate the variants, merge them into print-ready files in headless mode, and run the batch.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; generated copy needs guardrails. Length limits so it doesn’t overflow the text frame, tone and brand constraints, and a human-review gate for anything legally sensitive: pricing, claims, regulated industries.&lt;/p&gt;
&lt;h3 id=&quot;7-template-adaptation-and-auto-resize-across-skus&quot;&gt;7. Template adaptation and auto-resize across SKUs&lt;/h3&gt;
&lt;p&gt;A print catalog is the same design across many sizes and substrates: the same campaign as a postcard, a flyer, a poster, a social tile. &lt;a href=&quot;https://img.ly/demos/automated-resizing/&quot;&gt;AI-assisted resize and re-layout&lt;/a&gt; adapt one master design to every SKU’s dimensions, reflowing content intelligently instead of scaling it blindly. Combined with locked templates, you expand SKUs without re-authoring each one.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; reflow decisions still need brand rules. Lock what must stay fixed (logo size, safe area, mandatory legal text) so the AI rearranges within your constraints, not over them.&lt;/p&gt;
&lt;h3 id=&quot;8-guided-prepress-correction&quot;&gt;8. Guided prepress correction&lt;/h3&gt;
&lt;p&gt;Instead of rejecting a bad file after the fact, use AI to catch and &lt;em&gt;fix&lt;/em&gt; problems inside the editor: flag the low-res image and offer to upscale it, &lt;a href=&quot;https://img.ly/demos/design-validation/&quot;&gt;detect text outside the safe area&lt;/a&gt; and nudge it in, notice a near-white “white” that won’t print and correct it. The same problems your prepress team used to fix by hand get caught here, before the order is placed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; auto-correction must be transparent and reversible. Show the customer what changed and let them undo it. Silent “helpful” edits to someone’s artwork erode trust fast.&lt;/p&gt;
&lt;h3 id=&quot;9-localization-and-market-variants&quot;&gt;9. Localization and market variants&lt;/h3&gt;
&lt;p&gt;For franchise networks and multi-market brands, &lt;a href=&quot;https://img.ly/demos/language/&quot;&gt;AI translation&lt;/a&gt; plus regeneration of localized imagery and copy turns one approved master into 12 market versions, each print-ready and on-brand, without a translation agency or 12 design tickets. A 200-location franchise can ship 200 localized flyers with the same brand guideline enforced by the editor on every one.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; machine translation in regulated or legal copy needs human sign-off, and text expansion (German runs roughly 30% longer than English) will break tight layouts unless your template frames can flex.&lt;/p&gt;
&lt;h3 id=&quot;10-realistic-product-previews&quot;&gt;10. Realistic product previews&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/demos/mockup-editor/&quot;&gt;Generative and 3D-assisted mockups&lt;/a&gt; show the design &lt;strong&gt;on the actual product&lt;/strong&gt;: the shirt, the mug, the folded brochure, the box, not an abstract artboard. Operators consistently tie this to higher conversion and fewer “will this look right?” support tickets, because the preview matches what arrives on the doorstep.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Watch out:&lt;/strong&gt; a preview that’s prettier than the print sets up a disappointed customer and a reprint. Calibrate mockups to real output, including substrate color and finish.&lt;/p&gt;
&lt;h2 id=&quot;the-print-specific-realities&quot;&gt;The print-specific realities&lt;/h2&gt;
&lt;p&gt;Most generative models are built for screens. Print adds requirements the web doesn’t have, and the ones below are where projects most often break.&lt;/p&gt;
&lt;h3 id=&quot;resolution-and-dpi&quot;&gt;Resolution and DPI&lt;/h3&gt;
&lt;p&gt;Generated images typically arrive at around 1024px. At 300 DPI that’s about a 3.4-inch image: fine for a business card, not for a poster. &lt;strong&gt;Always validate output resolution against placed size&lt;/strong&gt;, and route undersized assets through upscaling before export. Treat raw generation resolution as a starting point, not a deliverable. Native output size is one of the axes we measure: see the &lt;a href=&quot;https://img.ly/ai-benchmarks/models/&quot;&gt;measured native output per model&lt;/a&gt; to know which models start closer to print size and which need the most upscaling.&lt;/p&gt;
&lt;h3 id=&quot;color-srgb-in-cmyk-out&quot;&gt;Color: sRGB in, CMYK out&lt;/h3&gt;
&lt;p&gt;Models generate in RGB. Print is CMYK, plus spot colors. The vivid blues, greens, and oranges that look great on screen sit outside the CMYK gamut and will shift on press. And even inside RGB, hitting an exact brand hex is hard: our &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/brand-color/&quot;&gt;brand-color finding&lt;/a&gt; measures color accuracy with CIEDE2000 and no model reliably lands the color in your brand book, which is one more reason correction belongs on the canvas. Your pipeline, not the model, owns this: convert with the right ICC profile, validate against the print provider’s color requirements, and preserve &lt;strong&gt;named spot colors&lt;/strong&gt; through to the PDF/X. Show the customer a soft-proof so the shift isn’t a surprise on delivery.&lt;/p&gt;
&lt;h3 id=&quot;vector-vs-raster-and-the-everything-got-rastered-trap&quot;&gt;Vector vs. raster (and the “everything got rastered” trap)&lt;/h3&gt;
&lt;p&gt;Most models output raster. For logos, type, and line art that need to scale and print crisply, raster is the wrong format, and a PDF that rasterizes all text and vectors is, in one operator’s words, &lt;em&gt;“death for the printer.”&lt;/em&gt; Use vector-native generation (e.g. Recraft) and vectorization for the elements that need it, and make sure your export keeps &lt;strong&gt;vector and text as vector&lt;/strong&gt; with CMYK values preserved (PDF/X-3), not flattened to pixels.&lt;/p&gt;
&lt;h3 id=&quot;text-rendering-inside-images&quot;&gt;Text rendering inside images&lt;/h3&gt;
&lt;p&gt;Generative models have a reputation for garbling text inside images, and our &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/text-reliability/&quot;&gt;text-reliability finding&lt;/a&gt; shows the top of the field has largely fixed it: a handful of flagships now render every required string in our suite exactly. The catch for print is that reputation and measurement no longer line up. Ideogram, the model best known for typography, ranks twelfth of fifteen on the measured column; several general-purpose flagships beat it. Reach for a specialist on reputation and you can end up with worse type than the default model you already route to. Put every model on the &lt;a href=&quot;https://img.ly/ai-benchmarks/prompts/t04-wordmark/&quot;&gt;same wordmark prompt&lt;/a&gt; to see the spread. For anything that must be readable in print, &lt;strong&gt;don’t bake text into a generated image.&lt;/strong&gt; Generate the imagery, then set real, editable, vectorizable type as a separate layer in the editor. If you must bake text in, pick the model from the measured column rather than from the marketing; the companion guide has the current ranking.&lt;/p&gt;
&lt;h3 id=&quot;brand-safety-and-guardrails&quot;&gt;Brand safety and guardrails&lt;/h3&gt;
&lt;p&gt;Let customers generate anything and you lose brand control. Box generation inside locked templates instead: it fills a defined placeholder, within fixed margins, alongside a logo and palette the customer can’t move. A franchisee can change the headline and the photo; the logo size, position, and safe area stay locked. AI expands what the customer can create without expanding what they can break.&lt;/p&gt;
&lt;h3 id=&quot;commercial-and-licensing-rights&quot;&gt;Commercial and licensing rights&lt;/h3&gt;
&lt;p&gt;A web-to-print product is sold. That makes the IP status of generated assets a question you have to answer before you ship. Does your model provider grant commercial-use rights to outputs? Are there indemnities? Does training-data provenance create exposure for a customer reselling the printed product? Pick models and providers whose commercial terms you’ve actually read, and surface usage terms where they matter.&lt;/p&gt;
&lt;h3 id=&quot;cost-and-latency-the-unit-economics&quot;&gt;Cost and latency (the unit economics)&lt;/h3&gt;
&lt;p&gt;Every generation costs money and time. In an interactive editor, a 30-second generation is a broken experience. In a high-volume VDP run, a few cents per asset multiplied by 100,000 recipients is a real line item. The spread is not small: the &lt;a href=&quot;https://img.ly/ai-benchmarks/models/&quot;&gt;measured $/image and p50 latency in the model rankings&lt;/a&gt; run from a fraction of a cent to tens of cents, and from a couple of seconds to over thirty. Match the model to the job: a fast, cheap model for interactive iteration, a higher-quality (slower, pricier) model for the final render, and use the &lt;a href=&quot;https://img.ly/ai-benchmarks/compare/&quot;&gt;compare tool&lt;/a&gt; to weigh the trade-off head-to-head. Cache aggressively. The companion guide scores models on that trade-off.&lt;/p&gt;
&lt;h3 id=&quot;data-privacy-and-sovereignty&quot;&gt;Data privacy and sovereignty&lt;/h3&gt;
&lt;p&gt;Customer-uploaded images can be sensitive: faces, IDs, confidential product designs. Browser-based processing (background removal, some editing) keeps data on-device. API-based generation sends it to a third party, which has GDPR and data-residency implications, especially for European customers and government buyers who &lt;em&gt;“prioritize a European vendor for data sovereignty.”&lt;/em&gt; Be explicit about what runs locally vs. in the cloud, and choose providers accordingly.&lt;/p&gt;
&lt;h3 id=&quot;hallucination-and-consistency&quot;&gt;Hallucination and consistency&lt;/h3&gt;
&lt;p&gt;Generative output is non-deterministic. The same prompt yields different results, and details drift. Our &lt;a href=&quot;https://img.ly/ai-benchmarks/findings/consistency/&quot;&gt;consistency finding&lt;/a&gt; measures the worst of it: no model holds a described character across scenes, so series work needs identity as a reusable asset, not a regeneration lottery. For brand assets, product likenesses, and anything a customer will compare against reality, constrain hard (reference images, seeds, image-to-image rather than free generation) and keep a human gate on the outputs that matter.&lt;/p&gt;
&lt;h2 id=&quot;how-imgly-approaches-it&quot;&gt;How IMG.LY approaches it&lt;/h2&gt;
&lt;p&gt;All of these constraints point the same way: &lt;strong&gt;AI is a step inside a print pipeline, not a feature bolted onto an editor.&lt;/strong&gt; A few principles shape how CE.SDK handles it:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Model-agnostic by design.&lt;/strong&gt; The &lt;a href=&quot;https://img.ly/use-cases/ai-editor/&quot;&gt;AI plugins&lt;/a&gt; connect to any third-party model or API: bring your own model. You’re not locked to one provider’s quality, price, or licensing terms. You route each task to the model that wins it, and swap models as better ones ship every few weeks.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The &lt;a href=&quot;https://img.ly/ai-gateway/&quot;&gt;AI Gateway&lt;/a&gt;&lt;/strong&gt; provides managed access to models for editors, so you’re not stitching together a dozen API integrations and key-management schemes yourself.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Generation lands inside the print pipeline.&lt;/strong&gt; Outputs flow into locked templates with bleed, safe areas, DPI thresholds, and brand constraints already enforced, then through a &lt;strong&gt;CMYK / spot-color / PDF/X-3&lt;/strong&gt; export that keeps vectors and text as vectors. The model never gets to bypass print correctness.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Drop-in generative features.&lt;/strong&gt; Background removal, generative fill, image-to-image, vectorize, and text-to-image are available as plugins, so you adopt the use cases above incrementally rather than rebuilding your editor.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Headless mode&lt;/strong&gt; runs the same generation and export server-side for batch and variable-data runs.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The result is the experience operators actually want to sell: the customer describes what they want, the AI does the part they were never qualified to do, and the file that reaches the press is print-correct because the pipeline enforced it, not because the customer got lucky. For the product-level view of this stack, see &lt;a href=&quot;https://img.ly/use-cases/ai-editor/print-on-demand/&quot;&gt;AI for Print on Demand&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;what-customers-are-doing-with-it&quot;&gt;What customers are doing with it&lt;/h2&gt;
&lt;p&gt;High-volume print and direct-mail platforms are already building on this foundation. IMG.LY powers print and &lt;a href=&quot;https://img.ly/use-cases/print-personalization/&quot;&gt;personalization workflows&lt;/a&gt; for operators including &lt;strong&gt;&lt;a href=&quot;https://img.ly/case-studies/postbuddy/&quot;&gt;Postbuddy&lt;/a&gt;&lt;/strong&gt; (personalized direct mail), &lt;strong&gt;Swiss Post&lt;/strong&gt;, &lt;strong&gt;&lt;a href=&quot;https://img.ly/case-studies/digitas/&quot;&gt;Digitas&lt;/a&gt;&lt;/strong&gt;, and &lt;strong&gt;HP&lt;/strong&gt;, alongside hundreds of smaller print, merch, and franchise platforms.&lt;/p&gt;
&lt;p&gt;The pattern that recurs in those conversations is narrow and specific: AI removed a bottleneck. The prepress team that fixed 40% of incoming files now lets the editor catch problems upstream. The customers who bounced at “I don’t have a good image” finish the order. The franchise network that needed a designer for every local variant generates them inside guardrails. It comes back to the point from the top: AI pays off where it deletes a step the customer couldn’t do and the operator didn’t want to.&lt;/p&gt;
&lt;h2 id=&quot;where-to-start&quot;&gt;Where to start&lt;/h2&gt;
&lt;p&gt;You don’t need all ten use cases on day one. The highest-ROI starting points, in order:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Background removal&lt;/strong&gt;: instant, in-browser, immediately useful for any merch or photo product.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Image correction and DPI validation&lt;/strong&gt;: catches and cleans up weak uploads before they ever reach the printer.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Generative expand / fill&lt;/strong&gt;: kills aspect-ratio and bleed rework.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Text-to-image into locked placeholders&lt;/strong&gt;: unblocks the “I don’t have an image” abandoner.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Vectorize&lt;/strong&gt;: rescues logos and protects your PDF/X output.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Each one removes a documented source of friction or rework. Add them inside locked templates and a real CMYK/PDF/X pipeline, and self-service web-to-print finally works the way it was always sold: the customer really can do it themselves.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Next: which models to actually use. See the companion guide, &lt;a href=&quot;https://img.ly/blog/best-generative-ai-models-for-web-to-print/&quot;&gt;&lt;strong&gt;The Best GenAI Models for Web-to-Print: A Buyer’s Guide&lt;/strong&gt;&lt;/a&gt;, for a criteria-based comparison across output quality, vector/SVG support, text rendering, color, responsiveness, and cost. For the measured evidence behind these pitfalls, see the &lt;a href=&quot;https://img.ly/ai-benchmarks/&quot;&gt;&lt;strong&gt;IMG.LY GenAI Benchmarks&lt;/strong&gt;&lt;/a&gt;: every model on the same prompts, scored per criterion, with the method and headline results written up in &lt;a href=&quot;https://img.ly/blog/introducing-imgly-ai-benchmarks/&quot;&gt;&lt;strong&gt;Introducing IMG.LY AI Benchmarks&lt;/strong&gt;&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="/_astro/cover.CHrckJHN.webp" medium="image"/><category>AI</category><category>Web-to-Print</category><category>Print</category><category>Creative Workflows</category></item><item><title>How 600+ Teams Shipped Creative Editing in 14 Days</title><link>https://img.ly/blog/imgly-impact-report/</link><guid isPermaLink="true">https://img.ly/blog/imgly-impact-report/</guid><description>A data report based on 600+ customers, 28 structured customer interviews, and a self-reported customer survey.</description><pubDate>Tue, 28 Apr 2026 21:15:17 GMT</pubDate><content:encoded>&lt;h2 id=&quot;the-imgly-impact-report&quot;&gt;The IMG.LY Impact Report&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;Based on a self-reported customer survey, 28 structured customer interviews, and IMG.LY’s published customer case studies. External benchmarks cited in context.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;600+ companies&lt;/strong&gt; from startups, over Fortune 500s to government organizations run creative editing on IMG.LY, collectively producing &lt;strong&gt;500 million creations every month&lt;/strong&gt;. The median customer goes from first SDK access to a &lt;strong&gt;working prototype in 2 days&lt;/strong&gt; and a &lt;strong&gt;production-ready feature shipped inside a single 14-day sprint&lt;/strong&gt;. Those self-reported medians run roughly &lt;strong&gt;13× faster&lt;/strong&gt; than the 6-month industry baseline for building an equivalent editor in-house, and around 6× faster than stitching together an open-source toolkit.&lt;/p&gt;
&lt;p&gt;The time-to-market number is the entry ticket. What those 600+ customers do &lt;em&gt;after&lt;/em&gt; shipping, how the editor changes their pricing, funnel, team velocity, and competitive position is what this report is about.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Methodology note.&lt;/strong&gt; Time-to-prototype and time-to-MVP figures are &lt;strong&gt;self-reported by IMG.LY customers&lt;/strong&gt; in a recent customer survey. All named customer outcomes are drawn from &lt;a href=&quot;https://img.ly/case-studies/&quot;&gt;IMG.LY’s published case studies&lt;/a&gt;, where the customer has reviewed and approved the specific figures cited. Quantitative ranges and anonymous quotes come from 30 structured customer interviews (2024–2026); individual customers are not identified.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;what-customers-unlock&quot;&gt;What Customers Unlock&lt;/h2&gt;
&lt;h3 id=&quot;1-ship-in-days-not-quarters&quot;&gt;1. Ship in days, not quarters&lt;/h3&gt;
&lt;p&gt;A working prototype takes an average of two days, a production feature often just a single sprint. Customer interviews put the alternative, an in-house build, at &lt;strong&gt;6–12 months with 2–3 senior developers&lt;/strong&gt;, consistent with the cost bands documented in IMG.LY’s &lt;a href=&quot;https://img.ly/blog/strategic-guide-to-creative-editing-when-to-build-when-to-buy/&quot;&gt;Build vs. Buy guide&lt;/a&gt; and independently with the &lt;a href=&quot;https://thestory.is/en/journal/chaos-report/&quot;&gt;Standish Group CHAOS Report&lt;/a&gt;, which finds only &lt;strong&gt;31% of software projects succeed&lt;/strong&gt; on time, on budget, and with full scope… and under &lt;strong&gt;10%&lt;/strong&gt; for large projects.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/case-studies/plai/&quot;&gt;Plai&lt;/a&gt; went live in about a month where their founder had estimated “months and months” for an in-house build. &lt;a href=&quot;https://img.ly/case-studies/optimizely/&quot;&gt;Optimizely&lt;/a&gt;, serving 1,000+ enterprise customers, described the integration as &lt;em&gt;“one of the quickest projects we’ve worked on.”&lt;/em&gt;&lt;/p&gt;
&lt;div class=&quot;imgly-chart imgly-chart--time&quot;&gt;&lt;p class=&quot;imgly-eyebrow&quot;&gt;Time to a production-ready MVP&lt;/p&gt;&lt;h3 class=&quot;imgly-title&quot;&gt;14 days vs. 6 months: how long it takes to ship creative editing&lt;/h3&gt;&lt;p class=&quot;imgly-subtitle&quot;&gt;Median time from project start to a production-ready creative editor, by integration path.&lt;/p&gt;&lt;div class=&quot;imgly-canvas-wrap&quot;&gt;&lt;canvas id=&quot;imgly-chart-time-canvas&quot;&gt;&lt;/canvas&gt;&lt;/div&gt;&lt;p class=&quot;imgly-footnote&quot;&gt;&lt;strong&gt;Sources.&lt;/strong&gt; CE.SDK figure (14 days) is the self-reported median from an IMG.LY customer survey; individual results vary with scope and team. In-house (180 days) and open-source (90 days) figures are midpoints of the ranges documented in IMG.LY&apos;s &lt;em&gt;Build vs. Buy guide&lt;/em&gt;, cross-referenced with customer interviews (N=28).&lt;/p&gt;&lt;/div&gt;
&lt;h3 id=&quot;2-open-new-pricing-and-market-tiers&quot;&gt;2. Open new pricing and market tiers&lt;/h3&gt;
&lt;p&gt;Embedding an editor changes the customer’s &lt;em&gt;own&lt;/em&gt; unit economics. It unlocks a self-serve tier where users previously needed a designer in the loop, lowers the price floor, and opens markets where manual production was the binding constraint.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/case-studies/plai/&quot;&gt;Plai&lt;/a&gt; doubled annual revenue and now generates &lt;strong&gt;30,000+ ad creatives monthly&lt;/strong&gt; on the platform. &lt;a href=&quot;https://img.ly/case-studies/postbuddy/&quot;&gt;Postbuddy&lt;/a&gt; measured a &lt;strong&gt;4× A/B-test improvement&lt;/strong&gt; after moving direct-mail creation in-app, a lift that funded expansion into Sweden and Norway. In an anonymized interview, one customer described the change plainly: &lt;em&gt;“We can lower our prices for the base product, it’s cheaper for clients who can use self-service.”&lt;/em&gt;&lt;/p&gt;
&lt;h3 id=&quot;3-lift-your-own-conversion-and-engagement&quot;&gt;3. Lift your own conversion and engagement&lt;/h3&gt;
&lt;p&gt;In-context editing keeps users inside the product, shortens time-to-value, and removes the “export to Photoshop” detour that quietly leaks trials. The effect shows up directly in sign-up and trial-to-paid numbers.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/case-studies/omneky/&quot;&gt;Omneky&lt;/a&gt; recorded a &lt;strong&gt;10× month-over-month increase in new sign-ups&lt;/strong&gt; after embedding CE.SDK. In a separate interview, one customer described &lt;strong&gt;400–500 organic trials per month with roughly $65–70K of opportunity evaporating&lt;/strong&gt; at an 8% trial-to-paid rate tied largely to editor friction, the number that made their own build-vs-buy case close itself.&lt;/p&gt;
&lt;h3 id=&quot;4-free-your-design-marketing-and-support-teams&quot;&gt;4. Free your design, marketing, and support teams&lt;/h3&gt;
&lt;p&gt;The cost saved by an editor isn’t only on the engineering line. It shows up in designer queues that disappear, support tickets that drop, and marketing teams that stop waiting on the bottleneck.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/case-studies/imagebank/&quot;&gt;ImageBank X&lt;/a&gt; compressed a recurring asset-prep task from &lt;strong&gt;15 minutes to 2 minutes, an 87% reduction per task&lt;/strong&gt;. &lt;a href=&quot;https://img.ly/case-studies/halio/&quot;&gt;Halio.ai&lt;/a&gt; enables financial advisors to produce &lt;strong&gt;30 days of branded content in 30 minutes&lt;/strong&gt;. Interview customers describe the same pattern in plainer language: &lt;em&gt;“happier marketing teams,”&lt;/em&gt; and, of the post-launch support burden, &lt;em&gt;“the things the sales team had to deal with all day are just gone.”&lt;/em&gt;&lt;/p&gt;
&lt;h3 id=&quot;5-compound-instead-of-freeze&quot;&gt;5. Compound instead of freeze&lt;/h3&gt;
&lt;p&gt;This is where the engineering-economics story earns its keep. An in-house editor is a snapshot: the team ships what’s feasible in the window, then expends time and money just to keep it alive. Atlassian’s &lt;a href=&quot;https://www.atlassian.com/software/compass/resources/state-of-developer-2024&quot;&gt;2024 State of Developer Experience Report&lt;/a&gt; (N=2,100+) found &lt;strong&gt;69% of developers lose eight or more hours per week to inefficiencies (about 20% of their time) with technical debt as the primary cause&lt;/strong&gt;. Stack Overflow’s &lt;a href=&quot;https://survey.stackoverflow.co/2024/&quot;&gt;2024 Developer Survey&lt;/a&gt; (N=65,437) independently found &lt;strong&gt;62% of developers rank technical debt as their single biggest frustration&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;CE.SDK customers don’t pay that tax. New capabilities such as AI background removal, text-to-image, object add/remove, style transfer arrive as modular plugins, not re-architectures. One interviewed customer, who had spent a year building their own editor before switching, described maintaining it as &lt;em&gt;“a full-time job and our main product isn’t editing.”&lt;/em&gt; Every month an in-house team spends on maintenance is a month IMG.LY customers get new capabilities shipped &lt;em&gt;to&lt;/em&gt; them. The feature gap widens, not narrows.&lt;/p&gt;
&lt;div class=&quot;imgly-chart imgly-chart--compound&quot;&gt;&lt;p class=&quot;imgly-eyebrow&quot;&gt;Editor capability over 24 months&lt;/p&gt;&lt;h3 class=&quot;imgly-title&quot;&gt;In-house builds freeze. CE.SDK customers compound.&lt;/h3&gt;&lt;p class=&quot;imgly-subtitle&quot;&gt;An in-house editor ships at the snapshot the team can afford and then plateaus under the weight of maintenance. CE.SDK customers absorb new capabilities (AI background removal, text-to-image, object remove, automatic subtitles) as modular plugins each quarter.&lt;/p&gt;&lt;div class=&quot;imgly-legend&quot;&gt;&lt;div class=&quot;imgly-legend-item&quot;&gt;&lt;span class=&quot;imgly-legend-line&quot;&gt;&lt;/span&gt;Build in-house&lt;/div&gt;&lt;div class=&quot;imgly-legend-item&quot;&gt;&lt;span class=&quot;imgly-legend-line&quot;&gt;&lt;/span&gt;CE.SDK customer&lt;/div&gt;&lt;/div&gt;&lt;div class=&quot;imgly-canvas-wrap&quot;&gt;&lt;canvas id=&quot;imgly-chart-compound-canvas&quot;&gt;&lt;/canvas&gt;&lt;/div&gt;&lt;p class=&quot;imgly-footnote&quot;&gt;&lt;strong&gt;Illustrative.&lt;/strong&gt; The CE.SDK trajectory reflects the cadence of major capability releases shipped to customers without re-architecture. The in-house trajectory reflects the typical build-then-maintain pattern documented in customer interviews and corroborated by the &lt;a href=&quot;https://www.atlassian.com/software/compass/resources/state-of-developer-2024&quot;&gt;Atlassian 2024 State of Developer Experience Report&lt;/a&gt; (69% of developers lose 8+ hours per week to inefficiencies, primarily technical debt).&lt;/p&gt;&lt;/div&gt;
&lt;h3 id=&quot;6-become-deal-eligible-and-privacy-credible&quot;&gt;6. Become deal-eligible and privacy-credible&lt;/h3&gt;
&lt;p&gt;Some categories require an editor to win the deal at all. One interviewed customer received a written RFP rejection that read: &lt;em&gt;“the other provider has an editor, you don’t.”&lt;/em&gt; That single moment approved their budget to buy.&lt;/p&gt;
&lt;p&gt;For regulated buyers, fintech, healthcare, government, enterprise, client-side processing is a procurement requirement, not a feature. CE.SDK’s default client-side execution means customer data never leaves the browser, which makes the editor defensible in categories where server-side alternatives are blocked.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Swiss Post&lt;/strong&gt; producing &lt;strong&gt;1 million+ personalized postcards annually&lt;/strong&gt; on the platform described CE.SDK as &lt;em&gt;“the only solution allowing a specialized, on-brand UI.”&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&quot;the-economics-in-one-table&quot;&gt;The economics, in one table&lt;/h2&gt;
&lt;p&gt;The cost gap between building and buying isn’t just wider on day one, it widens over time. Every month an in-house team spends maintaining the editor is a month IMG.LY customers compound new capability on top of what’s already shipped. The table below summarizes the 3-year picture. For the full cost methodology, including the open-source path and the senior-developer rate assumptions, see IMG.LY’s &lt;a href=&quot;https://img.ly/blog/strategic-guide-to-creative-editing-when-to-build-when-to-buy/&quot;&gt;Build vs. Buy guide&lt;/a&gt;.&lt;/p&gt;





















































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;&lt;/th&gt;&lt;th&gt;Build In-House&lt;/th&gt;&lt;th&gt;Open Source + Glue&lt;/th&gt;&lt;th&gt;CE.SDK&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Working prototype&lt;/td&gt;&lt;td&gt;2–6 months&lt;/td&gt;&lt;td&gt;1–2 months&lt;/td&gt;&lt;td&gt;&lt;strong&gt;2 days&lt;/strong&gt; (self-reported median)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Production MVP&lt;/td&gt;&lt;td&gt;6–12+ months&lt;/td&gt;&lt;td&gt;3–6 months&lt;/td&gt;&lt;td&gt;&lt;strong&gt;14 days&lt;/strong&gt; (self-reported median)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Team required&lt;/td&gt;&lt;td&gt;3–5 senior developers&lt;/td&gt;&lt;td&gt;1–2 senior developers&lt;/td&gt;&lt;td&gt;1 developer for integration&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Year-1 development cost&lt;/td&gt;&lt;td&gt;$100K–$500K+&lt;/td&gt;&lt;td&gt;$150K–$400K&lt;/td&gt;&lt;td&gt;License + integration&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Annual maintenance&lt;/td&gt;&lt;td&gt;$75K–$125K (15–25%)&lt;/td&gt;&lt;td&gt;$50K–$100K&lt;/td&gt;&lt;td&gt;Included&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;3-year TCO&lt;/td&gt;&lt;td&gt;$525K–$1.4M+&lt;/td&gt;&lt;td&gt;$300K–$700K+&lt;/td&gt;&lt;td&gt;Fraction of build cost&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Cross-platform parity&lt;/td&gt;&lt;td&gt;Build per platform&lt;/td&gt;&lt;td&gt;Typically web-only&lt;/td&gt;&lt;td&gt;Web, iOS, Android, Desktop, Server&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;&lt;em&gt;Source for build and open-source cost bands: IMG.LY’s&lt;/em&gt; &lt;a href=&quot;https://img.ly/blog/strategic-guide-to-creative-editing-when-to-build-when-to-buy/&quot;&gt;&lt;em&gt;Strategic Guide to Creative Editing: When to Build, When to Buy&lt;/em&gt;&lt;/a&gt;&lt;em&gt;. Cross-referenced against&lt;/em&gt; &lt;a href=&quot;https://www.bls.gov/ooh/computer-and-information-technology/software-developers.htm&quot;&gt;&lt;em&gt;BLS May 2024 wage data&lt;/em&gt;&lt;/a&gt; &lt;em&gt;- median software-developer wage $133,080, 90th percentile $211,450 - and&lt;/em&gt; &lt;a href=&quot;https://pegotec.net/software-maintenance-cost-percentage-2026-industry-benchmarks/&quot;&gt;&lt;em&gt;aggregated maintenance benchmarks (Pegotec, 2026)&lt;/em&gt;&lt;/a&gt;&lt;em&gt;. Time-to-prototype and time-to-MVP for CE.SDK are &lt;strong&gt;self-reported medians from an IMG.LY customer survey&lt;/strong&gt;; individual results vary with scope, team, and platform target.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&quot;when-building-in-house-still-makes-sense&quot;&gt;When building in-house still makes sense&lt;/h2&gt;
&lt;p&gt;This report defends one conclusion, so here is the honest counterweight.&lt;br&gt;
Build your own creative editor if &lt;strong&gt;all&lt;/strong&gt; of the following are true:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Creative editing is your core product, and your unique IP lives inside the editing experience itself.&lt;/li&gt;
&lt;li&gt;Your requirements are genuinely non-standard, i.e. no off-the-shelf SDK supports your rendering or interaction model.&lt;/li&gt;
&lt;li&gt;You have surplus senior engineering capacity that isn’t needed on differentiated work elsewhere.&lt;/li&gt;
&lt;li&gt;Time-to-market is not competitive in your category, nobody else is shipping similar capability.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If any one of these is false, the economics above generally make buying both cheaper and faster.&lt;/p&gt;
&lt;h2 id=&quot;methodology&quot;&gt;Methodology&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Time-to-prototype and time-to-MVP&lt;/strong&gt; (2 days, 14 days) are self-reported medians from an IMG.LY customer survey. Individual results vary with team experience, integration scope, and platform target.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;In-house and open-source timeline ranges&lt;/strong&gt; are drawn from 28 structured customer interviews conducted between 2024 and 2026, cross-referenced against IMG.LY’s &lt;a href=&quot;https://img.ly/blog/strategic-guide-to-creative-editing-when-to-build-when-to-buy/&quot;&gt;Build vs. Buy guide&lt;/a&gt;. Interview customers are not identified; quantitative ranges reflect the distribution of answers across the set.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Named customer outcomes&lt;/strong&gt; (Plai, Omneky, Optimizely, Swiss Post, ImageBank X, Halio.ai, Postbuddy) come from &lt;a href=&quot;https://img.ly/case-studies/&quot;&gt;IMG.LY’s published case studies&lt;/a&gt;, where the customer has reviewed and approved each cited figure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;External benchmarks&lt;/strong&gt; are drawn from:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://www.bls.gov/ooh/computer-and-information-technology/software-developers.htm&quot;&gt;U.S. Bureau of Labor Statistics, Occupational Outlook Handbook, Software Developers (May 2024)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.atlassian.com/software/compass/resources/state-of-developer-2024&quot;&gt;Atlassian State of Developer Experience Report (2024)&lt;/a&gt; (N = 2,100+ developers and engineering leaders, in partnership with DX and Wakefield Research)&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://survey.stackoverflow.co/2024/&quot;&gt;Stack Overflow 2024 Developer Survey&lt;/a&gt; (N = 65,437 developers across 185 countries)&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://thestory.is/en/journal/chaos-report/&quot;&gt;Standish Group CHAOS Report&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Aggregated software-maintenance benchmarks from Gartner, IEEE, and industry practitioners, summarized in &lt;a href=&quot;https://pegotec.net/software-maintenance-cost-percentage-2026-industry-benchmarks/&quot;&gt;Pegotec (2026)&lt;/a&gt; and &lt;a href=&quot;https://ventionteams.com/enterprise/software-maintenance-costs&quot;&gt;Vention Teams&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;frequently-asked-questions&quot;&gt;Frequently asked questions&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;How long does it take to integrate CE.SDK?&lt;/strong&gt;&lt;br&gt;
The median customer in our survey reports a working prototype within &lt;strong&gt;2 days&lt;/strong&gt; and a production-ready MVP within a single &lt;strong&gt;14-day sprint&lt;/strong&gt;. Self-reported; varies with scope and team.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What do IMG.LY customers unlock beyond faster shipping?&lt;/strong&gt;&lt;br&gt;
Five additional outcomes: new pricing and market tiers enabled by self-serve editing; lifted trial-to-paid and sign-up conversion; reduced design, marketing, and support load; a roadmap that compounds instead of freezes as AI features ship as plugins; and deal eligibility in procurement categories that require an embedded editor or client-side processing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can AI coding agents like Claude Code accelerate CE.SDK integration further?&lt;/strong&gt;&lt;br&gt;
Yes. IMG.LY publishes a set of specialized &lt;a href=&quot;https://github.com/imgly/agent-skills&quot;&gt;CE.SDK Agent Skills&lt;/a&gt; portable knowledge packs that give AI coding assistants expert-level understanding of the SDK before they write a line of code. The skills cover framework-specific documentation lookup (React, Vue, Svelte, Angular, Next.js, Nuxt.js, SvelteKit, Electron, Node.js, Vanilla JS), feature implementation (&lt;code&gt;/cesdk:build&lt;/code&gt;), conceptual explanations (&lt;code&gt;/cesdk:explain&lt;/code&gt;), and a builder agent that autonomously scaffolds complete CE.SDK applications from a natural-language description.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How much does it cost to build a creative editor in-house?&lt;/strong&gt;&lt;br&gt;
Published industry benchmarks and customer-reported estimates converge on &lt;strong&gt;$100K–$500K+ initial development&lt;/strong&gt; over &lt;strong&gt;6–12 months&lt;/strong&gt; with &lt;strong&gt;3–5 senior developers&lt;/strong&gt;, plus &lt;strong&gt;15–25% of that cost annually in maintenance&lt;/strong&gt;. 3-year TCO typically lands between &lt;strong&gt;$525K and $1.4M&lt;/strong&gt;. See our &lt;a href=&quot;https://img.ly/blog/strategic-guide-to-creative-editing-when-to-build-when-to-buy/&quot;&gt;Build vs. Buy guide&lt;/a&gt; for full methodology.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What percentage of developer time is spent on maintenance?&lt;/strong&gt;&lt;br&gt;
Atlassian’s 2024 &lt;em&gt;State of Developer Experience Report&lt;/em&gt; found &lt;strong&gt;69% of developers lose 8+ hours per week to inefficiencies, about 20% of their time, primarily to technical debt&lt;/strong&gt;. Gartner reports that &lt;strong&gt;55–80% of corporate IT budgets&lt;/strong&gt; go to maintaining existing systems.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What is the failure rate of custom software projects?&lt;/strong&gt;&lt;br&gt;
The Standish Group CHAOS research finds only &lt;strong&gt;31% of software projects succeed&lt;/strong&gt; (on time, on budget, full scope); &lt;strong&gt;50% are challenged&lt;/strong&gt;, &lt;strong&gt;19% cancelled outright&lt;/strong&gt;. For large projects specifically, success rates fall &lt;strong&gt;below 10%&lt;/strong&gt;.&lt;/p&gt;
&lt;div class=&quot;imgly-chart imgly-chart--success&quot;&gt;&lt;p class=&quot;imgly-eyebrow&quot;&gt;Custom software project outcomes&lt;/p&gt;&lt;h3 class=&quot;imgly-title&quot;&gt;Only 31% of custom software projects succeed&lt;/h3&gt;&lt;p class=&quot;imgly-subtitle&quot;&gt;Across tens of thousands of tracked projects, the Standish Group CHAOS research finds the majority of custom software is late, over budget, reduced in scope, or cancelled outright.&lt;/p&gt;&lt;div class=&quot;imgly-grid&quot;&gt;&lt;div class=&quot;imgly-canvas-wrap&quot;&gt;&lt;canvas id=&quot;imgly-chart-success-canvas&quot;&gt;&lt;/canvas&gt;&lt;/div&gt;&lt;div class=&quot;imgly-legend&quot;&gt;&lt;div class=&quot;imgly-legend-row&quot;&gt;&lt;span class=&quot;imgly-legend-swatch&quot;&gt;&lt;/span&gt;&lt;div&gt;&lt;div class=&quot;imgly-legend-pct&quot;&gt;31%&lt;/div&gt;&lt;div class=&quot;imgly-legend-label&quot;&gt;Successful&lt;/div&gt;&lt;div class=&quot;imgly-legend-detail&quot;&gt;On time, on budget, with full scope.&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;div class=&quot;imgly-legend-row&quot;&gt;&lt;span class=&quot;imgly-legend-swatch&quot;&gt;&lt;/span&gt;&lt;div&gt;&lt;div class=&quot;imgly-legend-pct&quot;&gt;50%&lt;/div&gt;&lt;div class=&quot;imgly-legend-label&quot;&gt;Challenged&lt;/div&gt;&lt;div class=&quot;imgly-legend-detail&quot;&gt;Delivered late, over budget, or with reduced features.&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;div class=&quot;imgly-legend-row&quot;&gt;&lt;span class=&quot;imgly-legend-swatch&quot;&gt;&lt;/span&gt;&lt;div&gt;&lt;div class=&quot;imgly-legend-pct&quot;&gt;19%&lt;/div&gt;&lt;div class=&quot;imgly-legend-label&quot;&gt;Failed&lt;/div&gt;&lt;div class=&quot;imgly-legend-detail&quot;&gt;Cancelled before delivery.&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;div class=&quot;imgly-callout&quot;&gt;&lt;p class=&quot;imgly-callout-headline&quot;&gt;~25% of IMG.LY customers failed at building a creative editor in-house&lt;/p&gt;&lt;p class=&quot;imgly-callout-body&quot;&gt;In our customer interviews (N=28), roughly one in four customers had attempted to build a creative editor in-house (or stitch one together from open-source), and abandoned the attempt before adopting CE.SDK.&lt;/p&gt;&lt;/div&gt;&lt;p class=&quot;imgly-footnote&quot;&gt;&lt;strong&gt;Source.&lt;/strong&gt; Standish Group CHAOS Report. For large projects specifically (the size band most in-house creative-editor builds fall into), success rates drop below &lt;strong&gt;10%&lt;/strong&gt;. IMG.LY customer figure is drawn from 28 structured customer interviews conducted between 2024 and 2026.&lt;/p&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;How many companies use IMG.LY CE.SDK?&lt;br&gt;
600+ customers&lt;/strong&gt; including startups, Fortune 500 companies, and government organizations collectively produce &lt;strong&gt;500 million creations per month&lt;/strong&gt; on the platform.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What platforms does CE.SDK support?&lt;/strong&gt;&lt;br&gt;
Web (React, Angular, Vue, Svelte, Next.js, Nuxt.js, SvelteKit, Vanilla JavaScript, Electron), Mobile (iOS/Swift, Android/Kotlin, React Native, Flutter, Ionic, Cordova), and Server (Node.js headless mode for batch automation).&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="https://blog.img.ly/2026/04/imgly-data-report.jpg" medium="image"/><category>Creative Editing</category></item><item><title>Building a CapCut-Like Video Editor with CE.SDK and AI</title><link>https://img.ly/blog/capcut-like-video-editor-web-react/</link><guid isPermaLink="true">https://img.ly/blog/capcut-like-video-editor-web-react/</guid><description>How to built a browser-based video editor with a dark CapCut-inspired UI, multi-track timeline, and AI-powered content generation... in under 200 lines of code.</description><pubDate>Thu, 26 Feb 2026 22:50:34 GMT</pubDate><content:encoded>&lt;h3 id=&quot;build-this-with-ai-in-minutes&quot;&gt;&lt;strong&gt;Build this with AI in minutes&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;This tutorial is a great way to understand how CE.SDK works under the hood. But if you just want to get to the result, you can build this entire editor using &lt;a href=&quot;https://img.ly/docs/cesdk/react/get-started/agent-skills-f7g8h9/&quot;&gt;IMG.LY Agent Skills&lt;/a&gt;, no manual setup required.&lt;/p&gt;
&lt;p&gt;Install the skills, then run:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;/cesdk:build Build a CapCut-like video editor with a dark theme, multi-track timeline, AI video/image/audio generation, background removal, and MP4 export.&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Want to understand a specific concept in depth? Use &lt;code&gt;/cesdk:explain&lt;/code&gt;, for example:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;/cesdk:explain How does the video timeline and block hierarchy work?&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;CapCut has set the standard for accessible video editing. Its dark UI, intuitive timeline, and AI-powered features make it feel like a professional tool that anyone can use. But what if you could build something similar embedded directly in your own web application in minutes? Yes, minutes!&lt;/p&gt;
&lt;p&gt;In this tutorial, we’ll walk through how we built a CapCut-like video editor using &lt;a href=&quot;https://img.ly&quot;&gt;IMG.LY’s CreativeEditor SDK (CE.SDK)&lt;/a&gt; for React. The finished editor includes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;multi-track timeline&lt;/strong&gt; for video, audio, text, and &lt;a href=&quot;https://img.ly/demos/video-captions/web/&quot;&gt;captions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Trim, split, and join&lt;/strong&gt; operations on video clips&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;CapCut-inspired dark theme&lt;/strong&gt; built with CSS custom properties&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI video generation&lt;/strong&gt; (text-to-video, image-to-video) via Minimax, Kling, and Pixverse&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI image generation&lt;/strong&gt; (text-to-image, image editing) via RecraftV3, IdeogramV3, and GPT Image&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI audio generation&lt;/strong&gt; (text-to-speech, sound effects) via ElevenLabs&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI text generation&lt;/strong&gt; (copywriting, translation) via Anthropic Claude&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Background removal&lt;/strong&gt; — client-side, powered by WebAssembly/WebGPU&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;MP4 export&lt;/strong&gt; directly from the browser&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The entire editor is a single React component with roughly 180 lines of JavaScript and 100 lines of CSS.&lt;/p&gt;
&lt;h2 id=&quot;architecture-overview&quot;&gt;Architecture Overview&lt;/h2&gt;
&lt;p&gt;&lt;img alt=&quot;App architecture&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 2000px) 2000px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;2000&quot; height=&quot;1969&quot; src=&quot;https://img.ly/_astro/Gemini_Generated_Image_6celpb6celpb6cel_4npMe.webp&quot; srcset=&quot;/_astro/Gemini_Generated_Image_6celpb6celpb6cel_Z2pUzxd.webp 640w, /_astro/Gemini_Generated_Image_6celpb6celpb6cel_1I72rW.webp 750w, /_astro/Gemini_Generated_Image_6celpb6celpb6cel_1YwWvd.webp 828w, /_astro/Gemini_Generated_Image_6celpb6celpb6cel_cWEce.webp 1080w, /_astro/Gemini_Generated_Image_6celpb6celpb6cel_22mCvu.webp 1280w, /_astro/Gemini_Generated_Image_6celpb6celpb6cel_Z1OLGCT.webp 1668w, /_astro/Gemini_Generated_Image_6celpb6celpb6cel_4npMe.webp 2000w&quot;&gt;&lt;/p&gt;
&lt;p&gt;CE.SDK has two main layers:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;CreativeEngine:&lt;/strong&gt; The headless core that manages scenes, blocks, assets, and rendering. It handles the video timeline, playback, and export entirely client-side.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CreativeEditor UI:&lt;/strong&gt; A pre-built, customizable UI layer that wraps the engine with a dock, inspector, timeline, canvas, and navigation bar.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Plugins extend both layers. The &lt;code&gt;AiApps&lt;/code&gt; plugin, for example, registers AI providers with the engine and injects UI components (dock buttons, canvas menu items, generation panels) into the editor.&lt;/p&gt;
&lt;h2 id=&quot;step-1-project-setup&quot;&gt;Step 1: Project Setup&lt;/h2&gt;
&lt;p&gt;We scaffolded a React project with Vite and installed CE.SDK alongside the AI plugin packages:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;npm&lt;/span&gt;&lt;span&gt; create&lt;/span&gt;&lt;span&gt; vite@latest&lt;/span&gt;&lt;span&gt; capcut-like-editor&lt;/span&gt;&lt;span&gt; --&lt;/span&gt;&lt;span&gt; --template&lt;/span&gt;&lt;span&gt; react&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;cd&lt;/span&gt;&lt;span&gt; capcut-like-editor&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;npm&lt;/span&gt;&lt;span&gt; install&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Core SDK&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;npm&lt;/span&gt;&lt;span&gt; install&lt;/span&gt;&lt;span&gt; @cesdk/cesdk-js&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# AI plugins&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;npm&lt;/span&gt;&lt;span&gt; install&lt;/span&gt;&lt;span&gt; @imgly/plugin-ai-apps-web&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;npm&lt;/span&gt;&lt;span&gt; install&lt;/span&gt;&lt;span&gt; @imgly/plugin-ai-video-generation-web&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;npm&lt;/span&gt;&lt;span&gt; install&lt;/span&gt;&lt;span&gt; @imgly/plugin-ai-image-generation-web&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;npm&lt;/span&gt;&lt;span&gt; install&lt;/span&gt;&lt;span&gt; @imgly/plugin-ai-audio-generation-web&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;npm&lt;/span&gt;&lt;span&gt; install&lt;/span&gt;&lt;span&gt; @imgly/plugin-ai-text-generation-web&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Background removal (runs locally via WASM/WebGPU)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;npm&lt;/span&gt;&lt;span&gt; install&lt;/span&gt;&lt;span&gt; @imgly/plugin-background-removal-web&lt;/span&gt;&lt;span&gt; onnxruntime-web@1.21.0&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;why-so-many-packages&quot;&gt;Why so many packages?&lt;/h3&gt;
&lt;p&gt;CE.SDK follows a modular plugin architecture. Each AI capability is a separate package with its own provider modules. This means you only bundle what you use if you don’t need audio generation, don’t install &lt;code&gt;@imgly/plugin-ai-audio-generation-web&lt;/code&gt;. The &lt;code&gt;@imgly/plugin-ai-apps-web&lt;/code&gt; package is the unifying layer that brings them all together into a single dock panel.&lt;/p&gt;
&lt;h2 id=&quot;step-2-the-react-component&quot;&gt;Step 2: The React Component&lt;/h2&gt;
&lt;p&gt;CE.SDK provides a first-class React wrapper via &lt;code&gt;@cesdk/cesdk-js/react&lt;/code&gt;. The &lt;code&gt;&amp;#x3C;CreativeEditor&gt;&lt;/code&gt; component handles mounting, initialization, and cleanup:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;jsx&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; CreativeEditor &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;@cesdk/cesdk-js/react&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; config&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  // license: &apos;YOUR_CESDK_LICENSE_KEY&apos;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;};&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; init&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; async&lt;/span&gt;&lt;span&gt; (&lt;/span&gt;&lt;span&gt;cesdk&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;=&gt;&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  // All setup happens here — theme, assets, plugins, UI customization&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;};&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;export&lt;/span&gt;&lt;span&gt; default&lt;/span&gt;&lt;span&gt; function&lt;/span&gt;&lt;span&gt; VideoEditor&lt;/span&gt;&lt;span&gt;() {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  return&lt;/span&gt;&lt;span&gt; (&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    &amp;#x3C;&lt;/span&gt;&lt;span&gt;CreativeEditor&lt;/span&gt;&lt;span&gt; config&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;{config} &lt;/span&gt;&lt;span&gt;init&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;{init} &lt;/span&gt;&lt;span&gt;width&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&quot;100vw&quot;&lt;/span&gt;&lt;span&gt; height&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&quot;100vh&quot;&lt;/span&gt;&lt;span&gt; /&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  );&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;}&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The &lt;code&gt;config&lt;/code&gt; object is passed to the engine at creation time. The &lt;code&gt;init&lt;/code&gt; callback fires once the &lt;code&gt;cesdk&lt;/code&gt; instance is ready. This is where all our customization lives.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Pitfall: Silent init errors.&lt;/strong&gt; The &lt;code&gt;&amp;#x3C;CreativeEditor&gt;&lt;/code&gt; component swallows errors thrown inside &lt;code&gt;init&lt;/code&gt;. If something fails, the editor loads but appears broken with no console output. Always wrap &lt;code&gt;init&lt;/code&gt; in a &lt;code&gt;try/catch&lt;/code&gt;:&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;step-3-creating-the-video-scene&quot;&gt;Step 3: Creating the Video Scene&lt;/h2&gt;
&lt;p&gt;A video editor needs three things at startup: asset sources, a video scene, and a timeline.&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;js&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Load built-in asset libraries (stickers, shapes, filters, typefaces, etc.)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;await&lt;/span&gt;&lt;span&gt; cesdk.&lt;/span&gt;&lt;span&gt;addDefaultAssetSources&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Load demo content (sample videos, images, audio) + enable upload slots&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;await&lt;/span&gt;&lt;span&gt; cesdk.&lt;/span&gt;&lt;span&gt;addDemoAssetSources&lt;/span&gt;&lt;span&gt;({&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  sceneMode: &lt;/span&gt;&lt;span&gt;&apos;Video&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  withUploadAssetSources: &lt;/span&gt;&lt;span&gt;true&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;});&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// Create a video scene with timeline&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;await&lt;/span&gt;&lt;span&gt; cesdk.&lt;/span&gt;&lt;span&gt;createVideoScene&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;code&gt;createVideoScene()&lt;/code&gt; sets up the scene hierarchy that powers the timeline:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Scene&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt; └── Page (represents the video canvas — 1920x1080 by default)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      ├── Track (video track — holds video clips in sequence)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      ├── Track (overlay track — text, stickers, images)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      ├── CaptionTrack (subtitles synced to playback)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      └── Audio (background music, voiceover, sound effects)&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Each element on the timeline is a &lt;strong&gt;block&lt;/strong&gt; with timing properties: &lt;code&gt;timeOffset&lt;/code&gt; (when it appears) and &lt;code&gt;duration&lt;/code&gt; (how long it plays). The engine handles rendering each frame, compositing layers, and synchronizing audio.&lt;/p&gt;
&lt;h3 id=&quot;browser-support&quot;&gt;Browser support&lt;/h3&gt;
&lt;p&gt;Video editing relies on modern web codecs (WebCodecs API), which are available in Chromium-based browsers (Chrome, Edge, Brave). Safari and Firefox support is limited.&lt;/p&gt;
&lt;h2 id=&quot;step-4-the-capcut-dark-theme&quot;&gt;Step 4: The CapCut Dark Theme&lt;/h2&gt;
&lt;p&gt;CapCut’s visual identity is defined by its deep charcoal backgrounds, teal accent colors, and subtle elevation layers. CE.SDK’s theming system maps perfectly to this through CSS custom properties.&lt;/p&gt;
&lt;h3 id=&quot;setting-the-base-theme&quot;&gt;Setting the base theme&lt;/h3&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;js&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;cesdk.ui.&lt;/span&gt;&lt;span&gt;setTheme&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;dark&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This activates CE.SDK’s built-in dark theme. But we want to go further. We need CapCut’s specific color palette.&lt;/p&gt;
&lt;h3 id=&quot;custom-css-overrides&quot;&gt;Custom CSS overrides&lt;/h3&gt;
&lt;p&gt;CE.SDK scopes all its UI under &lt;code&gt;.ubq-public&lt;/code&gt; with &lt;code&gt;data-ubq-theme&lt;/code&gt; and &lt;code&gt;data-ubq-scale&lt;/code&gt; attributes. We override the CSS custom properties to inject our palette:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;css&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;.ubq-public&lt;/span&gt;&lt;span&gt;[&lt;/span&gt;&lt;span&gt;data-ubq-theme&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&apos;dark&apos;&lt;/span&gt;&lt;span&gt;][&lt;/span&gt;&lt;span&gt;data-ubq-scale&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&apos;normal&apos;&lt;/span&gt;&lt;span&gt;],&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;.ubq-public&lt;/span&gt;&lt;span&gt;[&lt;/span&gt;&lt;span&gt;data-ubq-theme&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&apos;dark&apos;&lt;/span&gt;&lt;span&gt;][&lt;/span&gt;&lt;span&gt;data-ubq-scale&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&apos;modern&apos;&lt;/span&gt;&lt;span&gt;] {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  /* Deep charcoal backgrounds — darker than CE.SDK&apos;s default dark */&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  --ubq-canvas&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;hsl&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;220&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;15&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;8&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;!important&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  --ubq-elevation-1&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;hsl&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;220&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;13&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;12&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;!important&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  --ubq-elevation-2&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;hsl&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;220&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;12&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;15&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;!important&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  --ubq-elevation-3&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;hsl&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;220&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;10&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;18&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;!important&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  /* High-contrast white text */&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  --ubq-foreground-default&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;hsla&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;100&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;0.92&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;!important&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  --ubq-foreground-light&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;hsla&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;100&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;0.55&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;!important&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  /* CapCut&apos;s signature teal/cyan accent */&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  --ubq-interactive-accent-default&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;hsl&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;190&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;85&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;48&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;!important&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  --ubq-interactive-accent-hover&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;hsl&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;190&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;85&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;42&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;!important&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  --ubq-interactive-accent-pressed&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;hsl&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;190&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;85&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;36&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;!important&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  /* Subtle borders — barely visible, like CapCut */&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  --ubq-border-default&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;hsla&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;100&lt;/span&gt;&lt;span&gt;%&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;0.08&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;!important&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;}&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The key design decisions:&lt;/p&gt;






























&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Property&lt;/th&gt;&lt;th&gt;Value&lt;/th&gt;&lt;th&gt;Why&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;code&gt;--ubq-canvas&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;hsl(220, 15%, 8%)&lt;/code&gt;&lt;/td&gt;&lt;td&gt;Near-black with a slight blue tint — matches CapCut’s canvas area&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;code&gt;--ubq-elevation-1/2/3&lt;/code&gt;&lt;/td&gt;&lt;td&gt;12% → 15% → 18% lightness&lt;/td&gt;&lt;td&gt;Subtle elevation steps create depth without harsh contrast&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;code&gt;--ubq-interactive-accent-*&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;hsl(190, 85%, 48%)&lt;/code&gt;&lt;/td&gt;&lt;td&gt;CapCut’s teal — used for buttons, selections, and progress bars&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;code&gt;--ubq-border-default&lt;/code&gt;&lt;/td&gt;&lt;td&gt;8% opacity white&lt;/td&gt;&lt;td&gt;Nearly invisible borders that only appear on close inspection&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;h3 id=&quot;responsive-scale&quot;&gt;Responsive scale&lt;/h3&gt;
&lt;p&gt;We also configure responsive scaling so the editor adapts to touch devices:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;js&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;cesdk.ui.&lt;/span&gt;&lt;span&gt;setScale&lt;/span&gt;&lt;span&gt;(({ &lt;/span&gt;&lt;span&gt;containerWidth&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;isTouch&lt;/span&gt;&lt;span&gt; }) &lt;/span&gt;&lt;span&gt;=&gt;&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  if&lt;/span&gt;&lt;span&gt; ((containerWidth &lt;/span&gt;&lt;span&gt;&amp;#x26;&amp;#x26;&lt;/span&gt;&lt;span&gt; containerWidth &lt;/span&gt;&lt;span&gt;&amp;#x3C;&lt;/span&gt;&lt;span&gt; 768&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;||&lt;/span&gt;&lt;span&gt; isTouch) {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    return&lt;/span&gt;&lt;span&gt; &apos;large&apos;&lt;/span&gt;&lt;span&gt;; &lt;/span&gt;&lt;span&gt;// Bigger touch targets on small/touch screens&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  }&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  return&lt;/span&gt;&lt;span&gt; &apos;normal&apos;&lt;/span&gt;&lt;span&gt;; &lt;/span&gt;&lt;span&gt;// Standard desktop sizing&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;});&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;step-5-ai-plugin-integration&quot;&gt;Step 5: AI Plugin Integration&lt;/h2&gt;
&lt;p&gt;This is where the editor transforms from a basic video tool into something that feels like CapCut’s AI-powered experience. CE.SDK’s plugin system lets us add all AI capabilities through a single unified &lt;code&gt;AiApps&lt;/code&gt; plugin.&lt;/p&gt;
&lt;h3 id=&quot;the-unified-aiapps-approach&quot;&gt;The unified AiApps approach&lt;/h3&gt;
&lt;p&gt;Instead of registering each AI plugin separately, &lt;code&gt;@imgly/plugin-ai-apps-web&lt;/code&gt; provides a single entry point:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;js&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; AiApps &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;@imgly/plugin-ai-apps-web&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; FalAiVideo &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;@imgly/plugin-ai-video-generation-web/fal-ai&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; FalAiImage &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;@imgly/plugin-ai-image-generation-web/fal-ai&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; OpenAiImage &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;@imgly/plugin-ai-image-generation-web/open-ai&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; Elevenlabs &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;@imgly/plugin-ai-audio-generation-web/elevenlabs&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; Anthropic &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;@imgly/plugin-ai-text-generation-web/anthropic&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;await&lt;/span&gt;&lt;span&gt; cesdk.&lt;/span&gt;&lt;span&gt;addPlugin&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  AiApps&lt;/span&gt;&lt;span&gt;({&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    dryRun: &lt;/span&gt;&lt;span&gt;true&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;// Simulate for development — no API calls&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    providers: {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      text2text: Anthropic.&lt;/span&gt;&lt;span&gt;AnthropicProvider&lt;/span&gt;&lt;span&gt;({ proxyUrl: &lt;/span&gt;&lt;span&gt;PROXY_URL&lt;/span&gt;&lt;span&gt; }),&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      text2image: [FalAiImage.&lt;/span&gt;&lt;span&gt;RecraftV3&lt;/span&gt;&lt;span&gt;({ proxyUrl }) &lt;/span&gt;&lt;span&gt;/* ... */&lt;/span&gt;&lt;span&gt;],&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      image2image: [FalAiImage.&lt;/span&gt;&lt;span&gt;GeminiFlashEdit&lt;/span&gt;&lt;span&gt;({ proxyUrl }) &lt;/span&gt;&lt;span&gt;/* ... */&lt;/span&gt;&lt;span&gt;],&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      text2video: [FalAiVideo.&lt;/span&gt;&lt;span&gt;MinimaxVideo01Live&lt;/span&gt;&lt;span&gt;({ proxyUrl }) &lt;/span&gt;&lt;span&gt;/* ... */&lt;/span&gt;&lt;span&gt;],&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      image2video: [&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        FalAiVideo.&lt;/span&gt;&lt;span&gt;MinimaxVideo01LiveImageToVideo&lt;/span&gt;&lt;span&gt;({ proxyUrl }) &lt;/span&gt;&lt;span&gt;/* ... */&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      ],&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      text2speech: Elevenlabs.&lt;/span&gt;&lt;span&gt;ElevenMultilingualV2&lt;/span&gt;&lt;span&gt;({ proxyUrl }),&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      text2sound: Elevenlabs.&lt;/span&gt;&lt;span&gt;ElevenSoundEffects&lt;/span&gt;&lt;span&gt;({ proxyUrl }),&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  })&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;provider-categories-explained&quot;&gt;Provider categories explained&lt;/h3&gt;













































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Category&lt;/th&gt;&lt;th&gt;What it does&lt;/th&gt;&lt;th&gt;Models we configured&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;code&gt;text2text&lt;/code&gt;&lt;/td&gt;&lt;td&gt;AI copywriting — improve text, translate, change tone&lt;/td&gt;&lt;td&gt;Anthropic Claude&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;code&gt;text2image&lt;/code&gt;&lt;/td&gt;&lt;td&gt;Generate images from text prompts&lt;/td&gt;&lt;td&gt;RecraftV3, IdeogramV3, GPT Image&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;code&gt;image2image&lt;/code&gt;&lt;/td&gt;&lt;td&gt;Transform existing images with AI&lt;/td&gt;&lt;td&gt;Gemini Flash Edit, GPT Image&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;code&gt;text2video&lt;/code&gt;&lt;/td&gt;&lt;td&gt;Generate video clips from text descriptions&lt;/td&gt;&lt;td&gt;Minimax Video, Kling Video, Pixverse&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;code&gt;image2video&lt;/code&gt;&lt;/td&gt;&lt;td&gt;Animate a static image into video&lt;/td&gt;&lt;td&gt;Minimax Video, Kling Video&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;code&gt;text2speech&lt;/code&gt;&lt;/td&gt;&lt;td&gt;Convert text to spoken audio with voice selection&lt;/td&gt;&lt;td&gt;ElevenLabs Multilingual V2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;code&gt;text2sound&lt;/code&gt;&lt;/td&gt;&lt;td&gt;Generate sound effects from text&lt;/td&gt;&lt;td&gt;ElevenLabs Sound Effects&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;When multiple providers are configured in an array (like &lt;code&gt;text2image&lt;/code&gt;), the UI automatically shows a provider/model selection dropdown so users can choose which AI model to use.&lt;/p&gt;
&lt;h3 id=&quot;the-proxy-server-requirement&quot;&gt;The proxy server requirement&lt;/h3&gt;
&lt;p&gt;Every provider takes a &lt;code&gt;proxyUrl&lt;/code&gt; parameter. This is critical for production:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Browser  →  Your Proxy Server  →  AI Provider (fal.ai, ElevenLabs, etc.)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;                ↑&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;          Injects API keys&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;          server-side&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Your API keys should never be in client-side code. The proxy server receives requests from CE.SDK, attaches your API key, and forwards to the AI provider. During development, &lt;code&gt;dryRun: true&lt;/code&gt; simulates generation without any API calls.&lt;/p&gt;
&lt;h3 id=&quot;background-removal&quot;&gt;Background removal&lt;/h3&gt;
&lt;p&gt;Background removal is a separate plugin because it runs entirely client-side, no proxy needed:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;js&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; BackgroundRemovalPlugin &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;@imgly/plugin-background-removal-web&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;await&lt;/span&gt;&lt;span&gt; cesdk.&lt;/span&gt;&lt;span&gt;addPlugin&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  BackgroundRemovalPlugin&lt;/span&gt;&lt;span&gt;({&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    ui: { locations: [&lt;/span&gt;&lt;span&gt;&apos;canvasMenu&apos;&lt;/span&gt;&lt;span&gt;] },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  })&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This uses ONNX Runtime (WebAssembly + WebGPU) to run an AI segmentation model directly in the browser. The first run downloads ~40MB of model weights, which are then cached. Select any image on the canvas and click “Remove Background” in the context menu.&lt;/p&gt;
&lt;h2 id=&quot;step-6-ui-customization-with-the-component-order-api&quot;&gt;Step 6: UI Customization with the Component Order API&lt;/h2&gt;
&lt;p&gt;CE.SDK’s UI is built from five customizable areas: &lt;strong&gt;Dock&lt;/strong&gt;, &lt;strong&gt;Inspector Bar&lt;/strong&gt;, &lt;strong&gt;Canvas Menu&lt;/strong&gt;, &lt;strong&gt;Navigation Bar&lt;/strong&gt;, and &lt;strong&gt;Canvas Bar&lt;/strong&gt;. The Component Order API lets us insert, remove, and reorder components in each area.&lt;/p&gt;
&lt;h3 id=&quot;adding-the-ai-button-to-the-dock&quot;&gt;Adding the AI button to the dock&lt;/h3&gt;
&lt;p&gt;We want the AI Apps button to be the first thing users see in the dock:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;js&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;cesdk.ui.&lt;/span&gt;&lt;span&gt;insertOrderComponent&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  { in: &lt;/span&gt;&lt;span&gt;&apos;ly.img.dock&apos;&lt;/span&gt;&lt;span&gt;, position: &lt;/span&gt;&lt;span&gt;&apos;start&apos;&lt;/span&gt;&lt;span&gt; },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &apos;ly.img.ai.apps.dock&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;code&gt;insertOrderComponent&lt;/code&gt; takes a location specifier and the component(s) to insert. &lt;code&gt;position: &apos;start&apos;&lt;/code&gt; puts it at the top of the dock. Alternative positions include &lt;code&gt;&apos;end&apos;&lt;/code&gt;, a numeric index, or relative placement with &lt;code&gt;before&lt;/code&gt;/&lt;code&gt;after&lt;/code&gt; matchers.&lt;/p&gt;
&lt;h3 id=&quot;ai-options-in-the-canvas-context-menu&quot;&gt;AI options in the canvas context menu&lt;/h3&gt;
&lt;p&gt;When a user selects a text or image block and right-clicks, we want AI options available:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;js&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;cesdk.ui.&lt;/span&gt;&lt;span&gt;insertOrderComponent&lt;/span&gt;&lt;span&gt;({ in: &lt;/span&gt;&lt;span&gt;&apos;ly.img.canvas.menu&apos;&lt;/span&gt;&lt;span&gt;, position: &lt;/span&gt;&lt;span&gt;&apos;start&apos;&lt;/span&gt;&lt;span&gt; }, [&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &apos;ly.img.ai.text.canvasMenu&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &apos;ly.img.ai.image.canvasMenu&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;]);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Passing an array inserts multiple components at once.&lt;/p&gt;
&lt;h3 id=&quot;custom-export-button-in-the-navigation-bar&quot;&gt;Custom export button in the navigation bar&lt;/h3&gt;
&lt;p&gt;We add an Export button with a real click handler that triggers MP4 export:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;js&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;cesdk.ui.&lt;/span&gt;&lt;span&gt;insertOrderComponent&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  { in: &lt;/span&gt;&lt;span&gt;&apos;ly.img.navigation.bar&apos;&lt;/span&gt;&lt;span&gt;, position: &lt;/span&gt;&lt;span&gt;&apos;end&apos;&lt;/span&gt;&lt;span&gt; },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    id: &lt;/span&gt;&lt;span&gt;&apos;ly.img.action.navigationBar&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    key: &lt;/span&gt;&lt;span&gt;&apos;export&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    label: &lt;/span&gt;&lt;span&gt;&apos;Export&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    icon: &lt;/span&gt;&lt;span&gt;&apos;@imgly/Download&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    onClick&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;async&lt;/span&gt;&lt;span&gt; () &lt;/span&gt;&lt;span&gt;=&gt;&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      const&lt;/span&gt;&lt;span&gt; engine&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; cesdk.engine;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      const&lt;/span&gt;&lt;span&gt; page&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;findByType&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;page&apos;&lt;/span&gt;&lt;span&gt;)[&lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;];&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      if&lt;/span&gt;&lt;span&gt; (page) {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        const&lt;/span&gt;&lt;span&gt; blob&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; await&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;exportVideo&lt;/span&gt;&lt;span&gt;(page, {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;          mimeType: &lt;/span&gt;&lt;span&gt;&apos;video/mp4&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        });&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        const&lt;/span&gt;&lt;span&gt; url&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; URL&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;createObjectURL&lt;/span&gt;&lt;span&gt;(blob);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        const&lt;/span&gt;&lt;span&gt; a&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; document.&lt;/span&gt;&lt;span&gt;createElement&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;a&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        a.href &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; url;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        a.download &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; &apos;video.mp4&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        a.&lt;/span&gt;&lt;span&gt;click&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        URL&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;revokeObjectURL&lt;/span&gt;&lt;span&gt;(url);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      }&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  }&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The &lt;code&gt;exportVideo&lt;/code&gt; method renders every frame of the timeline (compositing video tracks, overlays, text, captions, and audio) into an MP4 blob entirely in the browser.&lt;/p&gt;
&lt;h2 id=&quot;step-7-wiring-generated-audio-into-the-asset-library&quot;&gt;Step 7: Wiring Generated Audio into the Asset Library&lt;/h2&gt;
&lt;p&gt;AI-generated audio (speech and sound effects) is stored in provider-specific history sources. To make this audio browsable alongside regular audio assets, we inject the history source into the audio asset library entry:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;js&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; audioEntry&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; cesdk.ui.&lt;/span&gt;&lt;span&gt;getAssetLibraryEntry&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;ly.img.audio&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;if&lt;/span&gt;&lt;span&gt; (audioEntry &lt;/span&gt;&lt;span&gt;!=&lt;/span&gt;&lt;span&gt; null&lt;/span&gt;&lt;span&gt;) {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  const&lt;/span&gt;&lt;span&gt; existingSourceIds&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; Array.&lt;/span&gt;&lt;span&gt;isArray&lt;/span&gt;&lt;span&gt;(audioEntry.sourceIds)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    ?&lt;/span&gt;&lt;span&gt; audioEntry.sourceIds&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    :&lt;/span&gt;&lt;span&gt; audioEntry.&lt;/span&gt;&lt;span&gt;sourceIds&lt;/span&gt;&lt;span&gt;({});&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  cesdk.ui.&lt;/span&gt;&lt;span&gt;updateAssetLibraryEntry&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;ly.img.audio&apos;&lt;/span&gt;&lt;span&gt;, {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    sourceIds: [&lt;/span&gt;&lt;span&gt;...&lt;/span&gt;&lt;span&gt;existingSourceIds, &lt;/span&gt;&lt;span&gt;&apos;ly.img.ai.audio-generation.history&apos;&lt;/span&gt;&lt;span&gt;],&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  });&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;}&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Now when users open the Audio panel in the dock, they’ll see their AI-generated audio alongside the sample library.&lt;/p&gt;
&lt;h2 id=&quot;step-8-custom-labels-with-i18n&quot;&gt;Step 8: Custom Labels with i18n&lt;/h2&gt;
&lt;p&gt;CE.SDK’s i18n system lets us customize any UI string. We used it to make the AI prompt fields more inviting:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;js&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;cesdk.i18n.&lt;/span&gt;&lt;span&gt;setTranslations&lt;/span&gt;&lt;span&gt;({&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  en: {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    &apos;ly.img.plugin-ai-video-generation-web.fal-ai/minimax/video-01-live.property.prompt&apos;&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      &apos;Describe your video...&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    &apos;ly.img.plugin-ai-image-generation-web.fal-ai/recraft-v3.property.prompt&apos;&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      &apos;Describe your image...&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;});&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Translation keys follow the pattern &lt;code&gt;{plugin-id}.{provider-id}.property.{field}&lt;/code&gt;. You can also add multi-language support by including keys for &lt;code&gt;de&lt;/code&gt;, &lt;code&gt;fr&lt;/code&gt;, &lt;code&gt;es&lt;/code&gt;, etc.&lt;/p&gt;
&lt;h2 id=&quot;the-complete-file-structure&quot;&gt;The Complete File Structure&lt;/h2&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;capcut-like-editor/&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;├── src/&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;│   ├── App.jsx              # Root — imports VideoEditor + theme CSS&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;│   ├── VideoEditor.jsx      # The entire editor (single component, ~180 lines)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;│   ├── capcut-theme.css     # CapCut dark theme overrides (~100 lines)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;│   ├── index.css            # Global reset (margin/padding/overflow)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;│   └── main.jsx             # React entry point&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;├── index.html&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;├── package.json&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;└── vite.config.js&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That’s it. The entire CapCut-like editor is three files: one React component, one CSS file, and a thin App wrapper.&lt;/p&gt;
&lt;h2 id=&quot;what-we-get-out-of-the-box&quot;&gt;What We Get Out of the Box&lt;/h2&gt;
&lt;p&gt;Because CE.SDK’s video UI is pre-built, we didn’t write any code for these features. They come from &lt;code&gt;createVideoScene()&lt;/code&gt; and the default UI:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Multi-track timeline&lt;/strong&gt; with drag-to-reorder, drag-to-resize&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Trim and split&lt;/strong&gt; — drag clip edges or use the split tool&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Join and arrange&lt;/strong&gt; — drag clips between tracks&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Transform&lt;/strong&gt; — crop, flip, rotate via the inspector&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Playback controls&lt;/strong&gt; — play, pause, seek, scrub&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Text overlays&lt;/strong&gt; — add styled text with the text tool&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Stickers and graphics&lt;/strong&gt; — from the asset library&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Filters and effects&lt;/strong&gt; — LUT-based color grading&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Undo/redo&lt;/strong&gt; — full history stack&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Zoom and pan&lt;/strong&gt; — standard canvas navigation&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The AI plugins add their own UI components (dock buttons, generation panels, context menu items) through the plugin system. We just positioned them where we wanted.&lt;/p&gt;
&lt;h2 id=&quot;going-to-production&quot;&gt;Going to Production&lt;/h2&gt;
&lt;p&gt;To take this from a prototype to production, you need three things:&lt;/p&gt;
&lt;h3 id=&quot;1-license-key&quot;&gt;1. License key&lt;/h3&gt;
&lt;p&gt;Get a free trial key at &lt;a href=&quot;https://img.ly/forms/contact-sales/&quot;&gt;https://img.ly/forms/contact-sales/&lt;/a&gt; and set it in the config (yes, if you want to ship to production you’ll have to talk to our lovely colleagues in sales):&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;js&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; config&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  license: &lt;/span&gt;&lt;span&gt;&apos;YOUR_CESDK_LICENSE_KEY&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;};&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;2-proxy-server&quot;&gt;2. Proxy server&lt;/h3&gt;
&lt;p&gt;Set up a server that forwards AI requests with your API keys:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;js&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; PROXY_URL&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; &apos;https://your-server.com/api/ai-proxy&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;See the &lt;a href=&quot;https://img.ly/docs/cesdk/react/user-interface/ai-integration/proxy-server-61f901/&quot;&gt;CE.SDK Proxy Server guide&lt;/a&gt; for Express.js and other server examples.&lt;/p&gt;
&lt;h3 id=&quot;3-disable-dry-run&quot;&gt;3. Disable dry run&lt;/h3&gt;
&lt;p&gt;Remove &lt;code&gt;dryRun: true&lt;/code&gt; from the AiApps configuration to enable real AI generation.&lt;/p&gt;
&lt;h2 id=&quot;key-takeaways&quot;&gt;Key Takeaways&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;CE.SDK does the heavy lifting.&lt;/strong&gt; The timeline, playback, rendering, and export are all handled by the engine. We wrote zero video processing code.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The plugin system is powerful.&lt;/strong&gt; Six AI capabilities were added with a single &lt;code&gt;addPlugin(AiApps({ ... }))&lt;/code&gt; call. Background removal was one more call.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CSS custom properties make theming painless.&lt;/strong&gt; We matched CapCut’s aesthetic by overriding ~25 CSS variables. No forking, no patching.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The Component Order API is the customization backbone.&lt;/strong&gt; &lt;code&gt;insertOrderComponent&lt;/code&gt; with position-based placement is the cleanest pattern for adding UI elements.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Wrap &lt;code&gt;init&lt;/code&gt; in try/catch.&lt;/strong&gt; CE.SDK’s &lt;code&gt;&amp;#x3C;CreativeEditor&gt;&lt;/code&gt; swallows errors silently. This is the single most important debugging tip.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Pin your package versions.&lt;/strong&gt; All &lt;code&gt;@imgly/*&lt;/code&gt; plugins must match the &lt;code&gt;@cesdk/cesdk-js&lt;/code&gt; version exactly. A version mismatch (like 1.68 vs 1.69) will cause peer dependency conflicts.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&quot;resources&quot;&gt;Resources&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://img.ly/docs/cesdk/react/starterkits/video-editor-e1nlor/&quot;&gt;CE.SDK Video Editor Documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://img.ly/docs/cesdk/react/user-interface/ai-integration/integrate-8e906c/&quot;&gt;AI Integration Guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://img.ly/docs/cesdk/react/user-interface/appearance/theming-4b0938/&quot;&gt;Theming Reference&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://img.ly/docs/cesdk/react/user-interface/customization/quick-start/reorder-components-f6g7h8/&quot;&gt;Rearrange Buttons (Component Order API)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://img.ly/demos/video-ui/web/&quot;&gt;Video Editor Web Demo&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/imgly/cesdk-web-examples&quot;&gt;GitHub Examples&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="https://blog.img.ly/2026/02/blogcover.png" medium="image"/><category>React</category><category>Video Editor</category><category>CE.SDK</category></item><item><title>From Prompt to Editor: Running CE.SDK Inside ChatGPT with the Apps SDK</title><link>https://img.ly/blog/from-prompt-to-editor-running-ce-sdk-inside-chatgpt-with-the-apps-sdk/</link><guid isPermaLink="true">https://img.ly/blog/from-prompt-to-editor-running-ce-sdk-inside-chatgpt-with-the-apps-sdk/</guid><description>We built a technical demo running CE.SDK directly inside ChatGPT using MCP, showing how AI chats can open real, interactive editors. It highlights strict MCP contracts, visual-first UX, and how chat becomes a coordination layer for doing creative work.</description><pubDate>Fri, 19 Dec 2025 15:43:16 GMT</pubDate><content:encoded>&lt;p&gt;With the new ChatGPT Apps SDK and Model Context Protocol (MCP), chat interfaces are starting to look less like Q&amp;#x26;A tools and more like places where work actually happens. To explore what that means for creative workflows, we built a small technical demo: &lt;a href=&quot;https://github.com/imgly/cesdk-web-examples/tree/main/cookbooks-chatgpt-app&quot;&gt;&lt;strong&gt;CE.SDK running directly inside ChatGPT&lt;/strong&gt;.&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;From a user’s perspective, the flow is almost trivial. You ask ChatGPT for something like an ecommerce template. ChatGPT searches our template catalog, selects a matching design, and opens it instantly in a fully interactive CE.SDK editor, right inside the chat interface. What looks like a preview is, in fact, a real editor loaded with a real template scene.&lt;/p&gt;
&lt;p&gt;This isn’t meant as a product announcement. It’s a technical proof of concept showing how creative SDKs can plug directly into AI-native interfaces.&lt;/p&gt;
&lt;h2 id=&quot;cesdk-as-a-chatgpt-app&quot;&gt;CE.SDK as a ChatGPT App&lt;/h2&gt;
&lt;figure class=&quot;kg-embed-card&quot;&gt;
&lt;iframe src=&quot;https://www.youtube.com/embed/AoDqUNVLLJ4?feature=oembed&quot; title=&quot;ChatGPT Opens a Real Design Editor: CE.SDK Inside MCP App (Tech Demo)&quot; loading=&quot;lazy&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; referrerpolicy=&quot;strict-origin-when-cross-origin&quot; allowfullscreen&gt;&lt;/iframe&gt;
&lt;/figure&gt;
&lt;p&gt;The integration is built around a custom MCP server that exposes CE.SDK to ChatGPT as a tool. The server speaks OpenAI’s JSON-RPC–style MCP and implements the standard lifecycle methods (&lt;code&gt;initialize&lt;/code&gt;, &lt;code&gt;tools.list&lt;/code&gt;, &lt;code&gt;tools.call&lt;/code&gt;, &lt;code&gt;resources.read&lt;/code&gt;). It knows about our premium template catalog and emits structured payloads that the frontend understands.&lt;/p&gt;
&lt;p&gt;On the client side, a Next.js app listens to tool output events streamed from ChatGPT, renders CE.SDK widgets, and hydrates them with the payloads returned by the tool, such as a scene URL, placeholder values, or export permissions. Templates are loaded via CE.SDK’s Template API, either from a URL or from a serialized scene string.&lt;/p&gt;
&lt;p&gt;Under the hood, the stack is fairly conventional:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Next.js 15 (App Router)&lt;/li&gt;
&lt;li&gt;CE.SDK Web / CreativeEngine&lt;/li&gt;
&lt;li&gt;A custom MCP handler to normalize JSON-RPC&lt;/li&gt;
&lt;li&gt;Vercel for hosting&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;What’s new is not the technology itself, but the context in which it runs.&lt;/p&gt;
&lt;h2 id=&quot;working-with-mcp-in-practice&quot;&gt;Working with MCP in Practice&lt;/h2&gt;
&lt;p&gt;The hardest part of the demo wasn’t CE.SDK. It was MCP.&lt;/p&gt;
&lt;p&gt;OpenAI’s MCP implementation is extremely strict. Even the smallest schema mismatch can trigger the infamous “TaskGroup 424” error, usually without any hint as to what went wrong. In many cases, the HTTP response is technically successful, but the JSON structure doesn’t match the expected schema closely enough.&lt;/p&gt;
&lt;p&gt;The key lesson here is to treat MCP responses as hard contracts:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Validate every response against a schema (for example with zod).&lt;/li&gt;
&lt;li&gt;Mirror OpenAI’s field names exactly, even for empty or optional capabilities.&lt;/li&gt;
&lt;li&gt;Assume that a 424 almost always means “your JSON shape is wrong”.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Another important insight was how critical visual context is in chat-based tools. If your MCP responses don’t include thumbnails or preview images, ChatGPT will often fall back to rendering links. For creative tools, that immediately breaks the experience. In a chat UI, visuals aren’t an enhancement. They &lt;em&gt;are&lt;/em&gt; the interface.&lt;/p&gt;
&lt;p&gt;State handling also requires a shift in mindset. ChatGPT can replay tool calls, and each prompt effectively creates a new widget instance. You can’t rely on mutating an existing editor. The frontend needs to be idempotent: load scenes from serialized state first, then apply changes. Every tool call should be treated as a fresh render.&lt;/p&gt;
&lt;h2 id=&quot;why-this-pattern-matters&quot;&gt;Why This Pattern Matters&lt;/h2&gt;
&lt;p&gt;This demo points to a broader change in how creative software may be accessed. Chat becomes a coordination layer, not just a conversational one. Instead of explaining how something could be designed, the AI opens the actual editor and lets the user continue from there.&lt;/p&gt;
&lt;p&gt;For CE.SDK, this fits naturally. Editors become embeddable capabilities rather than standalone applications, and AI systems become the entry point into creative workflows. Prompting turns into doing.&lt;/p&gt;
&lt;h2 id=&quot;beyond-openai-the-mcp-ui-standard&quot;&gt;Beyond OpenAI: The MCP UI Standard&lt;/h2&gt;
&lt;p&gt;Although this demo uses OpenAI’s MCP, the architecture maps cleanly to the new MCP UI standard recently introduced by Anthropic. That standard aims to make tool definitions and UI rendering more consistent across models and platforms.&lt;/p&gt;
&lt;p&gt;Because this integration already separates tool logic from UI rendering and relies on structured, explicit payloads, transferring it to Anthropic’s MCP UI model is conceptually straightforward. CE.SDK can act as a reusable creative surface across ChatGPT, Claude, and future AI app ecosystems.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://blog.modelcontextprotocol.io/posts/2025-11-21-mcp-apps/&quot;&gt;You can read more about Anthropic’s MCP UI direction here.&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;This demo is intentionally small and technical, but it highlights a meaningful shift: AI systems that don’t just describe creative outcomes, but open the tools to actually create them.&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><dc:creator>Sven</dc:creator><media:content url="https://blog.img.ly/2025/12/chatgpt-open-design-editor-tempalte.jpg" medium="image"/><category>AI</category><category>MCP</category><category>CE.SDK</category></item><item><title>Introducing CE.SDK Renderer: The Missing Piece of Creative Automation</title><link>https://img.ly/blog/ce-sdk-renderer-creative-automation/</link><guid isPermaLink="true">https://img.ly/blog/ce-sdk-renderer-creative-automation/</guid><description>The CE.SDK Renderer brings fast, reliable, fully licensed server-side rendering to your backend unlocking true end-to-end creative automation with GPU acceleration, perfect fidelity, and scalable image, PDF, and video export.</description><pubDate>Mon, 24 Nov 2025 11:05:15 GMT</pubDate><content:encoded>&lt;p&gt;For years, teams have used &lt;strong&gt;CE.SDK&lt;/strong&gt; to power rich, customizable editing experiences across web, mobile, and desktop. But one critical component of true end-to-end creative automation remained challenging: &lt;strong&gt;server-side rendering&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Until now, export workflows depended on client devices, brittle FFmpeg scripts, or slow headless browser setups. While these approaches worked, they often fell short for enterprise scale workflows, AI pipelines, or high-volume creative automation.&lt;/p&gt;
&lt;p&gt;Today, we’re excited to introduce the &lt;a href=&quot;https://img.ly/docs/cesdk/renderer/get-started/overview-e18f40/&quot;&gt;&lt;strong&gt;CE.SDK Renderer&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;,&lt;/strong&gt; the native, GPU-accelerated server-side rendering engine that completes the CE.SDK ecosystem.&lt;/p&gt;
&lt;p&gt;The Renderer brings CE.SDK’s design engine to your backend with &lt;strong&gt;fast, compliant, enterprise-ready&lt;/strong&gt; export for images, PDFs, and video, enabling organizations to generate media at scale with full fidelity and predictable performance.&lt;/p&gt;
&lt;h2 id=&quot;why-the-cesdk-renderer-matters&quot;&gt;&lt;strong&gt;Why the CE.SDK Renderer Matters&lt;/strong&gt;&lt;/h2&gt;
&lt;h3 id=&quot;rendering-has-been-the-last-barrier-to-full-automation&quot;&gt;&lt;strong&gt;Rendering Has Been the Last Barrier to Full Automation&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Teams could design and generate templates programmatically, but producing the final asset (especially video) was constrained by:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Lack of GPU acceleration&lt;/li&gt;
&lt;li&gt;No native CE.SDK export on servers&lt;/li&gt;
&lt;li&gt;Legal requirements for H.264/H.265&lt;/li&gt;
&lt;li&gt;Unstable browser virtualization&lt;/li&gt;
&lt;li&gt;Node.js memory and performance limits&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The CE.SDK Renderer solves all of this in one fell swoop with a native Linux binary packaged as a Docker container, built explicitly for backend media generation.&lt;/p&gt;
&lt;p&gt;Many of our SaaS customers have told us the same thing: they want a safe, supported alternative to FFmpeg that they don’t have to maintain themselves.&lt;/p&gt;
&lt;h2 id=&quot;what-the-cesdk-renderer-delivers&quot;&gt;&lt;strong&gt;What the CE.SDK Renderer Delivers&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;The CE.SDK Renderer fully unlocks its value in the &lt;a href=&quot;https://img.ly/&quot;&gt;IMG.LY&lt;/a&gt; ecosystem by providing a creative automation infrastructure layer for the user facing portion of &lt;a href=&quot;https://img.ly/&quot;&gt;IMG.LY&lt;/a&gt;’s editor SDK. It enables two core capabilities that modern creative automation depends on: &lt;strong&gt;scalable rendering&lt;/strong&gt; and &lt;strong&gt;fully compliant video export:&lt;/strong&gt;&lt;/p&gt;
&lt;h3 id=&quot;gpu-accelerated-export-performance&quot;&gt;&lt;strong&gt;GPU-Accelerated Export Performance&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Native code and GPU acceleration allow extremely fast rendering. Ideal for heavy scenes, 4K+ requests, and large batch pipelines.&lt;/p&gt;
&lt;h3 id=&quot;fully-licensed-video-on-the-backend-h264h265&quot;&gt;&lt;strong&gt;Fully Licensed Video on the Backend (H.264/H.265)&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;You can render MP4 video entirely on your servers with full licensing compliance. Powered by &lt;a href=&quot;https://fluendo.com/&quot;&gt;Fluendo&lt;/a&gt;’s AV stack, the Renderer provides enterprise-grade legal protection for H.264/H.265, eliminating the risks associated with unlicensed or open-source encoding. This gives teams the confidence to run video-heavy pipelines at scale.&lt;/p&gt;
&lt;h3 id=&quot;built-for-scale&quot;&gt;&lt;strong&gt;Built for Scale&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Perfect for workloads generating hundreds or thousands of creative variations: ads, postcards, dynamic social assets, catalogs, and more on stable maintainable infrastructure supplanting FFmpeg or brittle pipelines based on headless browsers.&lt;/p&gt;
&lt;h3 id=&quot;perfect-visual-fidelity&quot;&gt;&lt;strong&gt;Perfect Visual Fidelity&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Because it uses the same CE.SDK engine that powers the editing experience, the Renderer produces output that matches exactly what users see in the editor. This fidelity is critical for automated pipelines, ensuring every exported asset is accurate, consistent, and visually correct.&lt;/p&gt;
&lt;p&gt;Together, these unlock a complete creative automation workflow: &lt;strong&gt;design → generate → render&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id=&quot;how-it-works&quot;&gt;How It Works&lt;/h2&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 2000px) 2000px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;2000&quot; height=&quot;1125&quot; src=&quot;https://img.ly/_astro/Serverside-rendering_--1-_Z9YgrH.webp&quot; srcset=&quot;/_astro/Serverside-rendering_--1-_Z1iOkFH.webp 640w, /_astro/Serverside-rendering_--1-_Z2nSVco.webp 750w, /_astro/Serverside-rendering_--1-_Z7rv4f.webp 828w, /_astro/Serverside-rendering_--1-_c5Alb.webp 1080w, /_astro/Serverside-rendering_--1-_Z1YOsnU.webp 1280w, /_astro/Serverside-rendering_--1-_2fY3Ah.webp 1668w, /_astro/Serverside-rendering_--1-_Z9YgrH.webp 2000w&quot;&gt;&lt;/p&gt;
&lt;h3 id=&quot;async-exports&quot;&gt;Async Exports&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;A CE.SDK scene is created via browser editor, mobile app, or Node.js.&lt;/li&gt;
&lt;li&gt;The scene is sent to the &lt;strong&gt;Renderer&lt;/strong&gt;, running in Docker.&lt;/li&gt;
&lt;li&gt;The Renderer outputs: PNG/JPEG, PDF, MP4 (H.264/H.265)&lt;/li&gt;
&lt;li&gt;The rendered asset is returned to your application or batch pipeline.&lt;/li&gt;
&lt;li&gt;Users can continue editing while export happens asynchronously.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This removes export load from the client entirely and opens the door for sophisticated backend media generation.&lt;/p&gt;
&lt;h3 id=&quot;batch-rendering&quot;&gt;Batch Rendering&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;Multiple CE.SDK templates are prepared for automation.&lt;/li&gt;
&lt;li&gt;An external data source (CSV, JSON, API) is connected to the templates to generate all design variants.&lt;/li&gt;
&lt;li&gt;Each resulting scene is sent as a separate job to the &lt;strong&gt;Renderer&lt;/strong&gt; running in Docker.&lt;/li&gt;
&lt;li&gt;The Renderer outputs PNG/JPEG, PDF, or MP4 for every variant.&lt;/li&gt;
&lt;li&gt;All assets are returned to your backend or stored in your pipeline, enabling large-scale automated creative production.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&quot;example-use-case-hyper-personalized-video-ads-at-massive-scale&quot;&gt;&lt;strong&gt;Example Use Case: Hyper-Personalized Video Ads at Massive Scale&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;One of our customers runs large demographic-targeted campaigns where a single master video must transform into tens of thousands of personalized variants. Instead of showing the same creative to everyone, they tailor each video to the viewer’s profile think &lt;em&gt;German, mid-30s, dog owner&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;From one base creative, their system dynamically generates up to &lt;strong&gt;100,000 unique versions&lt;/strong&gt;. Their backend assembles all scene variations in Node.js, then hands them off to the CE.SDK Renderer to produce the final video files.&lt;/p&gt;
&lt;p&gt;Before adopting CE.SDK Renderer, this workflow was painful: brittle FFmpeg scripts, unpredictable render results across machines, codec licensing concerns, and processing times that stretched over days.&lt;/p&gt;
&lt;p&gt;Maintaining this setup consumed engineering time that would have been better spent on improving the creative logic itself.&lt;/p&gt;
&lt;p&gt;With CE.SDK Renderer, the entire pipeline became both scalable and predictable. They now render thousands of variants with consistent results, fully licensed codecs, and throughput high enough to keep up with real-world campaign demands. Instead of fighting their rendering stack, their team can focus on creative automation and campaign performance.&lt;/p&gt;
&lt;p&gt;This is just one example of how CE.SDK Renderer enables production-grade, high-volume personalized media at scale something that previously required complex custom infrastructure.&lt;/p&gt;
&lt;h2 id=&quot;licensing-options&quot;&gt;&lt;strong&gt;Licensing Options&lt;/strong&gt;&lt;/h2&gt;
&lt;h3 id=&quot;open-source-renderer&quot;&gt;&lt;strong&gt;Open-Source Renderer&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;A lightweight version using open-source codecs. It’s fast, simple to run, and great for prototyping or early-stage products. Because it doesn’t include licensed H.264/H.265 encoding, it isn’t suitable for enterprise video production or legally compliant distribution.&lt;/p&gt;
&lt;h3 id=&quot;av-licensed-renderer-fluendo&quot;&gt;&lt;strong&gt;AV-Licensed Renderer (Fluendo)&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;This edition includes fully licensed H.264/H.265 encoding powered by &lt;a href=&quot;https://fluendo.com/&quot;&gt;Fluendo&lt;/a&gt;, ensuring your video exports meet all patent and compliance requirements. It uses concurrency-based licensing and is designed for production workloads, regulated industries, and any platform generating video at scale.&lt;/p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;By bringing fast, compliant, and scalable rendering to the backend, the CE.SDK Renderer removes one of the biggest engineering hurdles in creative automation. Whether you’re generating hundreds or hundreds of thousands of assets, it gives you the tools to build stable, modern media pipelines without the maintenance burden. The simplest path to production-grade creative automation.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/docs/cesdk/&quot;&gt;Start your free trial&lt;/a&gt; today and get started in minutes.&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="https://blog.img.ly/2025/11/render-automation-nodejs-alternative-creative-sdk-imgly-export-video.jpg" medium="image"/><category>Creative Automation</category><category>Server-side Video</category></item><item><title>How to Embed an Editable InDesign Template in Your Website</title><link>https://img.ly/blog/embed-edit-automate-indesign-files-in-the-browser-complete-guide-2025/</link><guid isPermaLink="true">https://img.ly/blog/embed-edit-automate-indesign-files-in-the-browser-complete-guide-2025/</guid><description>InDesign files hold valuable design structure - yet remain stuck offline. This article shows how CE.SDK turns IDML templates into web-editable layouts with full fidelity, unlocking collaboration and automation impossible in desktop-only workflows.</description><pubDate>Tue, 28 Oct 2025 12:29:03 GMT</pubDate><content:encoded>&lt;h2 id=&quot;why-editing-indesign-files-in-the-browser-matters-now&quot;&gt;&lt;strong&gt;Why Editing InDesign Files in the Browser Matters Now&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Adobe InDesign has long been the standard for high-fidelity layout and print design. Yet for many teams, its power comes with friction: it’s desktop-bound, collaboration-limited, and inaccessible to clients or non-designers who simply need to make minor edits.&lt;/p&gt;
&lt;p&gt;Creative work requires ever shorter feedback loops and is becoming more and more accessible to the average users, hence organizations are looking for ways to &lt;strong&gt;bring InDesign workflows into the browser&lt;/strong&gt; to make templates editable, collaborative, and automatable.&lt;/p&gt;
&lt;p&gt;At the same time, businesses are under pressure to modernize creative production. Marketing teams need to localize campaigns at scale, SaaS platforms want to let users personalize assets, and agencies aim to deliver editable templates instead of static files. The question naturally arises:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;“Can I edit InDesign files in a browser?”&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Until recently, the answer was “not really.” Adobe offers &lt;em&gt;Share for Review&lt;/em&gt; for commenting, &lt;em&gt;InCopy on the Web&lt;/em&gt; for limited text changes, and &lt;em&gt;Adobe Express&lt;/em&gt; for simplified exports, but none provide full-fidelity, browser-native editing or the ability to embed such functionality into your own product.&lt;/p&gt;
&lt;p&gt;That’s where &lt;strong&gt;CE.SDK&lt;/strong&gt; enters the picture.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;CE.SDK is an embeddable creative editor that powers photo, video, and design workflows directly in the browser. It’s deeply customizable and extensible, enabling developers to tailor every aspect of the editing experience. The same SDK works cross-platform, Web, iOS, Android, Desktop, and Server, so teams can build consistent creative tools across environments.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;By combining CE.SDK’s robust editing engine with the &lt;a href=&quot;https://img.ly/demos/indesign-template-import/web/&quot;&gt;&lt;strong&gt;InDesign Importer&lt;/strong&gt;&lt;/a&gt;, you can now &lt;strong&gt;bring InDesign templates (IDML files) into a fully fledged web-based design editor&lt;/strong&gt; while preserving essential layout, style, and asset information. The result: true browser editing of InDesign content, without the limitations of traditional desktop software.&lt;/p&gt;
&lt;h2 id=&quot;the-current-landscape-whats-possible-and-what-isnt-with-indesign-on-the-web&quot;&gt;&lt;strong&gt;The Current Landscape: What’s Possible (and What Isn’t) with InDesign on the Web&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;While creative teams increasingly expect collaborative, browser-based tools, Adobe InDesign remains deeply rooted in its desktop heritage. Its powerful layout engine and proprietary file structure were never designed for real-time, cloud-native editing. As a result, teams who rely on InDesign often face friction when trying to make designs accessible to clients or other stakeholders online.&lt;/p&gt;
&lt;p&gt;Adobe has made incremental steps toward the web, but these tools still serve limited purposes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://helpx.adobe.com/indesign/using/share-for-review.html?utm_source=chatgpt.com&quot;&gt;&lt;strong&gt;Share for Review&lt;/strong&gt;&lt;/a&gt; – enables commenting and approval workflows in the browser, but doesn’t allow editing or layout changes.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://helpx.adobe.com/indesign/using/incopy-web.html?utm_source=chatgpt.com&quot;&gt;&lt;strong&gt;InCopy on the Web (beta)&lt;/strong&gt;&lt;/a&gt; – offers browser-based text editing within locked layouts. It’s useful for copy review, yet visual elements remain untouchable.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://helpx.adobe.com/indesign/using/export-to-express.html&quot;&gt;&lt;strong&gt;Adobe Express export&lt;/strong&gt;&lt;/a&gt; – lets designers repurpose InDesign layouts as simplified templates for lightweight editing, but the process is one-way and loses much of InDesign’s fidelity and control.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For teams that need to &lt;strong&gt;deliver editable templates, enable client-side personalization, or embed creative workflows inside their own platforms&lt;/strong&gt;, these options aren’t sufficient. They lack extensibility, API access, and the ability to maintain brand-level control in a web environment.&lt;/p&gt;
&lt;h3 id=&quot;beyond-adobe-existing-alternatives&quot;&gt;&lt;strong&gt;Beyond Adobe: Existing Alternatives&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Several third-party tools have tried to fill the gap:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://viva.systems/designer&quot;&gt;&lt;strong&gt;VivaDesigner&lt;/strong&gt;&lt;/a&gt; mirrors parts of InDesign’s functionality in a browser, but operates as a self-contained product rather than an embeddable SDK.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.siliconpublishing.com/designer/&quot;&gt;&lt;strong&gt;Silicon Designer&lt;/strong&gt;&lt;/a&gt; builds on InDesign Server to power web-to-print solutions, but depends on heavy backend infrastructure and costly licensing.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.photopea.com/&quot;&gt;&lt;strong&gt;Photopea&lt;/strong&gt;&lt;/a&gt; provides impressive browser editing for layered graphics and basic IDML files, yet lacks enterprise-grade extensibility or workflow integration.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Each of these solutions demonstrates what’s possible, but none offer a developer-friendly foundation for building editing workflows and experience on-top of InDesign in the browser.&lt;/p&gt;
&lt;h3 id=&quot;where-imglys-cesdk-fits-in&quot;&gt;&lt;strong&gt;Where IMG.LY’s CE.SDK Fits In&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;This is the gap that CE.SDK fills.&lt;/p&gt;
&lt;p&gt;Instead of emulating InDesign’s desktop application, CE.SDK focuses on &lt;strong&gt;data translation and browser-native rendering&lt;/strong&gt;. Its &lt;strong&gt;InDesign Importer&lt;/strong&gt; converts the open IDML format into CE.SDK’s optimized scene format retaining layout structure, typography, and key design elements so users can edit and export designs directly in the browser and all other platforms supported by CE.SDK.&lt;/p&gt;
&lt;p&gt;For developers, this approach unlocks a new level of flexibility:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Embed an editable InDesign experience in any web platform.&lt;/li&gt;
&lt;li&gt;Integrate design editing into DAMs, CMSs, or creative automation workflows.&lt;/li&gt;
&lt;li&gt;Customize UI, behaviors, and integrations to match existing systems.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In short, CE.SDK transforms what used to be static, desktop-bound InDesign files into &lt;strong&gt;interactive, web-based design templates&lt;/strong&gt;, without compromising control or scalability.&lt;/p&gt;
&lt;h2 id=&quot;introducing-the-cesdk-indesign-importer&quot;&gt;&lt;strong&gt;Introducing the CE.SDK InDesign Importer&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;video src=&quot;https://blog.img.ly/2025/11/indesign-importer-creative-sdk-martech-saas-whitelabel-imgly.mp4&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;p&gt;Moving professional InDesign layouts into the browser isn’t just about file conversion, it’s about &lt;strong&gt;accurately translating complex design data&lt;/strong&gt; into a web-native format that can be rendered, edited, and automated.&lt;/p&gt;
&lt;p&gt;That’s exactly what the &lt;strong&gt;CE.SDK InDesign Importer&lt;/strong&gt; does.&lt;/p&gt;
&lt;p&gt;The importer acts as a bridge between &lt;strong&gt;Adobe InDesign’s IDML format&lt;/strong&gt; and &lt;strong&gt;CE.SDK’s scene model&lt;/strong&gt;, transforming desktop-authored layouts into editable, browser-ready projects. Once an &lt;code&gt;.idml&lt;/code&gt; file is exported from InDesign, the importer reconstructs its layers, assets, and properties, packaging them into a CE.SDK scene archive that can be opened instantly inside any CE.SDK instance.&lt;/p&gt;
&lt;p&gt;Explore the &lt;a href=&quot;https://img.ly/demos/indesign-template-import/web/&quot;&gt;InDesign Template Import Demo&lt;/a&gt; for a comprehensive example.&lt;/p&gt;
&lt;h3 id=&quot;a-stand-alone-module-built-for-integration&quot;&gt;&lt;strong&gt;A Stand-Alone Module, Built for Integration&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Unlike the core CE.SDK editor, the InDesign Importer is distributed as a &lt;strong&gt;stand-alone package&lt;/strong&gt; that you can integrate into any workflow. It’s available via npm as&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://www.npmjs.com/package/@imgly/idml-importer/v/1.1.1&quot;&gt;&lt;code&gt;@imgly/idml-importer&lt;/code&gt;&lt;/a&gt;,&lt;/p&gt;
&lt;p&gt;allowing developers to run imports in their own build systems, servers, or client-side applications before loading the resulting scene into CE.SDK.&lt;/p&gt;
&lt;p&gt;This separation makes it easy to slot the importer into existing pipelines (for example, automated template ingestion systems, DAM integrations, or internal pre-processing tools) without requiring the full editor runtime.&lt;/p&gt;
&lt;h3 id=&quot;how-it-fits-into-the-cesdk-ecosystem&quot;&gt;&lt;strong&gt;How It Fits into the CE.SDK Ecosystem&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;CE.SDK&lt;/strong&gt; (CreativeEditor SDK) is an &lt;strong&gt;embeddable creative editor&lt;/strong&gt; powering photo, video, and design workflows across &lt;strong&gt;Web, iOS, Android, Desktop, and Server&lt;/strong&gt;. It offers a modular, extensible engine and UI framework that teams can tailor to any brand or use case.&lt;/p&gt;
&lt;p&gt;The InDesign Importer extends that ecosystem by unlocking compatibility with one of the most widely used layout tools in the world. Together, they enable a complete pipeline:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;InDesign (IDML) → @imgly/idml-importer → CE.SDK Scene File → Browser Editing &amp;#x26; Automation&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This means existing InDesign templates can become live, editable browser experiences, without rebuilding designs manually or deploying heavy server infrastructure.&lt;/p&gt;
&lt;h3 id=&quot;what-the-importer-delivers&quot;&gt;&lt;strong&gt;What the Importer Delivers&lt;/strong&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;File-Format Translation&lt;/strong&gt; – Converts IDML files into CE.SDK scene archives while preserving layout hierarchy, positioning, and grouping.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Asset Bundling&lt;/strong&gt; – Packages fonts, embedded images, and color data for immediate use in CE.SDK.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Color Mapping&lt;/strong&gt; – Converts CMYK values into RGB for web rendering (native CMYK support is in development).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Element Preservation&lt;/strong&gt; – Maintains grouping, rotation, shapes (rectangles, ovals, polygons, lines), gradients, transparency, and strokes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Developer Flexibility&lt;/strong&gt; – Import locally or at scale, feed the resulting scene into CE.SDK’s API, or integrate into automated asset pipelines.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;real-world-use-cases--workflows&quot;&gt;&lt;strong&gt;Real-World Use Cases &amp;#x26; Workflows&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Once an InDesign file becomes editable in the browser, entirely new workflows open up, from collaborative editing to automated content generation.&lt;/p&gt;
&lt;p&gt;The CE.SDK InDesign Importer enables organizations to extend proven InDesign templates into scalable, web-native experiences:&lt;/p&gt;
&lt;h3 id=&quot;web-to-print-platforms&quot;&gt;&lt;strong&gt;Web-to-Print Platforms&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Allow end users to personalize marketing collateral, business cards, or packaging directly in a browser editor while maintaining the designer’s original layout integrity. This is the core of a &lt;a href=&quot;https://img.ly/use-cases/web-to-print-design-tool/&quot;&gt;web-to-print design tool&lt;/a&gt;: the template stays locked where it matters, and the customer only edits what you let them edit.&lt;/p&gt;
&lt;h3 id=&quot;brand-template-portals&quot;&gt;&lt;strong&gt;Brand Template Portals&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Empower distributed teams, agencies, or franchise partners to create on-brand materials without ever touching desktop software. Designers upload InDesign templates once; users edit and export variations on demand.&lt;/p&gt;
&lt;h3 id=&quot;creative-automation-systems&quot;&gt;&lt;strong&gt;Creative Automation Systems&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Combine CE.SDK with data pipelines to automatically generate localized or personalized assets at scale, replacing time-consuming manual layout work with programmable design workflows.&lt;/p&gt;
&lt;h3 id=&quot;client-collaboration&quot;&gt;&lt;strong&gt;Client Collaboration&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Deliver interactive proofing experiences where clients can adjust copy, swap images, or approve layouts in a controlled browser environment, eliminating the “export–review–revise” loop typical of InDesign-based projects.&lt;/p&gt;
&lt;p&gt;Each of these use cases builds on the same foundation: reliable IDML translation plus CE.SDK’s flexible editing engine.&lt;/p&gt;
&lt;p&gt;That combination makes the Importer not just a conversion tool, but a bridge to &lt;strong&gt;entirely new creative business models&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id=&quot;cesdk-vs-traditional-indesign-server--other-alternatives&quot;&gt;&lt;strong&gt;CE.SDK vs. Traditional InDesign Server &amp;#x26; Other Alternatives&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;For teams exploring browser-based design editing, the landscape typically centers on three paths: &lt;strong&gt;InDesign Server&lt;/strong&gt;, &lt;strong&gt;web-to-print middleware&lt;/strong&gt;, or &lt;strong&gt;browser SDKs&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The CE.SDK InDesign Importer offers a modern alternative to all three.&lt;/p&gt;















































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Feature / Capability&lt;/th&gt;&lt;th&gt;&lt;strong&gt;InDesign Server&lt;/strong&gt;&lt;/th&gt;&lt;th&gt;&lt;strong&gt;Third-Party Tools (VivaDesigner, Silicon Designer)&lt;/strong&gt;&lt;/th&gt;&lt;th&gt;&lt;strong&gt;CE.SDK + InDesign Importer&lt;/strong&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Editing Fidelity&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Full but desktop-rendered&lt;/td&gt;&lt;td&gt;Partial; varies by implementation&lt;/td&gt;&lt;td&gt;High; layout preserved via IDML&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Web Accessibility&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Limited; server-side only&lt;/td&gt;&lt;td&gt;Browser UI, but closed systems&lt;/td&gt;&lt;td&gt;Fully client-side, browser-native&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Embeddable / SDK&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;No&lt;/td&gt;&lt;td&gt;Proprietary&lt;/td&gt;&lt;td&gt;Yes + modular npm packages&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Infrastructure&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Requires Adobe licensing &amp;#x26; server setup&lt;/td&gt;&lt;td&gt;Vendor-hosted&lt;/td&gt;&lt;td&gt;Lightweight; deploy anywhere&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Extensibility&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Restricted scripting&lt;/td&gt;&lt;td&gt;Limited&lt;/td&gt;&lt;td&gt;Full API &amp;#x26; UI customization&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Cost / Licensing&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;High, per-instance&lt;/td&gt;&lt;td&gt;Varies; often enterprise-only&lt;/td&gt;&lt;td&gt;Predictable developer friendly licensing&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;CE.SDK’s approach eliminates the dependency on server-side rendering and proprietary hosting, providing &lt;strong&gt;a developer-first, browser-native foundation&lt;/strong&gt; for creative editing.&lt;/p&gt;
&lt;p&gt;By translating InDesign layouts into open CE.SDK scenes, it combines &lt;strong&gt;professional-grade fidelity&lt;/strong&gt; with the &lt;strong&gt;flexibility of modern web architecture&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id=&quot;conclusion--a-new-era-for-indesign-workflows&quot;&gt;&lt;strong&gt;Conclusion – A New Era for InDesign Workflows&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;For years, creative teams have struggled to bridge the gap between &lt;strong&gt;InDesign’s print-grade precision&lt;/strong&gt; and the &lt;strong&gt;web’s flexibility and scalability&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The &lt;strong&gt;CE.SDK InDesign Importer&lt;/strong&gt; closes that gap, turning traditional &lt;code&gt;.indd&lt;/code&gt; projects into browser-ready, editable templates that can live inside any modern application.&lt;/p&gt;
&lt;p&gt;Whether you’re building a &lt;strong&gt;web-to-print platform&lt;/strong&gt;, empowering clients through &lt;strong&gt;self-service editing&lt;/strong&gt;, or connecting templates to &lt;strong&gt;creative-automation pipelines&lt;/strong&gt;, CE.SDK provides the foundation to make it happen, with full developer control and a consistent experience across Web, Mobile, and Desktop.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Explore the live demo: &lt;a href=&quot;https://img.ly/demos/indesign-template-import/web/&quot;&gt;InDesign Template Import Demo&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Try the importer: &lt;a href=&quot;https://www.npmjs.com/package/@imgly/idml-importer/v/1.1.1&quot;&gt;&lt;code&gt;@imgly/idml-importer&lt;/code&gt; on npm&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Learn more about CE.SDK: &lt;a href=&quot;https://img.ly/products/creative-sdk/&quot;&gt;https://img.ly/products/creative-sdk/&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;frequently-asked-questions&quot;&gt;&lt;strong&gt;Frequently Asked Questions&lt;/strong&gt;&lt;/h2&gt;
&lt;h3 id=&quot;can-i-edit-an-indesign-file-directly-in-a-browser&quot;&gt;&lt;strong&gt;Can I edit an InDesign file directly in a browser?&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Not with Adobe’s native tools alone. &lt;em&gt;Share for Review&lt;/em&gt; and &lt;em&gt;InCopy on the Web&lt;/em&gt; only allow commenting or text changes. With the &lt;strong&gt;CE.SDK InDesign Importer&lt;/strong&gt;, however, you can convert an exported IDML file into a browser-editable format that retains layout, fonts, and key visual elements.&lt;/p&gt;
&lt;h3 id=&quot;what-file-formats-does-the-importer-support&quot;&gt;&lt;strong&gt;What file formats does the importer support?&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;The importer reads &lt;strong&gt;IDML&lt;/strong&gt; files exported from Adobe InDesign and converts them into &lt;strong&gt;CE.SDK scene archives&lt;/strong&gt;. These can then be opened in CE.SDK for full browser editing and exported again to formats such as &lt;strong&gt;PDF, PNG, or JSON&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id=&quot;does-it-require-adobe-indesign-server&quot;&gt;&lt;strong&gt;Does it require Adobe InDesign Server?&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;No. The &lt;strong&gt;@imgly/idml-importer&lt;/strong&gt; runs independently. It’s a standalone npm package and doesn’t depend on InDesign Server or any Adobe infrastructure. You only need an IDML export from InDesign.&lt;/p&gt;
&lt;h3 id=&quot;is-cmyk-color-supported&quot;&gt;&lt;strong&gt;Is CMYK color supported?&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Currently, CMYK values are automatically translated into RGB for accurate web rendering. Native CMYK support is planned in future updates.&lt;/p&gt;
&lt;h3 id=&quot;can-i-automate-bulk-imports&quot;&gt;&lt;strong&gt;Can I automate bulk imports?&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Yes. Because the importer is installable via npm, you can integrate it into scripts or pipelines to process large template libraries automatically before loading them into CE.SDK.&lt;/p&gt;
&lt;h3 id=&quot;do-imported-indesign-templates-remain-editable-for-non-designers&quot;&gt;&lt;strong&gt;Do imported InDesign templates remain editable for non-designers?&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Absolutely. Once loaded into CE.SDK, templates can be edited through a customizable browser interface, ideal for client portals, marketing platforms, or self-service brand editors.&lt;/p&gt;
</content:encoded><dc:creator>Jan</dc:creator><media:content url="https://blog.img.ly/2025/10/edit-indesign-file-in-browser.jpg" medium="image"/><category>Insights</category><category>Design Editor</category><category>Creative Workflows</category></item><item><title>Animate Between Images - AI-Native Video Workflows with CE.SDK and Veo 3</title><link>https://img.ly/blog/animate-between-images-ai-native-video-workflows-with-ce-sdk-and-veo-3/</link><guid isPermaLink="true">https://img.ly/blog/animate-between-images-ai-native-video-workflows-with-ce-sdk-and-veo-3/</guid><description>See how we integrated Veo 3.1 into CE.SDK to animate between images in seconds. With generation times as low as 9 s, this demo shows how easily you can embed AI-native workflows from stills to smooth video clips directly inside your Creative Editor.</description><pubDate>Wed, 22 Oct 2025 14:10:31 GMT</pubDate><content:encoded>&lt;p&gt;With the release of &lt;a href=&quot;https://aistudio.google.com/models/veo-3&quot;&gt;&lt;strong&gt;Veo 3.1&lt;/strong&gt;,&lt;/a&gt; we wanted to show just how effortless it is to embed generative AI capabilities directly into creative workflows. In this quick demo, we integrated Veo 3 into &lt;strong&gt;CreativeEditor SDK (CE.SDK)&lt;/strong&gt; enabling users to &lt;strong&gt;animate between two still images&lt;/strong&gt; in just a few clicks.&lt;/p&gt;
&lt;p&gt;&lt;video src=&quot;https://storage.googleapis.com/imgly-static-assets/static/blog/videos/ai-launch/IMG-LY-CE-SDK-veo-3-1-video-design-editor-sdk.mp4&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;h3 id=&quot;from-still-images-to-motion&quot;&gt;From Still Images to Motion&lt;/h3&gt;
&lt;p&gt;In our demo, we start with two images of the same person, one wearing a hat, the other without. Inside the editor, users simply select both images, click the &lt;strong&gt;AI context button&lt;/strong&gt;, and choose &lt;strong&gt;“Animate between images.”&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The images are loaded into a side panel where users can optionally add a &lt;strong&gt;text prompt&lt;/strong&gt; to guide the transition. Once generated, the resulting short video is placed directly on the canvas ready for editing, compositing, or export.&lt;/p&gt;
&lt;p&gt;What’s particularly impressive is the &lt;strong&gt;generation speed&lt;/strong&gt;. In this example, Veo 3.1 produced a smooth &lt;strong&gt;8-second transition in just 9 seconds&lt;/strong&gt; a major improvement compared to earlier versions. This speed makes iterative creative workflows feel fluid and responsive, bridging the gap between prompt-driven generation and real-time editing.&lt;/p&gt;
&lt;h3 id=&quot;try-it-out&quot;&gt;Try It Out&lt;/h3&gt;
&lt;p&gt;&lt;img alt=&quot;Try out Veo 3.1&amp;#39;s magic in CE.SDK.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1312px) 1312px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1312&quot; height=&quot;310&quot; src=&quot;https://img.ly/_astro/veo-fal-ai-video-editor-sdk-ce_sdk-imgly_sD9cv.webp&quot; srcset=&quot;/_astro/veo-fal-ai-video-editor-sdk-ce_sdk-imgly_fSI3j.webp 640w, /_astro/veo-fal-ai-video-editor-sdk-ce_sdk-imgly_Z2fext8.webp 750w, /_astro/veo-fal-ai-video-editor-sdk-ce_sdk-imgly_Z1SRm7j.webp 828w, /_astro/veo-fal-ai-video-editor-sdk-ce_sdk-imgly_Hlv7V.webp 1080w, /_astro/veo-fal-ai-video-editor-sdk-ce_sdk-imgly_1Vk21l.webp 1280w, /_astro/veo-fal-ai-video-editor-sdk-ce_sdk-imgly_sD9cv.webp 1312w&quot;&gt;&lt;/p&gt;
&lt;p&gt;You can check out the implementation on &lt;a href=&quot;https://github.com/imgly/plugins/blob/main/examples/ai/src/pages/Veo31Example.tsx&quot;&gt;GitHub&lt;/a&gt; and give Veo 3.1 a spin inside your CE.SDK instance (sign up for a &lt;a href=&quot;https://img.ly/docs/cesdk/&quot;&gt;free trial&lt;/a&gt; if you haven’t already).&lt;/p&gt;
&lt;h3 id=&quot;a-glimpse-of-ai-native-editing&quot;&gt;A Glimpse of AI-Native Editing&lt;/h3&gt;
&lt;p&gt;This simple feature highlights how easy it is to make your editor &lt;strong&gt;AI-native,&lt;/strong&gt; combining traditional editing tools with generative intelligence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Practical use cases include:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;E-commerce&lt;/strong&gt;&lt;br&gt;
Show products “in action” or animate between styles and configurations.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Marketing&lt;/strong&gt;&lt;br&gt;
Create quick product reveal animations from static assets.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Content creation&lt;/strong&gt;&lt;br&gt;
Generate short motion clips or “tween” between creative scenes.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And if users want longer clips, they can simply &lt;strong&gt;add another 8-second track&lt;/strong&gt; and &lt;strong&gt;transition seamlessly&lt;/strong&gt; into the next.&lt;/p&gt;
&lt;h3 id=&quot;the-future-of-creative-workflows&quot;&gt;The Future of Creative Workflows&lt;/h3&gt;
&lt;p&gt;With Veo 3 integrated, CE.SDK becomes a powerful playground for AI-driven creativity from image-to-video to scene interpolation and contextual animation.&lt;/p&gt;
&lt;p&gt;We already empowered over 600 innovative startups, government entities, and Fortune 500 companies to add powerful design, video, and photo editing workflows to their products. &lt;a href=&quot;https://img.ly/forms/contact-sales/&quot;&gt;Get in touch&lt;/a&gt;, to see how we can do the same for you.&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><dc:creator>Marcin</dc:creator><media:content url="https://blog.img.ly/2025/10/veo-3_1-creative-editor-sdk-video-editor-imgly-2.jpg" medium="image"/><category>AI</category><category>Video Editing</category><category>Image2Video</category></item><item><title>A Guide to Creative Automation with CE.SDK &amp; Javascript</title><link>https://img.ly/blog/a-guide-to-creative-automation-with-ce-sdk-javascript/</link><guid isPermaLink="true">https://img.ly/blog/a-guide-to-creative-automation-with-ce-sdk-javascript/</guid><description>Discover how to build end-to-end creative automation with CE.SDK and React. This guide shows you how to create templates with placeholders, connect them to external data, and automatically generate personalized assets at scale.</description><pubDate>Fri, 05 Sep 2025 11:14:31 GMT</pubDate><content:encoded>&lt;figure class=&quot;kg-embed-card&quot;&gt;
&lt;iframe src=&quot;https://www.youtube.com/embed/GtORNyzjOq0?feature=oembed&quot; title=&quot;Creative Automation with CE.SDK and React / JavaScript – Full Guide&quot; loading=&quot;lazy&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&quot; referrerpolicy=&quot;strict-origin-when-cross-origin&quot; allowfullscreen&gt;&lt;/iframe&gt;
&lt;/figure&gt;
&lt;h2 id=&quot;what-is-creative-automation&quot;&gt;What is Creative Automation?&lt;/h2&gt;
&lt;p&gt;Creative Automation is the process of generating visual content, images, videos, designs, programmatically, often in response to structured inputs like product data, user attributes, or campaign parameters.&lt;/p&gt;
&lt;p&gt;Instead of designing each asset manually, creative automation enables you to define templates and rules that the system uses to produce content at scale.&lt;/p&gt;
&lt;p&gt;This approach accelerates production workflows and enables personalization at scale, multivariate testing, and omni-channel consistency without stretching design resources.&lt;/p&gt;
&lt;h2 id=&quot;creative-automation--generative-ai&quot;&gt;Creative Automation &amp;#x26; Generative AI&lt;/h2&gt;
&lt;p&gt;Of course, you cannot talk about creative automation in this day and age without also discussing the role of generative AI. Gen AI introduces intelligent content generation into the pipeline making it easy to generate headlines, product descriptions and entire design elements like illustrations, videos or voiceovers and background music.&lt;/p&gt;
&lt;p&gt;When paired with a robust editing and rendering engine like &lt;strong&gt;CE.SDK&lt;/strong&gt;, generative AI becomes even more powerful. A design editor that sits at the intersection between the human decision maker and raw AI generated creatives ensures that those creatives can be curated, refined and deployed in an orderly fashion. This usually works by employing design templates that provide extension points for dynamic data through variables and placeholders and that AI generated content can be merged into for a brand-safe final creative.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/case-studies/omneky/&quot;&gt;Our customer Omneky&lt;/a&gt; provides an excellent example of how these components can be orchestrated for maximum efficiency and productivity. Omneky is an ad tech platform that uses generative AI to automatically generate creatives and ad copy for a company based on their product and positioning, these are then interpolated into CE.SDK template to create ad variations. These variations are then tested across ad channels and the results used to inform further refinement and ad selection. This is a prime example of how to combine creative automation and editors to tame the chaotic fountain of AI creatives.&lt;/p&gt;
&lt;h2 id=&quot;how-cesdk-enables-creative-automation&quot;&gt;How CE.SDK Enables Creative Automation&lt;/h2&gt;
&lt;p&gt;Let us get into the meat and potatoes of creative automation and sketch out the requirements of a creative automation solution and an end to end example with CE.SDK. &lt;a href=&quot;https://img.ly/products/creative-sdk/&quot;&gt;IMG.LY’s &lt;strong&gt;CreativeEditor SDK (CE.SDK)&lt;/strong&gt;&lt;/a&gt; offers both a user interface for design-time flexibility and a graphics processing engine API for runtime automation.&lt;/p&gt;
&lt;h3 id=&quot;design-once-automate-everywhere&quot;&gt;Design Once, Automate Everywhere&lt;/h3&gt;
&lt;p&gt;With CE.SDK, you create &lt;strong&gt;design templates&lt;/strong&gt; that define placeholders, variables, and lockable design elements. These templates ensure brand consistency while allowing dynamic content generation. Whether you’re building social ads, team cards, or product visuals, the structure remains stable while the content varies. The diagram below illustrates these components using a classical mail merge example. The user creates a design template within the editor and defines certain variables like &lt;code&gt;first_name&lt;/code&gt; and &lt;code&gt;address&lt;/code&gt;, a data source of address data can now be merged with this template using the engine API: &lt;code&gt;variable: setString&lt;/code&gt; to create an individualized postcard for every data point in the set:&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 2106px) 2106px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;2106&quot; height=&quot;1142&quot; src=&quot;https://img.ly/_astro/Screenshot-2025-09-02-at-10.37.28_6SrSA.webp&quot; srcset=&quot;/_astro/Screenshot-2025-09-02-at-10.37.28_Z6guvR.webp 640w, /_astro/Screenshot-2025-09-02-at-10.37.28_Z2uEnjH.webp 750w, /_astro/Screenshot-2025-09-02-at-10.37.28_ZQPQrt.webp 828w, /_astro/Screenshot-2025-09-02-at-10.37.28_Z22tx6s.webp 1080w, /_astro/Screenshot-2025-09-02-at-10.37.28_17hMcX.webp 1280w, /_astro/Screenshot-2025-09-02-at-10.37.28_Z2vSrpD.webp 1668w, /_astro/Screenshot-2025-09-02-at-10.37.28_1TL7CY.webp 2048w, /_astro/Screenshot-2025-09-02-at-10.37.28_6SrSA.webp 2106w&quot;&gt;&lt;/p&gt;
&lt;h2 id=&quot;under-the-hood-cesdk-architecture&quot;&gt;Under the Hood: CE.SDK Architecture&lt;/h2&gt;
&lt;p&gt;In order to fully understand how to implement creative automation workflows with CE.SDK we need a basic mental model of its architecture. CE.SDK is built on two distinct layers, giving you total flexibility and control.&lt;/p&gt;
&lt;p&gt;The &lt;strong&gt;UI Layer&lt;/strong&gt; is the customizable frontend allowing users to design and edit different media types from photos, design compositions such as collages, videos or mixed media.&lt;/p&gt;
&lt;p&gt;The &lt;strong&gt;Engine Layer&lt;/strong&gt; exposes the Creative Engine, CE.SDK’s core rendering and editing engine, through an API. The architecture of CE.SDK is built around a layered approach that separates concerns while ensuring cross-platform consistency:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;**Engine Core (C++)**At the foundation lies the &lt;strong&gt;Creative Engine&lt;/strong&gt;, a high-performance rendering and editing core written in C++. This is where all processing, layouting, and rendering logic happens, guaranteeing precision and fidelity.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;API Layer&lt;/strong&gt;On top of the core sits the &lt;strong&gt;API Layer&lt;/strong&gt;, which exposes the engine’s capabilities in a consistent, high-level interface. This abstraction allows developers to focus on creative logic without dealing directly with low-level rendering details.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cross-Platform Targets&lt;/strong&gt;The same engine and API are &lt;strong&gt;cross-compiled to iOS, Android, Web, and server environments&lt;/strong&gt;, ensuring that rendering results remain identical regardless of platform. This uniformity is critical for automation workflows, where assets may be produced or consumed across multiple devices and services.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Headless Mode&lt;/strong&gt;Beyond powering interactive UIs, the engine can also be run &lt;strong&gt;headlessly&lt;/strong&gt;. In this mode, CE.SDK operates without a frontend, enabling automated scenarios such as bulk rendering of images or videos, dynamic template generation, and data-driven personalization at scale.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This modular architecture ensures that whether you are building a &lt;strong&gt;custom editing UI&lt;/strong&gt; or running &lt;strong&gt;automated creative workflows on the server&lt;/strong&gt;, every output is backed by the same reliable, cross-platform engine.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1926px) 1926px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1926&quot; height=&quot;1880&quot; src=&quot;https://img.ly/_astro/CESDK-architecture_Zef7fB.webp&quot; srcset=&quot;/_astro/CESDK-architecture_Z5SLOS.webp 640w, /_astro/CESDK-architecture_Z19j53A.webp 750w, /_astro/CESDK-architecture_JQ3Ih.webp 828w, /_astro/CESDK-architecture_FyyKh.webp 1080w, /_astro/CESDK-architecture_ZVpPPO.webp 1280w, /_astro/CESDK-architecture_Z5oLSf.webp 1668w, /_astro/CESDK-architecture_Zef7fB.webp 1926w&quot;&gt;&lt;/p&gt;
&lt;h2 id=&quot;how-to-integrate-creative-automation-into-your-workflow&quot;&gt;How to Integrate Creative Automation into Your Workflow&lt;/h2&gt;
&lt;h3 id=&quot;step-1-integrate-cesdk&quot;&gt;Step 1: Integrate CE.SDK&lt;/h3&gt;
&lt;p&gt;Even if you don’t plan to use the CE.SDK UI for end users, you still need to spin it up locally or within internal tooling to create templates. Let’s install CE.SDK and configure it to the advanced UI, in Creator mode.&lt;/p&gt;
&lt;p&gt;CE.SDK by default comes with two UI flavors. The design or default UI config exposes all the essentials for productive design editing by most users, ideal for adapting templates. An advanced UI fine grained editing controls and the ability to define placeholder elements and control what users can change.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The “Creator” role allows setting constraints on template elements, while the “Adopter” role is focused on adapting these elements.&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Creator: Set constraints and manage template settings.&lt;/li&gt;
&lt;li&gt;Adopter: Edit elements within the bounds set by the Creator.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Install CE.SDK following &lt;a href=&quot;https://img.ly/docs/cesdk/js/get-started/overview-e18f40/&quot;&gt;one of our guides&lt;/a&gt; and explore the entire &lt;a href=&quot;https://img.ly/demos/headless-design/web/&quot;&gt;demo here.&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You can find the complete demo here&lt;/p&gt;
&lt;h3 id=&quot;configuring-cesdk-for-template-creation&quot;&gt;Configuring CE.SDK for Template Creation&lt;/h3&gt;
&lt;p&gt;Below is a sample configuration. Notice how we set &lt;code&gt;role: &apos;Creator&apos;&lt;/code&gt; and &lt;code&gt;view: &apos;advanced&apos;&lt;/code&gt; to unlock template authoring features. We also configure an &lt;code&gt;onSave&lt;/code&gt; callback to export complete &lt;strong&gt;template archives&lt;/strong&gt;. These include all necessary resources (images, fonts, stickers, etc.) in a &lt;strong&gt;self-contained package&lt;/strong&gt;. Archives are what you’ll deliver to adopters or use for data-driven automation later.&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; config&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  theme: &lt;/span&gt;&lt;span&gt;&apos;dark&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  role: &lt;/span&gt;&lt;span&gt;&apos;Creator&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  callbacks: {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    onExport: &lt;/span&gt;&lt;span&gt;&apos;download&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    onUpload: &lt;/span&gt;&lt;span&gt;&apos;local&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    onSave&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;async&lt;/span&gt;&lt;span&gt; (&lt;/span&gt;&lt;span&gt;sceneString&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; string&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;=&gt;&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      try&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        // assumes engine is available in scope&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        const&lt;/span&gt;&lt;span&gt; blob&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; await&lt;/span&gt;&lt;span&gt; engine.scene.&lt;/span&gt;&lt;span&gt;saveToArchive&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        const&lt;/span&gt;&lt;span&gt; formData&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; new&lt;/span&gt;&lt;span&gt; FormData&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        formData.&lt;/span&gt;&lt;span&gt;append&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;file&apos;&lt;/span&gt;&lt;span&gt;, blob);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        await&lt;/span&gt;&lt;span&gt; fetch&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;/upload&apos;&lt;/span&gt;&lt;span&gt;, {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;          method: &lt;/span&gt;&lt;span&gt;&apos;POST&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;          body: formData,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        });&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      } &lt;/span&gt;&lt;span&gt;catch&lt;/span&gt;&lt;span&gt; (error) {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        console.&lt;/span&gt;&lt;span&gt;error&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;Save failed&apos;&lt;/span&gt;&lt;span&gt;, error);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      }&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  ui: {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    elements: {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      view: &lt;/span&gt;&lt;span&gt;&apos;advanced&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      panels: {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        inspector: {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;          show: &lt;/span&gt;&lt;span&gt;true&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;          position: &lt;/span&gt;&lt;span&gt;&apos;right&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      dock: {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        iconSize: &lt;/span&gt;&lt;span&gt;&apos;normal&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        hideLabels: &lt;/span&gt;&lt;span&gt;true&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      navigation: {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        action: {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;          export: {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;            show: &lt;/span&gt;&lt;span&gt;true&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;            format: [&lt;/span&gt;&lt;span&gt;&apos;image/png&apos;&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;&apos;application/pdf&apos;&lt;/span&gt;&lt;span&gt;],&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;          },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;};&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;step-2-build-smart-templates&quot;&gt;Step 2: Build Smart Templates&lt;/h3&gt;
&lt;p&gt;With CE.SDK running in &lt;strong&gt;Creator mode&lt;/strong&gt;, we can now define &lt;strong&gt;editable templates&lt;/strong&gt;. These templates can include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Named blocks&lt;/strong&gt; → Easy-to-reference design regions (e.g., &lt;code&gt;PodcastCover&lt;/code&gt;, &lt;code&gt;PodcastBadge&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Variables&lt;/strong&gt; → Text placeholders enclosed in double curly braces (e.g., &lt;code&gt;{{Message}}&lt;/code&gt;, &lt;code&gt;{{PodcastName}}&lt;/code&gt;).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For example, here’s a &lt;strong&gt;podcast cover template&lt;/strong&gt; with two text variables and named blocks. The variables make it possible to inject data programmatically later.&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Podcast template:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;- Variables: {{Message}}, {{PodcastName}}&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;- Named blocks: PodcastCover, PodcastBadge&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Since we’re in Creator mode, we can also enforce &lt;strong&gt;brand constraints,&lt;/strong&gt; for instance, locking background colors, logos, or fonts so that adopters (or automated processes) don’t accidentally override brand-critical assets.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 2000px) 2000px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;2000&quot; height=&quot;1053&quot; src=&quot;https://img.ly/_astro/Screenshot-2025-09-04-at-13.01.25_ZJuH0A.webp&quot; srcset=&quot;/_astro/Screenshot-2025-09-04-at-13.01.25_Z2v8kLa.webp 640w, /_astro/Screenshot-2025-09-04-at-13.01.25_1uCzrP.webp 750w, /_astro/Screenshot-2025-09-04-at-13.01.25_Z16Kj3X.webp 828w, /_astro/Screenshot-2025-09-04-at-13.01.25_1HahGR.webp 1080w, /_astro/Screenshot-2025-09-04-at-13.01.25_IOBh7.webp 1280w, /_astro/Screenshot-2025-09-04-at-13.01.25_Z1hGL7j.webp 1668w, /_astro/Screenshot-2025-09-04-at-13.01.25_ZJuH0A.webp 2000w&quot;&gt;&lt;/p&gt;
&lt;h3 id=&quot;step-3-automate-asset-creation&quot;&gt;Step 3: Automate Asset Creation&lt;/h3&gt;
&lt;p&gt;Once you have a template, the next step is automation. Using the &lt;strong&gt;headless Creative Engine&lt;/strong&gt;, you can:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Populate placeholders with &lt;strong&gt;external data&lt;/strong&gt; (e.g., from APIs).&lt;/li&gt;
&lt;li&gt;Swap images, update text variables, or adjust styling dynamically.&lt;/li&gt;
&lt;li&gt;Export the results at scale into formats like PNG, PDF, or even video.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;example-dynamic-podcast-covers&quot;&gt;Example: Dynamic Podcast Covers&lt;/h3&gt;
&lt;p&gt;In this demo, we connect to the &lt;strong&gt;Apple iTunes API&lt;/strong&gt; to fetch podcast data. Whenever a podcast is selected (or user input changes), our &lt;code&gt;fillTemplate&lt;/code&gt; function updates the CE.SDK scene.&lt;/p&gt;
&lt;p&gt;&lt;video src=&quot;https://blog.img.ly/2025/09/podcast-data-source.mov&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;p&gt;Key concepts:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Blocks API&lt;/strong&gt; → Query template blocks by name (&lt;code&gt;engine.block.findByName(&apos;PodcastBadge&apos;)&lt;/code&gt;) and replace their fills.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fills&lt;/strong&gt; → Blocks can have images, solid colors, gradients, or videos. For podcast artwork, we set &lt;code&gt;imageFileURI&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Variables API&lt;/strong&gt; → Set text variables directly (&lt;code&gt;engine.variable.setString(&apos;Message&apos;, &apos;Hello World&apos;)&lt;/code&gt;).&lt;/li&gt;
&lt;/ul&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt; const&lt;/span&gt;&lt;span&gt; fillTemplate&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; (&lt;/span&gt;&lt;span&gt;engine&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; CreativeEngine&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;page&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; number&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;=&gt;&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    let&lt;/span&gt;&lt;span&gt; { r, g, b } &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; backgroundColorRGBA;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    engine.block.&lt;/span&gt;&lt;span&gt;setColor&lt;/span&gt;&lt;span&gt;(page, &lt;/span&gt;&lt;span&gt;&apos;fill/solid/color&apos;&lt;/span&gt;&lt;span&gt;, { r, g, b, a: &lt;/span&gt;&lt;span&gt;1&lt;/span&gt;&lt;span&gt; });&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    if&lt;/span&gt;&lt;span&gt; (podcast) {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      const&lt;/span&gt;&lt;span&gt; photoBlocks&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;findByName&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;PodcastCover&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      photoBlocks.&lt;/span&gt;&lt;span&gt;forEach&lt;/span&gt;&lt;span&gt;((&lt;/span&gt;&lt;span&gt;photoBlock&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; number&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;=&gt;&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        const&lt;/span&gt;&lt;span&gt; photoFillBlock&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;getFill&lt;/span&gt;&lt;span&gt;(photoBlock);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        engine.block.&lt;/span&gt;&lt;span&gt;setString&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;          photoFillBlock,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;          &apos;fill/image/imageFileURI&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;          podcast.artworkUrl600&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        );&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      });&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    }&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    const&lt;/span&gt;&lt;span&gt; colorTheme&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; getThemeColorFromBackgroundColor&lt;/span&gt;&lt;span&gt;(backgroundColorRGBA);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    const&lt;/span&gt;&lt;span&gt; [&lt;/span&gt;&lt;span&gt;badgeBlock&lt;/span&gt;&lt;span&gt;] &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;findByName&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;PodcastBadge&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    engine.block.&lt;/span&gt;&lt;span&gt;setString&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      engine.block.&lt;/span&gt;&lt;span&gt;getFill&lt;/span&gt;&lt;span&gt;(badgeBlock),&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      &apos;fill/image/imageFileURI&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      caseAssetPath&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        `/podcast-badge-${&lt;/span&gt;&lt;span&gt;colorTheme&lt;/span&gt;&lt;span&gt; ===&lt;/span&gt;&lt;span&gt; &apos;light&apos;&lt;/span&gt;&lt;span&gt; ?&lt;/span&gt;&lt;span&gt; &apos;black&apos;&lt;/span&gt;&lt;span&gt; :&lt;/span&gt;&lt;span&gt; &apos;white&apos;}.png`&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      )&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    );&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    // set text variables&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    engine.variable.&lt;/span&gt;&lt;span&gt;setString&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;Message&apos;&lt;/span&gt;&lt;span&gt;, debouncedMessage &lt;/span&gt;&lt;span&gt;||&lt;/span&gt;&lt;span&gt; &apos;&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    engine.variable.&lt;/span&gt;&lt;span&gt;setString&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      &apos;PodcastName&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      podcast &lt;/span&gt;&lt;span&gt;?&lt;/span&gt;&lt;span&gt; podcast.collectionName &lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; &apos;&apos;&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      // set text colors based on background theme&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      const&lt;/span&gt;&lt;span&gt; [&lt;/span&gt;&lt;span&gt;messageBlock&lt;/span&gt;&lt;span&gt;] &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;findByName&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;Message &amp;#x26; Name&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      const&lt;/span&gt;&lt;span&gt; rgb&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      colorTheme &lt;/span&gt;&lt;span&gt;===&lt;/span&gt;&lt;span&gt; &apos;dark&apos;&lt;/span&gt;&lt;span&gt; ?&lt;/span&gt;&lt;span&gt; { r: &lt;/span&gt;&lt;span&gt;1&lt;/span&gt;&lt;span&gt;, g: &lt;/span&gt;&lt;span&gt;1&lt;/span&gt;&lt;span&gt;, b: &lt;/span&gt;&lt;span&gt;1&lt;/span&gt;&lt;span&gt; } &lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; { r: &lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;, g: &lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;, b: &lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt; };&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      engine.block.&lt;/span&gt;&lt;span&gt;setTextColor&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      messageBlock,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      { &lt;/span&gt;&lt;span&gt;...&lt;/span&gt;&lt;span&gt;rgb, a: &lt;/span&gt;&lt;span&gt;0.75&lt;/span&gt;&lt;span&gt; },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      0&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      &quot;{{Message}}&quot;&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;length&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    );&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    engine.block.&lt;/span&gt;&lt;span&gt;setTextColor&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      messageBlock,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      { &lt;/span&gt;&lt;span&gt;...&lt;/span&gt;&lt;span&gt;rgb, a: &lt;/span&gt;&lt;span&gt;1.0&lt;/span&gt;&lt;span&gt; },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;       &quot;{{Message}}&quot;&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;length&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    );&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  };&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;At this point, you’ve created a &lt;strong&gt;data-driven creative automation pipeline&lt;/strong&gt;:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Templates authored with &lt;strong&gt;placeholders and constraints&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Dynamic population of text and image fields via external data.&lt;/li&gt;
&lt;li&gt;Export to &lt;strong&gt;high-quality assets&lt;/strong&gt; at scale.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This workflow is the foundation for &lt;strong&gt;personalized marketing campaigns, automated content production, and AI-driven design pipelines&lt;/strong&gt; with CE.SDK.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Creative Automation isn’t just about speed it’s about &lt;strong&gt;scalability, consistency, and customization&lt;/strong&gt;. &lt;a href=&quot;https://img.ly/&quot;&gt;IMG.LY&lt;/a&gt;’s CE.SDK gives you the tools to build powerful, flexible automation workflows tailored to your brand and users.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/docs/cesdk/&quot;&gt;Start you free trial today&lt;/a&gt; and integrate content automation into your application.&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="https://blog.img.ly/2025/09/Creative-Automation-Cover.png" medium="image"/><category>Creative Automation</category></item><item><title>The Top 7 Video Editing SDKs in 2026</title><link>https://img.ly/blog/top-7-video-editing-sdks-in-2025/</link><guid isPermaLink="true">https://img.ly/blog/top-7-video-editing-sdks-in-2025/</guid><description>In this guide, we break down the seven leading video editing SDK solutions: IMG.LY, Meishe, Banuba, Shotstack, BytePlus, Rendley, and Picsart, covering their features, platforms, use cases, and trade-offs. </description><pubDate>Tue, 02 Sep 2025 11:19:08 GMT</pubDate><content:encoded>&lt;p&gt;With dozens of video editing SDKs on the market, choosing the right one can feel like a daunting task.&lt;/p&gt;
&lt;p&gt;To help you cut through the noise, we’ve put together a breakdown of the &lt;strong&gt;7 best video editing SDKs in 2026:&lt;/strong&gt; IMG.LY, Meishe SDK, Banuba, Shotstack, BytePlus Effects, Rendley, and Picsart. We’ll cover:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Key features and capabilities&lt;/li&gt;
&lt;li&gt;Supported platforms and customization options&lt;/li&gt;
&lt;li&gt;Ideal use cases and trade-offs&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This guide is unbiased, but we’ll also highlight how IMG.LY compares and where it stands out as the only solution that combines a polished editor with scalable automation in a single package.&lt;/p&gt;
&lt;h2 id=&quot;1-imgly-creativeeditor-sdk-cesdk&quot;&gt;&lt;strong&gt;1. IMG.LY CreativeEditor SDK (CE.SDK)&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Choosing IMG.LY means selecting a platform that offers both: a Canva-grade editor for your users and a powerful Engine API for automation. That dual focus is what sets CE.SDK apart.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Features &amp;#x26; capabilities:&lt;/strong&gt; CE.SDK works best when you see it as a bridge between what end users expect and what your product team needs behind the scenes. On the surface, users get familiar tools including multi‑track editing, ready‑made templates, and a growing set of AI features like text‑to‑image, text‑to‑video, and AI voiceovers. Behind the scenes, developers can extend and customise with a plugin system so the editor adapts as your product evolves.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Platforms &amp;#x26; customization:&lt;/strong&gt; This flexibility carries over to platforms. CE.SDK works across Web, iOS, Android, React Native, Flutter, Electron, and Node.js. It’s also fully white‑label, which means your customers see a seamless extension of your brand, not a third‑party tool.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Implementation &amp;#x26; use cases:&lt;/strong&gt; Getting started with our tool doesn’t require a huge lift. Teams can spin up a proof of concept quickly, then scale into enterprise deployment when ready. That’s why you’ll see CE.SDK inside SaaS platforms, e‑commerce personalization tools, DAM systems, and even social networks.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Future‑proofing, support &amp;#x26; pricing&lt;/strong&gt;: What reassures teams is the long‑term roadmap: regular feature releases, AI integrations, and direct support from solution engineers. Enterprise SLAs and SSO come standard. Pricing depends on the usage, which makes it easier to scale without committing to rigid licensing.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Who it’s for:&lt;/strong&gt; CE.SDK is valuable for SaaS, MarTech, e-commerce, and DAM platforms. These teams need to give users a simple, branded editing experience while also managing complex automation in the background. Instead of stitching together multiple tools or sacrificing one capability for another, CE.SDK provides an all-in-one foundation: polished editing, scalable automation, enterprise-grade support, and ongoing AI-driven innovation.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;key-differentiator&quot;&gt;&lt;strong&gt;Key differentiator&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;CE.SDK doesn’t force you to choose between usability and power. On one side, it delivers a polished, user-facing editor that feels familiar to end users. On the other hand, it provides a powerful automation engine that scales with your product. Most SDKs lean heavily in one direction. They either focus on end-user editing or specialize in backend automation. CE.SDK is different because it combines both in a single platform.&lt;/p&gt;
&lt;h2 id=&quot;2-meishe-sdk-meicam&quot;&gt;&lt;strong&gt;2. Meishe SDK (Meicam)&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Meishe SDK, also known as Meicam, is a mature Chinese SDK that has carved out a space in many professional apps across Asia. Think of it as a toolkit built for industries that need stability and advanced editing power, though it comes with a few trade‑offs that teams should weigh carefully.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Features &amp;#x26; capabilities:&lt;/strong&gt; Meishe offers multi‑track timelines, real‑time effects, and plugin support across mobile, web, and PC. Industries like automotive, broadcasting, and education have already used Meishe successfully, showing how it handles complex and demanding video workflows.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Platforms &amp;#x26; customization:&lt;/strong&gt; It supports full UI customization, white‑labeling, and plugins. And while customization depth is solid, integration tends to be heavier compared to other tools. Documentation can feel dense, which sometimes slows teams down as they work toward launch.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Implementation &amp;#x26; use cases:&lt;/strong&gt; Meishe SDK is beginning to layer in AI integrations like intelligent editing and motion tracking. For large organizations in Asia with strong internal dev resources, these capabilities can be appealing.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Future‑proofing, support &amp;#x26; pricing:&lt;/strong&gt; Meishe continues to expand with new AI features like facial recognition and virtual anchors. It also offers custom enterprise support agreements. However, teams often struggle with the integration process, and the developer experience feels less polished. Setting up complex projects can take longer, and the documentation isn’t as approachable, which leaves some developers hesitant to adopt it widely. Pricing typically follows an enterprise licensing model, which can be less flexible for fast‑scaling SaaS companies.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Who it’s for:&lt;/strong&gt; Meishe SDK is best suited for large-scale apps in Asia and companies with internal dev resources.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;how-it-compares-to-imgly&quot;&gt;&lt;strong&gt;How it compares to IMG.LY&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Meishe delivers solid editing capabilities and is a proven choice in professional and regional contexts. But the trade‑offs become clear when you look at integration and support. It can feel heavy to implement, and documentation isn’t as approachable.&lt;/p&gt;
&lt;p&gt;IMG.LY pulls ahead here with stronger web support, compatibility with creative file types like PSD, AI, and INDD, and a faster onboarding process that makes it easier for teams to move from idea to production.&lt;/p&gt;
&lt;p&gt;Learn more about how &lt;a href=&quot;https://img.ly/meishesdk-alternative/&quot;&gt;IMG.LY compares to Meishe SDK here&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;3-banuba&quot;&gt;&lt;strong&gt;3. Banuba&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Banuba is a mobile‑first SDK with a strong focus on augmented reality (AR). It has built its reputation on powering social and UGC apps with effects that mimic the TikTok experience. If your product relies on immersive filters, beauty effects, or virtual try‑on, Banuba often appears at the top of the evaluation list.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Features &amp;#x26; capabilities:&lt;/strong&gt; The SDK provides a TikTok‑like editor with AR masks, beauty filters, face tracking, background removal, and even virtual try‑on features. These tools are designed to keep users engaged with interactive and shareable content.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Platforms &amp;#x26; customization:&lt;/strong&gt; Banuba supports iOS, Android, Flutter, and React Native. It offers modular components with a preset UI, which makes implementation faster for mobile. That said, customization options are limited when compared to solutions designed for enterprise-level flexibility.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Implementation &amp;#x26; use cases:&lt;/strong&gt; Integration is relatively quick for mobile teams, and the SDK has proven popular among social media platforms, beauty apps, and e‑commerce companies offering AR try‑on. However, support for web or large‑scale enterprise environments is limited, which narrows its scope.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Future‑proofing, support &amp;#x26; pricing:&lt;/strong&gt; The company continues to expand its AR and AI features, focusing on personalization and beauty‑driven effects. Banuba operates with a smaller team of about 70 employees, which can affect the pace of enterprise‑grade support. Licensing costs are typically high and not publicly disclosed, which may be a barrier for smaller companies.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Who it’s for:&lt;/strong&gt; Banuba is best suited for mobile apps that need advanced AR features like beauty filters, face tracking, or virtual try‑on to keep users engaged.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;how-it-compares-to-imgly-1&quot;&gt;&lt;strong&gt;How it compares to IMG.LY&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Banuba leads when it comes to AR‑driven capabilities, making it a natural fit for beauty and social platforms. But when teams look beyond filters to enterprise readiness, automation, and cross‑platform support, other options such as IMG.LY offer broader possibilities.&lt;/p&gt;
&lt;p&gt;Learn more about how &lt;a href=&quot;https://img.ly/banuba-alternative/&quot;&gt;IMG.LY compares to Banuba here&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;4-shotstack&quot;&gt;&lt;strong&gt;4. Shotstack&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Shotstack focuses squarely on automation. It’s a cloud‑based video editing API designed for bulk rendering rather than end‑user editing. Teams often evaluate it when they need to generate thousands of videos quickly and at scale.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Features &amp;#x26; capabilities:&lt;/strong&gt; Shotstack runs on a JSON‑driven API that lets developers programmatically create and assemble video content. With this approach, teams can handle automated ad generation, personalization, or large‑scale marketing campaigns. While powerful for automation, it doesn’t include a native editor, which means you must build a custom UI if you want end‑user editing.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Platforms &amp;#x26; customization:&lt;/strong&gt; Shotstack is built as a REST API in the cloud, so it fits naturally into modern development stacks. It allows teams to design workflows programmatically, though the absence of prebuilt UI components makes it harder to use the tool right away as compared to other SDKs.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Implementation &amp;#x26; use cases:&lt;/strong&gt; Developers appreciate the smooth onboarding experience and clear API structure. Typical use cases include bulk rendering of ads, dynamic personalization for campaigns, and other scenarios where scaling output matters more than in‑app editing.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Future‑proofing, support &amp;#x26; pricing:&lt;/strong&gt; Shotstack continues to invest in automation and AI‑powered generation, offering tools that personalize video creation at scale. The vendor provides commercial SLAs, but support resources remain limited given the team size. Pricing follows a pay‑as‑you‑go model ($0.30 per minute) or subscription plans, which can work well for predictable, high‑volume usage.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Who it’s for:&lt;/strong&gt; Shotstack is best suited for teams that need scalable rendering pipelines and automation rather than in‑app editing features.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;how-it-compares-to-imgly-2&quot;&gt;&lt;strong&gt;How it compares to IMG.LY&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Shotstack gives you powerful rendering automation and programmatic workflows. IMG.LY takes it further by combining those automation strengths with a complete, user‑facing editor, so you can address both your development needs and your end‑users’ experiences in one solution.&lt;/p&gt;
&lt;p&gt;Get a detailed &lt;a href=&quot;https://img.ly/shotstack-alternative/&quot;&gt;comparison of IMG.LY CreativeEditor SDK and Shotstack&lt;/a&gt; here.&lt;/p&gt;
&lt;h2 id=&quot;5-byteplus-effects&quot;&gt;&lt;strong&gt;5. BytePlus Effects&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;BytePlus Effects is modelled after TikTok and CapCut. It’s designed to give app developers the same kind of short‑form, social‑first experience that has made those platforms successful. As a result, it appeals to teams that want to replicate TikTok‑like engagement inside their own apps.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Features &amp;#x26; capabilities:&lt;/strong&gt; The SDK offers more than 80,000 effects and filters, including AR stickers and beautification tools. These features aim to keep users engaged by giving them endless ways to personalize and enhance their content.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Platforms &amp;#x26; customization:&lt;/strong&gt; BytePlus integrates with iOS, Android, React Native, and Flutter, making it a strong option for mobile‑first development. You can customise the UI to match your brand, but customization beyond theming is limited, which can be a constraint if you need deeper control.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Implementation &amp;#x26; use cases:&lt;/strong&gt; Since it’s optimized for mobile integration, you can embed it relatively quickly into short‑form consumer video apps. Common use cases include social platforms, beauty apps, and any consumer‑facing product that thrives on AR‑enhanced video.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Future‑proofing, support &amp;#x26; pricing:&lt;/strong&gt; With the backing of TikTok’s ecosystem, BytePlus is evolving rapidly, adding new AI‑driven AR and beauty effects frequently. However, commercial support feels less transparent, and licensing comes through enterprise negotiations with annual fees. This structure can work for large consumer platforms, but may be harder for smaller teams.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Who it’s for:&lt;/strong&gt; BytePlus is best suited for apps that want to recreate TikTok‑style user experiences and rely heavily on prebuilt AR effects to keep users engaged.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;how-it-compares-to-imgly-3&quot;&gt;&lt;strong&gt;How it compares to IMG.LY&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;BytePlus takes a step ahead when it comes to prebuilt effects and AR‑driven experiences. But if you need more flexibility in customization, automation capabilities, or enterprise‑grade workflows, IMG.LY offers a broader solution.&lt;/p&gt;
&lt;p&gt;For a detailed comparison, take a look at our guide on &lt;a href=&quot;https://img.ly/byteplus-alternative/&quot;&gt;IMG.LY vs BytePlus Video Editor.&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;6-rendley&quot;&gt;&lt;strong&gt;6. Rendley&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Rendley is a newer player in the SDK space and has gained attention because of its browser‑only approach. It’s JavaScript‑based and runs entirely in the browser, which means you don’t need a server to get started. This makes it attractive for teams that want a lightweight, client‑only setup without managing infrastructure.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Features &amp;#x26; capabilities:&lt;/strong&gt; The SDK supports multi‑track editing, keyframes, and even integrates with After Effects (AE). These features give frontend developers enough flexibility to build interactive editing experiences directly in the browser.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Platforms &amp;#x26; customization:&lt;/strong&gt; Rendley runs entirely in the browser, which makes it a natural fit for lightweight web apps and UGC platforms. Developers get flexibility through an open, in‑browser approach. However, without enterprise‑level frameworks, scaling to larger organizations becomes more difficult.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Implementation &amp;#x26; use cases:&lt;/strong&gt; Setting up Rendley is straightforward for frontend teams, which makes it appealing for lightweight UGC applications or smaller web projects. That ease of entry, however, comes with limits: AI integrations aren’t a core focus, so its potential for more advanced future applications remains restricted.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Future‑proofing, support &amp;#x26; pricing:&lt;/strong&gt; Rendley is still early‑stage and hasn’t yet proven itself at enterprise scale. Support is provided by a small vendor team, which can make long‑term stability less certain. Pricing details are not publicly available, leaving prospective users to negotiate directly.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Who it’s for:&lt;/strong&gt; Rendley fits developers who want a browser‑native, client‑only editing tool without the need for server infrastructure.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;how-it-compares-to-imgly-4&quot;&gt;&lt;strong&gt;How it compares to IMG.LY&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Rendley works well as a lightweight, browser‑native solution. But if you need enterprise‑grade robustness, automation, or dedicated support, IMG.LY offers more comprehensive coverage.&lt;/p&gt;
&lt;p&gt;Learn more about how &lt;a href=&quot;https://img.ly/rendley-alternative/&quot;&gt;IMG.LY is a great alternative for Rendley&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;7-picsart-video-sdk&quot;&gt;&lt;strong&gt;7. Picsart Video SDK&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Picsart Video SDK is an embeddable video editor SDK that brings Picsart’s creative editing tools into third-party apps or websites.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Features &amp;#x26; capabilities:&lt;/strong&gt; This SDK brings the company’s well‑known creative editing tools into third‑party products. Out of the box, you get features like trimming, multi‑scene editing, transitions, dynamic audio, subtitles, and text overlays, enough to cover the basics of video creation. AI‑driven tools extend these capabilities with smarter image and video editing options.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Platforms &amp;#x26; customization:&lt;/strong&gt; The SDK is React Native‑based but optimized for web embedding, making it appealing for web‑first platforms. You can configure the UI, apply theming, white‑label the experience, toggle features on or off, and integrate your own asset library so the editor matches your product’s design and workflow.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Implementation &amp;#x26; use cases:&lt;/strong&gt; One of Picsart’s strengths is ease of setup: installation is fast, with a React editor installation process that reduces complexity, documentation is detailed, and the editor is designed to embed with minimal lift. This makes it a good choice for advertising or marketing platforms, edu‑tech products, and media platforms that want to give users creative tools quickly.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Future‑proofing, support &amp;#x26; pricing:&lt;/strong&gt; The SDK is backed by Picsart’s large ecosystem and its AI research arm (PAIR), so teams can expect regular updates, new features, and scalable performance. Support comes through developer documentation, a support team, and enterprise collaboration options. Pricing isn’t public, but it’s likely offered as enterprise or usage‑based licensing.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Who it’s for:&lt;/strong&gt; Picsart’s Video SDK is best for teams that need a web‑based video editor to power product features or educational video production without investing heavily in custom development.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;how-it-compares-to-imgly-5&quot;&gt;&lt;strong&gt;How it compares to IMG.LY&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Picsart offers a fast, embeddable editor with AI effects and a strong ecosystem. However, IMG.LY provides deeper customization, more automation capabilities, cross‑platform support, and enterprise‑grade integrations.&lt;/p&gt;
&lt;h2 id=&quot;overview-table-use-cases-vs-solutions&quot;&gt;&lt;strong&gt;Overview Table (Use Cases vs. Solutions)&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Here is how the seven options line up on the four questions teams ask most often during an evaluation.&lt;/p&gt;





























































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;SDK&lt;/th&gt;&lt;th&gt;Best for&lt;/th&gt;&lt;th&gt;Platforms&lt;/th&gt;&lt;th&gt;Customization depth&lt;/th&gt;&lt;th&gt;Pricing model&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;IMG.LY CE.SDK&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;SaaS, MarTech, e-commerce and DAM platforms that need a user-facing editor and backend automation&lt;/td&gt;&lt;td&gt;Web, iOS, Android, React Native, Flutter, Electron, Node.js&lt;/td&gt;&lt;td&gt;Fully white-label, extensible through a plugin system&lt;/td&gt;&lt;td&gt;Usage-based&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Meishe (Meicam)&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Large-scale apps in Asia with strong internal dev resources&lt;/td&gt;&lt;td&gt;Mobile, web, PC&lt;/td&gt;&lt;td&gt;Full UI customization, white-labeling and plugins, though integration is heavier&lt;/td&gt;&lt;td&gt;Enterprise licensing&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Banuba&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Mobile apps built around AR: beauty filters, face tracking, virtual try-on&lt;/td&gt;&lt;td&gt;iOS, Android, Flutter, React Native&lt;/td&gt;&lt;td&gt;Modular components with a preset UI, limited beyond that&lt;/td&gt;&lt;td&gt;Not public, typically high&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Shotstack&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Bulk rendering and automation pipelines rather than in-app editing&lt;/td&gt;&lt;td&gt;Cloud REST API&lt;/td&gt;&lt;td&gt;No prebuilt UI components, you build the editor yourself&lt;/td&gt;&lt;td&gt;Pay-as-you-go at $0.30 per minute, or subscription&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;BytePlus Effects&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Consumer apps recreating TikTok-style experiences from prebuilt AR effects&lt;/td&gt;&lt;td&gt;iOS, Android, React Native, Flutter&lt;/td&gt;&lt;td&gt;Theming to match your brand, limited past that&lt;/td&gt;&lt;td&gt;Enterprise negotiation with annual fees&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Rendley&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Browser-native, client-only editing with no server infrastructure&lt;/td&gt;&lt;td&gt;Browser (JavaScript)&lt;/td&gt;&lt;td&gt;Open in-browser approach, no enterprise-level frameworks&lt;/td&gt;&lt;td&gt;Not public&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Picsart Video SDK&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Web-first product, edu-tech and media teams that want a fast embed&lt;/td&gt;&lt;td&gt;React Native based, optimized for web embedding&lt;/td&gt;&lt;td&gt;Configurable UI, theming, white-label, feature toggles, own asset library&lt;/td&gt;&lt;td&gt;Not public, likely enterprise or usage-based&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;&lt;strong&gt;Conclusion&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;When you line these SDKs up side by side, it’s clear that each one has its own niche strengths and limitations. Here’s a quick overview:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Banuba / BytePlus:&lt;/strong&gt; These are the strongest options for AR and social apps, with beauty filters, effects, and TikTok-like workflows.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Meishe:&lt;/strong&gt; A professional, Asia-centric SDK with heavy native capabilities, but a steeper implementation curve.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Shotstack / Rendley:&lt;/strong&gt; Both shine in automation pipelines or lightweight browser-based setups, but they don’t offer polished editor UIs.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Picsart:&lt;/strong&gt; A strong web-first, prebuilt editor with AI effects that’s simple to embed, though limited in customization and automation.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;where-does-that-leave-you&quot;&gt;Where does that leave you?&lt;/h3&gt;
&lt;p&gt;If you’re evaluating SDKs, you need to balance two priorities: delivering a great editor UX for your users and scaling automation behind the scenes. Only IMG.LY covers both ends: offering a polished, enterprise-ready editor and a scalable automation engine.&lt;/p&gt;
&lt;p&gt;That makes it the go-to option for teams planning not only their current projects but also long-term growth. With IMG.LY, you can give users a polished editing experience today while building the automation backbone that will let your product scale tomorrow&lt;/p&gt;
&lt;p&gt;If you’d like to see how IMG.LY can fit into your platform, &lt;a href=&quot;https://img.ly/forms/contact-sales/&quot;&gt;get in touch with our experts today&lt;/a&gt; for a deeper look and tailored guidance.&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="https://blog.img.ly/2025/09/top-video-editor-sdks-in-2025--1-.png" medium="image"/><category>Video Editor</category><category>Insights</category></item><item><title>What is Visual Prompting?</title><link>https://img.ly/blog/what-is-visual-prompting/</link><guid isPermaLink="true">https://img.ly/blog/what-is-visual-prompting/</guid><description>Visual Prompting is a new way to guide AI using visual input instead of just text. By composing layouts with images, annotations, and design cues directly on the canvas, creators can prompt AI more intuitively. Learn how IMG.LY’s CE.SDK brings this paradigm to life.</description><pubDate>Tue, 29 Jul 2025 13:06:11 GMT</pubDate><content:encoded>&lt;h2 id=&quot;a-new-paradigm-for-creative-ai-built-by-imgly&quot;&gt;A New Paradigm for Creative AI, Built by IMG.LY&lt;/h2&gt;
&lt;p&gt;To say it’s trite to refer to the impact of AI in this or that domain as disruptive or groundbreaking would be an understatement. Yet, few areas have been as profoundly affected as the creative process. With just a text prompt, anyone can produce stunning images, remix visual styles, and explore design possibilities at a scale and speed never seen before. AI has inserted itself so quickly into this process that its gone from curious novelty to an essential part of the creator toolchain.&lt;/p&gt;
&lt;p&gt;The more serious adoption we see, however, the more key limitations of today’s AI tooling come into focus: &lt;strong&gt;the prompt itself&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Text alone, for all its expressive power, struggles to capture the essence of visual intent. Most creative work doesn’t begin with a sentence it begins with a sketch, a layout, a mood board, or an arrangement of elements. Visual ideas are shared by pointing, placing, showing.&lt;/p&gt;
&lt;p&gt;At &lt;strong&gt;IMG.LY&lt;/strong&gt;, we have begun to think about better ways to direct AI for visual generation, the term we use is Visual Prompting.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Visual Prompting: the practice of composing a visual scene or layout as input for a generative model.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Instead of describing what you want with paragraphs of text, you show it directly using a canvas of images, text, spatial cues, and annotations. This visual composition then becomes the prompt for the AI to generate new content in return. It’s a more natural, intuitive, and powerful way to collaborate with AI, especially when integrated directly into the creative process.&lt;/p&gt;
&lt;h2 id=&quot;problem-the-chat-disconnect&quot;&gt;Problem: the Chat Disconnect&lt;/h2&gt;
&lt;p&gt;The current generation of AI tools has largely been shaped by language-first interfaces. Whether it’s ChatGPT for writing or Midjourney for image generation, the assumption is the same: the user will type a descriptive prompt, and the AI will generate a result based on it.&lt;/p&gt;
&lt;p&gt;But when it comes to &lt;strong&gt;design&lt;/strong&gt;, this workflow quickly runs into friction. Visual ideas are inherently spatial and non-linear. Trying to express layout, balance, mood, or specific spatial relationships through text can feel like trying to describe a painting over the phone. It’s possible but unnecessarily cumbersome.&lt;/p&gt;
&lt;p&gt;A designer might want to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Indicate that a certain area in the image should be blue.&lt;/li&gt;
&lt;li&gt;Replace a background with a texture sample.&lt;/li&gt;
&lt;li&gt;Position a character precisely in a composition.&lt;/li&gt;
&lt;li&gt;Annotate which parts of a scene to preserve or modify.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;All of these are difficult to express fluently in text. But they’re &lt;strong&gt;effortless&lt;/strong&gt; in a visual interface. The truth is: &lt;strong&gt;an image is worth more than a thousand words when prompting an image.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id=&quot;what-is-visual-prompting&quot;&gt;What Is Visual Prompting?&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Visual Prompting&lt;/strong&gt; is a multimodal approach to generative AI, where the &lt;strong&gt;input to the model is not just text, but a full visual composition&lt;/strong&gt;: images, text, annotations, and layout.&lt;/p&gt;
&lt;p&gt;Rather than prompting AI in isolation, the user builds their intent on a canvas. This might include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Reference images that communicate mood or style.&lt;/li&gt;
&lt;li&gt;Text blocks indicating desired copy or instructions.&lt;/li&gt;
&lt;li&gt;Annotations pointing to specific areas with notes like “make this glow” or “replace this object.”&lt;/li&gt;
&lt;li&gt;Spatial composition: where elements are arranged meaningfully to convey intent.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The visual prompt is then interpreted by a multimodal model such as OpenAI’s &lt;code&gt;gpt-image-1&lt;/code&gt; to generate new visual content that reflects not only the textual description, but also the visual context.&lt;/p&gt;
&lt;h2 id=&quot;how-visual-prompting-works-in-cesdk&quot;&gt;How Visual Prompting Works in CE.SDK&lt;/h2&gt;
&lt;p&gt;About time for an example. As part of our recent AI released &lt;a href=&quot;https://img.ly/demos/ai-editor/web/&quot;&gt;we demoed how to use OpenAIs &lt;code&gt;gpt-image-1&lt;/code&gt; model&lt;/a&gt; to build visual prompting into &lt;strong&gt;CreativeEditor SDK (CE.SDK)&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Here’s what the process looks like inside CE.SDK:&lt;/p&gt;
&lt;p&gt;&lt;video src=&quot;https://storage.googleapis.com/imgly-static-assets/static/blog/videos/imgly-ai/visualprompt_05.mp4&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Compose Visually:&lt;/strong&gt; The user creates a layout with reference content, uploaded images, icons, color schemes, design elements, placeholder text, and annotations. This composition represents the “prompt” in visual form.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Add AI Layers:&lt;/strong&gt; With a single click, the user can trigger image generation using CE.SDK’s built-in AI plugin. The plugin sends the visual context (alongside any optional text input) to a multimodal model capable of interpreting both.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Refine and Iterate:&lt;/strong&gt; Users can adjust the layout, reposition elements, change annotations, or layer in new references, then prompt again. Because the canvas is interactive and editable, the feedback loop is tight.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Build Up Complexity:&lt;/strong&gt; Over time, users can layer generated images with manually designed components or other generated outputs, creating rich compositions that blend AI creativity with human direction.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This workflow turns the traditional prompt/response cycle into a &lt;strong&gt;conversation between the designer and the model&lt;/strong&gt;, with the canvas acting as the shared language.&lt;/p&gt;
&lt;h2 id=&quot;who-is-visual-prompting-for&quot;&gt;Who Is Visual Prompting For?&lt;/h2&gt;
&lt;p&gt;The use cases for Visual Prompting extend across industries:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Creative teams&lt;/strong&gt; can go from reference to generation in seconds, iterating visually instead of wrangling prompts.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Marketing teams&lt;/strong&gt; can generate regionalized or personalized creative variants from a shared layout.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Product designers&lt;/strong&gt; can prototype in context, turning layouts into realistic screens without leaving the editor.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Storytellers and content creators&lt;/strong&gt; can use annotated sketches to generate detailed illustrations or scene variations.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;E-commerce platforms&lt;/strong&gt; can give sellers the power to visually customize their brand materials with AI assistance.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In every case, Visual Prompting replaces friction with flow and text-based prompting with something more expressive, more reliable, and more fun.&lt;/p&gt;
&lt;h2 id=&quot;built-for-this-multimodal-models-and-cesdks-plugin-system&quot;&gt;Built for This: Multimodal Models and CE.SDK’s Plugin System&lt;/h2&gt;
&lt;p&gt;Visual Prompting is only possible because of two parallel advancements:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Multimodal AI models&lt;/strong&gt;, such as OpenAI’s &lt;code&gt;gpt-image-1&lt;/code&gt;, that can interpret both images and text, understand spatial relationships, and respond to annotated cues.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A flexible, composable editor SDK&lt;/strong&gt; like CE.SDK, which enables the construction of visual prompts on a live canvas, and makes it easy to integrate AI models directly into the design flow.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Our SDK was built from the ground up to support &lt;strong&gt;AI-first creative workflows&lt;/strong&gt;. Its plugin architecture allows you to add any model or API, image generation, video generation, captioning, text rewriting and use it natively inside the editor without the need to switch tools or copy/paste.&lt;/p&gt;
&lt;p&gt;Generative AI’s full potential is only unlocked when it is embedded directly into the tools creatives use not siloed in chatbots or separate interfaces. Visual Prompting allows that embedding to go even deeper, aligning the &lt;strong&gt;mode of input (visual)&lt;/strong&gt; with the &lt;strong&gt;desired output (visual)&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id=&quot;explore-it-yourself&quot;&gt;Explore It Yourself&lt;/h3&gt;
&lt;p&gt;🎨 Try out Visual Prompting in our &lt;a href=&quot;https://img.ly/demos/ai-editor/web/&quot;&gt;AI Editor demo&lt;/a&gt;&lt;br&gt;
📘 &lt;a href=&quot;https://img.ly/docs/cesdk/js/user-interface/ai-integration/integrate-8e906c/&quot;&gt;Learn How to Integrate AI into CE.SDK&lt;/a&gt;&lt;br&gt;
💬 &lt;a href=&quot;https://img.ly/forms/contact-sales/&quot;&gt;Contact Us&lt;/a&gt; to Bring Visual Prompting to Your Product&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="https://blog.img.ly/2025/07/cesdk-2025-07-25T12_20_58.735Z--1-.png" medium="image"/><category>AI</category><category>CE.SDK</category><category>Creativity</category><category>Creative Workflows</category></item><item><title>Build vs. Buy: Is Fabric.js Right for You</title><link>https://img.ly/blog/build-vs-buy-is-fabric-js-right-for-you/</link><guid isPermaLink="true">https://img.ly/blog/build-vs-buy-is-fabric-js-right-for-you/</guid><description>This blog article provides a detailed comparison between Fabric.js and IMG.LY’s CreativeEditor SDK (CE.SDK) for teams building in-product design editors. It outlines the strengths of Fabric.js as a low-level canvas library, but focuses on the growing limitations teams face as they scale.</description><pubDate>Thu, 15 May 2025 09:31:43 GMT</pubDate><content:encoded>&lt;p&gt;When you’re building a design editor into your product whether for social content, marketing assets, video, or design workflows choosing the right foundation matters. Many teams start with Fabric.js, an open-source canvas library, attracted by its flexibility and permissive license. But as the project scope grows, so do the limitations.&lt;/p&gt;
&lt;p&gt;This article explores the advantages, but also the pitfalls of building a creative tool with Fabric.js and why many product teams ultimately decide to switch or start with a commercial SDK like &lt;a href=&quot;https://IMG.LY&quot;&gt;IMG.LY&lt;/a&gt;’s CreativeEditor SDK (CE.SDK).&lt;/p&gt;
&lt;p&gt;Consult our &lt;a href=&quot;https://img.ly/fabricjs-alternative/&quot;&gt;comparison page of Fabric.js and IMG.LY&lt;/a&gt; for a feature by feature breakdown of the differences between our SDKs and Fabric.js.&lt;/p&gt;
&lt;h2 id=&quot;why-teams-switch-from-fabricjs&quot;&gt;&lt;strong&gt;Why Teams Switch from Fabric.js&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;We’ve spoken with dozens of companies, from startups to enterprise clients, who evaluated Fabric.js and in many cases opted for IMG.LY instead. Here’s what we consistently hear:&lt;/p&gt;
&lt;h3 id=&quot;1-time-to-market-vs-reinventing-the-wheel&quot;&gt;&lt;strong&gt;1. Time-to-Market vs. Reinventing the Wheel&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Open-source is attractive at first. But once the excitement of building wears off, teams often face a sobering reality: stitching together a decent editing experience from Fabric.js is a months-long effort.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;“We used Fabric.js before, but if IMG.LY’s SDK is cost-effective, we’d rather not reinvent the wheel.”&lt;/p&gt;
&lt;p&gt;— B2B SaaS Prospect&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;They want a modern editor UI, fast iteration cycles, and a stable base not to spend their next quarter building the basics like zoom, text editing, object grouping, layer management, or responsive templates.&lt;/p&gt;
&lt;h3 id=&quot;2-lack-of-advanced-features&quot;&gt;&lt;strong&gt;2. Lack of Advanced Features&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Fabric.js excels at low-level canvas manipulation&lt;/strong&gt;, offering a robust API for working directly with objects on the HTML5 canvas - shapes, images, text, and transformations. It’s a great starting point for developers who want fine-grained control over rendering and interactivity. &lt;strong&gt;However, it falls short when it comes to modern editing paradigms&lt;/strong&gt;. Fabric.js provides no native concept of reusable templates, lacks support for layout and design constraints like alignment guides or snapping behavior, and does not include higher-level abstractions for content-aware or AI-powered editing. &lt;strong&gt;In short, Fabric.js gives you a raw toolkit, not a polished, plug-and-play editing solution.&lt;/strong&gt; If you’re building a sophisticated design application, you’ll need to layer these capabilities yourself or integrate additional frameworks to bridge the gap.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;“We needed a template system and advanced photo editing. Open-source libraries couldn’t do it out of the box.”&lt;/p&gt;
&lt;p&gt;— Marketplace Platform Prospect&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Teams chasing parity with Canva or Adobe tools quickly hit walls. What starts as a proof-of-concept becomes a rebuild of the tables stakes of design editing.&lt;/p&gt;
&lt;h2 id=&quot;real-world-use-cases--vertical-specific-challenges&quot;&gt;&lt;strong&gt;Real-World Use Cases &amp;#x26; Vertical-Specific Challenges&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;We’ve seen teams from various industries reach the same conclusion: Fabric.js doesn’t scale to meet their creative, technical, or business needs. A few common themes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;E-commerce &amp;#x26; Web-to-Print&lt;/strong&gt;: Retailers building product customization flows need template constraints, consistent output quality, and export formats that go beyond canvas. They can’t afford rendering inconsistencies or unpredictable export fidelity.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;MarTech &amp;#x26; Social Media Tools&lt;/strong&gt;: Marketers need batch generation, AI-assisted creative workflow, brand constraints and consistent creative output across devices. Fabric.js lacks built-in support for these workflows.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Mobile-first Apps&lt;/strong&gt;: Teams building hybrid apps or web apps that need to be functional on mobile struggle with Fabric.js’ &lt;a href=&quot;https://github.com/fabricjs/fabric.js/issues/6980&quot;&gt;lacklustre mobile support&lt;/a&gt;. The inability to deliver a consistent UX across iOS, Android, and Web leads to dropped features or split tech stacks.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Video and Multimedia Platforms&lt;/strong&gt;: Fabric.js has no native video support, track-based editing, or multi-frame logic. Prospects in this space frequently abandon their Fabric.js POCs once real constraints emerge.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A major pain point across all verticals is &lt;strong&gt;rendering consistency&lt;/strong&gt;. CE.SDK renders identically across browser, server, and mobile thanks to a shared rendering core (&lt;a href=&quot;https://img.ly/docs/cesdk/node/get-started/overview-e18f40/&quot;&gt;CreativeEngine&lt;/a&gt;). This is critical for:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Workflows that begin on web and finish on mobile&lt;/li&gt;
&lt;li&gt;Server-side generation (e.g., previews, PDFs, batch exports)&lt;/li&gt;
&lt;li&gt;Feature parity and UX reliability across platforms&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Check out our &lt;a href=&quot;https://img.ly/demos/&quot;&gt;demo page&lt;/a&gt; to explore these use cases.&lt;/p&gt;
&lt;h2 id=&quot;technical-debt-the-hidden-cost-of-fabricjs&quot;&gt;&lt;strong&gt;Technical Debt: The Hidden Cost of Fabric.js&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;On paper, Fabric.js seems like a fast way to get started. But its real cost emerges over time in performance bottlenecks, missing architecture, and maintenance drag.&lt;/p&gt;
&lt;h3 id=&quot;aging-codebase-and-patchy-maintenance&quot;&gt;&lt;strong&gt;Aging Codebase and Patchy Maintenance&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;As of 2025, Fabric.js has over 400 open issues on GitHub, some dating back years. Many feature requests such as better performance for large object sets, async rendering, or text-on-path improvements are either unresolved or uncertainly prioritized.&lt;/p&gt;
&lt;p&gt;The core maintainers do their best, but progress is slow. You’re relying on volunteer effort for critical infrastructure.&lt;/p&gt;
&lt;h3 id=&quot;slow-issue-resolution&quot;&gt;&lt;strong&gt;Slow Issue Resolution&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Many of the GitHub issues (e.g., &lt;a href=&quot;https://github.com/fabricjs/fabric.js/issues/5885&quot;&gt;#5885&lt;/a&gt;, &lt;a href=&quot;https://github.com/fabricjs/fabric.js/issues/6582&quot;&gt;#6582&lt;/a&gt;) highlight long-standing rendering bugs and performance problems. These are hard to fix, and updates can break your own hacks or workarounds.&lt;/p&gt;
&lt;p&gt;Unlike a commercial SDK, Fabric.js doesn’t guarantee backward compatibility. No SLAs. Little roadmap visibility.&lt;/p&gt;
&lt;h2 id=&quot;customer-feedback-why-they-walked-away&quot;&gt;&lt;strong&gt;Customer Feedback: Why They Walked Away&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Here are some soundbites from our customers and prospects on why they decided against building with Fabric.js:&lt;/p&gt;
&lt;p&gt;&lt;em&gt;“The quality and UX didn’t meet our standards.”&lt;/em&gt;&lt;br&gt;
Product Customizer&lt;/p&gt;
&lt;p&gt;&lt;em&gt;“Developer experience was poor. We spent more time debugging than building features.”&lt;/em&gt;&lt;br&gt;
B2B SaaS Ad Design&lt;/p&gt;
&lt;p&gt;&lt;em&gt;“We needed enterprise support, scalability, and documentation. Open source didn’t cut it.”&lt;/em&gt;&lt;br&gt;
Web to print Customer&lt;/p&gt;
&lt;p&gt;&lt;em&gt;“Cross-platform support and minimum SDK version constraints were a major blocker.”&lt;/em&gt;&lt;br&gt;
Claim Management Application&lt;/p&gt;
&lt;h2 id=&quot;voice-of-the-dev&quot;&gt;&lt;strong&gt;Voice of the Dev&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Beyond enterprise evaluations, developers working with Fabric.js often share their struggles publicly on GitHub, Stack Overflow, and developer forums. Their commentary offers real-world insight into day-to-day frustrations:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;“FabricJS doesn’t seem to work well on Mobile.”&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/fabricjs/fabric.js/issues/6980&quot;&gt;Github issue #6980&lt;/a&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;“Fabric.js object controls don’t work until after a common selection is made. Had to hack around it with extra event listeners.”&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a href=&quot;https://stackoverflow.com/questions/64697215/fabricjs-object-controls-not-working-until-common-selection&quot;&gt;StackOverflow&lt;/a&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;“Trying to integrate Fabric.js with Next.js + Rollup is a mess. Unexpected tokens, config rewrites — it doesn’t play well with modern bundlers.”&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/fabricjs/fabric.js/issues/8444&quot;&gt;GitHub Issue #8444&lt;/a&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;“We’re blocked by the use of ‘unsafe-eval’ due to our content security policy. Fabric.js needs a rewrite to be CSP-compliant.”&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/fabricjs/fabric.js/issues/9666&quot;&gt;GitHub Issue #9666&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Concrete takeaways from developer feedback:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Basic UI control quirks&lt;/strong&gt;: Common issues with selection handling and control behavior require manual workarounds.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Compatibility friction&lt;/strong&gt;: Fabric.js does not play well with modern build tools like Next.js and Rollup without additional configuration or patching.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Security blockers&lt;/strong&gt;: The reliance on &lt;code&gt;unsafe-eval&lt;/code&gt; creates CSP conflicts, making it unsuitable for projects with strict security requirements.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Uncertain roadmap&lt;/strong&gt;: Critical features like async rendering may sit open for years, and niche features important to your use case may never land if they don’t align with community priorities.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;While every open source project faces these difficulties, they do represent systemic issues with performance, extensibility, and production-readiness. For teams building long-lived, customer-facing tools, this developer feedback is often a leading indicator of future roadblocks.&lt;/p&gt;
&lt;h2 id=&quot;back-of-the-envelope-cost-of-building-with-fabricjs&quot;&gt;&lt;strong&gt;Back-of-the-Envelope Cost of Building with Fabric.js&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Many teams underestimate the cost of building a design editor from scratch. Here’s a rough breakdown of the time investment we’ve heard from teams who tried:&lt;/p&gt;





























&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;&lt;strong&gt;Task&lt;/strong&gt;&lt;/th&gt;&lt;th&gt;&lt;strong&gt;Estimated Engineering Time&lt;/strong&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Core canvas-based editor UI (zoom, drag, resize, selection, tool switching)&lt;/td&gt;&lt;td&gt;6–8 weeks&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;SVG support, snapping, grouping, and layer management&lt;/td&gt;&lt;td&gt;4–6 weeks&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Template constraints&lt;/td&gt;&lt;td&gt;3–5 weeks&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Cross-platform adaptation (ironing out issues on mobile browsers, defining fallbacks)&lt;/td&gt;&lt;td&gt;4–6 weeks&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Bug triage, ongoing maintenance, and refactors (first 12–18 months)&lt;/td&gt;&lt;td&gt;6–9 weeks&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Total: ~6–8 months of senior dev time&lt;/strong&gt; just to get to feature parity with what CE.SDK offers out of the box.&lt;/p&gt;
&lt;p&gt;This doesn’t include the cost of QA, product management, or future extensibility. Nor the opportunity cost of what your team could be building instead.&lt;/p&gt;
&lt;h2 id=&quot;imglys-cesdk-a-ready-made-solution-built-for-growth&quot;&gt;&lt;strong&gt;IMG.LY’s CE.SDK: A Ready-Made Solution Built for Growth&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;By contrast, IMG.LY offers a production-grade SDK built for cross-platform creative tools, backed by a dedicated engineering team and used by major apps across industries.&lt;/p&gt;
&lt;div&gt;&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;&lt;strong&gt;Category&lt;/strong&gt;&lt;/th&gt;&lt;th&gt;&lt;strong&gt;&lt;a href=&quot;https://IMG.LY&quot;&gt;IMG.LY&lt;/a&gt; (CE.SDK)&lt;/strong&gt;&lt;/th&gt;&lt;th&gt;&lt;strong&gt;Fabric.js&lt;/strong&gt;&lt;/th&gt;&lt;th&gt;&lt;strong&gt;Notes&lt;/strong&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Out-of-the-box UI&lt;/td&gt;&lt;td&gt;✅ Prebuilt modern UI&lt;/td&gt;&lt;td&gt;❌ Build from scratch&lt;/td&gt;&lt;td&gt;CE.SDK ships with full UI/UX patterns&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Plugin System&lt;/td&gt;&lt;td&gt;✅ Native plugin architecture&lt;/td&gt;&lt;td&gt;⚠️ Manual code extension&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://IMG.LY&quot;&gt;IMG.LY&lt;/a&gt; supports formal plugin APIs&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Free Drawing Tools&lt;/td&gt;&lt;td&gt;⚠️ Requires integration&lt;/td&gt;&lt;td&gt;✅ Built-in pencil + shapes&lt;/td&gt;&lt;td&gt;Fabric.js excels at shape-level control&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Advanced Editing&lt;/td&gt;&lt;td&gt;✅ Vector + raster support, layering, filters&lt;/td&gt;&lt;td&gt;⚠️ Basic vector only&lt;/td&gt;&lt;td&gt;CE.SDK includes pro-grade editing capabilities&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Templating System&lt;/td&gt;&lt;td&gt;✅ Dynamic placeholders + constraints&lt;/td&gt;&lt;td&gt;❌ Manual implementation&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://IMG.LY&quot;&gt;IMG.LY&lt;/a&gt; supports automation workflows&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Video Editing&lt;/td&gt;&lt;td&gt;✅ Multi-track timeline editor&lt;/td&gt;&lt;td&gt;❌ Not supported&lt;/td&gt;&lt;td&gt;No video support in Fabric.js&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Cross-Platform&lt;/td&gt;&lt;td&gt;✅ Web, iOS, Android, Node, Flutter&lt;/td&gt;&lt;td&gt;⚠️ Web only&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://IMG.LY&quot;&gt;IMG.LY&lt;/a&gt; has native SDKs and server tools&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Asset Integrations&lt;/td&gt;&lt;td&gt;✅ Unsplash, Getty, Brandfolder&lt;/td&gt;&lt;td&gt;❌ Manual setup&lt;/td&gt;&lt;td&gt;CE.SDK integrates assets out-of-the-box&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Enterprise Support&lt;/td&gt;&lt;td&gt;✅ SLAs, onboarding, deployment help&lt;/td&gt;&lt;td&gt;❌ Community support only&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://IMG.LY&quot;&gt;IMG.LY&lt;/a&gt; offers dedicated engineering help&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Design File Import Support&lt;/td&gt;&lt;td&gt;✅ PSD, AI, INDD&lt;/td&gt;&lt;td&gt;⚠️ None&lt;/td&gt;&lt;td&gt;CE.SDK supports importing common design file formats&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;AI Editing&lt;/td&gt;&lt;td&gt;✅ Background removal, integrate any AI model for image/video/text/audio gen&lt;/td&gt;&lt;td&gt;❌ Not supported&lt;/td&gt;&lt;td&gt;AI-native editing pipeline available&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Rendering Consistency&lt;/td&gt;&lt;td&gt;✅ Unified engine across platforms&lt;/td&gt;&lt;td&gt;⚠️ Browser dependent&lt;/td&gt;&lt;td&gt;CE.SDK ensures pixel-parity on all targets&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;h2 id=&quot;when-does-fabricjs-still-make-sense&quot;&gt;&lt;strong&gt;When Does Fabric.js Still Make Sense?&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;This article might read like a put-down, but it’s not supposed to be that. Until now we were simply evaluating Fabric.js with respect to building fully-featured design editing tools. There are, however, a number of well-scoped use cases where Fabric.js excels. These include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Building an &lt;strong&gt;internal tool&lt;/strong&gt;, proof of concept, or educational app&lt;/li&gt;
&lt;li&gt;Creating a &lt;strong&gt;custom, single-platform sketching or annotation layer&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Developing a &lt;strong&gt;low-fidelity prototype&lt;/strong&gt; with full code-level control&lt;/li&gt;
&lt;li&gt;Building canvas collaboration tools such as &lt;strong&gt;drag and drop editors&lt;/strong&gt; or interactive canvases, such as Miro.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;An example of where Fabric.js worked well is a tool such as &lt;a href=&quot;https://mockola.com/&quot;&gt;Mockola&lt;/a&gt;, a drag-and-drop diagram editor, where shape manipulation, lightweight rendering, and canvas-level control are more important than cross-platform design fidelity. Mockola’s is a good use case, because its emphasis on real-time collaboration and planning rather than high-fidelity design. This means that features like free drawing, drag-and-drop, and JSON-based serialization are a great fit, all of which Fabric.js supports out of the box.&lt;/p&gt;
&lt;p&gt;For teams willing to invest in building and maintaining their own editor infrastructure and whose use case does not require cross-platform parity, advanced media editing, or enterprise scaling Fabric.js continues to be a compelling foundation.&lt;/p&gt;
&lt;p&gt;However, for products that need to ship quickly, scale reliably, and support modern creative workflows across platforms, IMG.LY’s CE.SDK offers the infrastructure and polish that customers now expect.&lt;/p&gt;
&lt;h2 id=&quot;the-bottom-line&quot;&gt;&lt;strong&gt;The Bottom Line&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;As &lt;a href=&quot;https://www.youtube.com/watch?v=TBOazhxtq5s&quot;&gt;MIT CIO Symposium speaker&lt;/a&gt; Mark Holst-Knudsen put it:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;“You shouldn’t build anything that’s available off the shelf because it’s not a source of competitive advantage if everybody else can avail themselves of it.”&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;For most companies, building an editor from scratch is not a competitive edge, it’s a distraction from shipping and scaling. And Fabric.js, while capable in the hands of specialists, isn’t a shortcut to the finish line.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://IMG.LY&quot;&gt;IMG.LY&lt;/a&gt;’s CE.SDK gives you a head start with advanced editing features, AI automation, template support, plugin extensibility, rendering consistency, and battle-tested scalability across platforms. We put together a &lt;a href=&quot;https://img.ly/blog/imgly-impact-report/&quot;&gt;data report based on 600+ customers, 28 structured customer interviews, and a self-reported customer survey&lt;/a&gt; that speaks for itself when it comes to real-life outcomes.&lt;/p&gt;
&lt;p&gt;So before you dive into Fabric.js and start reinventing core functionality, ask yourself: What competitive advantage are you going to build &lt;strong&gt;on top&lt;/strong&gt; of a design editor and which functionality are you better served buying from a trusted vendor?&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="https://blog.img.ly/2025/05/build-vs-buy-imgly-fabricjs-sdk.jpg" medium="image"/><category>Fabric.js</category><category>CE.SDK</category><category>OpenSource</category></item><item><title>OpenAI GPT-4o Image Generation (gpt-image-1) API: A Complete Guide for Creative Workflows for 2025</title><link>https://img.ly/blog/openai-gpt-4o-image-generation-api-gpt-image-1-a-complete-guide-for-creative-workflows-for-2025/</link><guid isPermaLink="true">https://img.ly/blog/openai-gpt-4o-image-generation-api-gpt-image-1-a-complete-guide-for-creative-workflows-for-2025/</guid><description>Learn how to integrate OpenAI’s gpt-image-1 API into modern creative applications. This complete 2025 guide covers technical setup, CE.SDK integration, prompt engineering, and tips for building real multimodal creative workflows.</description><pubDate>Mon, 28 Apr 2025 07:55:48 GMT</pubDate><content:encoded>&lt;h2 id=&quot;update-ai-first-visual-editing&quot;&gt;Update: AI-first Visual Editing&lt;/h2&gt;
&lt;p&gt;A day after the release of the &lt;code&gt;gpt-image-1&lt;/code&gt; API, we took it for a spin and integrated it into CreativeEditor SDK. Users can now generate images, create variants and use the canvas to compose visual prompts with our design editor. See it in action:&lt;/p&gt;
&lt;div class=&quot;cta-button-wrapper&quot;&gt;&lt;a href=&quot;https://cdn.img.ly/demo/gpt-image-1/v1/&quot; target=&quot;_blank&quot; class=&quot;cta-button&quot;&gt;Open AI Editor Demo Page&lt;/a&gt;&lt;/div&gt;
&lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;The release of OpenAI’s &lt;code&gt;gpt-image-1&lt;/code&gt; model signals a pivotal shift in the creative developer landscape, one that moves beyond static, one-shot image generation and toward a more dynamic, multimodal interaction model. Until recently, most image APIs followed a predictable pattern: submit a prompt, receive a finished image. The process was useful, but flat. What’s changing now is not just image quality or style fidelity, but the shape of the workflow itself. With &lt;code&gt;gpt-image-1&lt;/code&gt;, built on the GPT-4o foundation, developers can start designing creative tools that feel conversational and iterative. This evolution invites a new kind of interface where prompting, tweaking, and refining happen inside the canvas, not outside of it.&lt;/p&gt;
&lt;p&gt;For teams building creative editing experience into their app, this moment coincides with the release of &lt;a href=&quot;https://img.ly/demos/ai-editor/web/&quot;&gt;IMG.LY’s AI Editor SDK&lt;/a&gt;, a powerful, fully integrated toolkit designed for generative workflows. The SDK is already equipped to support interactive image generation, contextual editing, and multimodal inputs, and you can try it today through &lt;a href=&quot;https://img.ly/demos/&quot;&gt;this live demo&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;This guide is a comprehensive introduction to the &lt;code&gt;gpt-image-1&lt;/code&gt; API, but it also goes further. It’s not just about wiring up an endpoint, it’s about rethinking what image generation means in a user-centric product.&lt;/p&gt;
&lt;p&gt;From prompt handling to interactive iteration, we’ll walk through how to design creative cycles, not just outputs. This guide explores how to make that shift, how to go from generating images to integrating &lt;code&gt;gpt-image-1&lt;/code&gt; into real creative cycles, where AI becomes a tool that bends to user intent, not the other way around.&lt;/p&gt;
&lt;h2 id=&quot;overview-of-gpt-image-1&quot;&gt;Overview of &lt;code&gt;gpt-image-1&lt;/code&gt;&lt;/h2&gt;
&lt;p&gt;OpenAI’s &lt;code&gt;gpt-image-1&lt;/code&gt; model, released in April 2025, is the latest evolution in the company’s generative image lineup and marks a turning point in how developers approach visual creation inside applications. Built on the same multimodal foundation as GPT-4o, this model allows applications to move beyond one-shot static generation and instead build toward more conversational, iterative image workflows.&lt;/p&gt;
&lt;h3 id=&quot;model-architecture-and-capabilities&quot;&gt;Model Architecture and Capabilities&lt;/h3&gt;
&lt;p&gt;&lt;code&gt;gpt-image-1&lt;/code&gt; is rooted in GPT-4o’s ability to understand and generate across modalities. It is designed to produce high-resolution images (up to 4096×4096 pixels) based on natural language prompts. The model handles complex scenes with more fidelity than previous iterations and provides improved consistency in how it interprets detailed descriptions. This is particularly relevant for tools that need reliability when turning prompt inputs into design elements.&lt;/p&gt;
&lt;h3 id=&quot;parameter-control&quot;&gt;Parameter Control&lt;/h3&gt;
&lt;p&gt;Developers working with &lt;code&gt;gpt-image-1&lt;/code&gt; have access to a streamlined set of parameters, here is a subset of the most important ones:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;prompt&lt;/code&gt;: The primary text input describing the desired image.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;size&lt;/code&gt;: Choose between “1024x1024”, “1024x1536” (portrait), “1536x1024” (landscape), or “auto” (default, based on prompt).&lt;/li&gt;
&lt;li&gt;&lt;code&gt;n&lt;/code&gt;: Number of images to generate (default is 1).&lt;/li&gt;
&lt;li&gt;&lt;code&gt;response_format&lt;/code&gt;: Always returns &lt;code&gt;b64_json&lt;/code&gt;. URL outputs are not supported.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Unlike DALL·E 3, &lt;code&gt;gpt-image-1&lt;/code&gt; does not accept &lt;code&gt;style&lt;/code&gt; modifiers or &lt;code&gt;quality&lt;/code&gt; settings. It is designed for straightforward, high-fidelity image creation driven purely by the text prompt and size selection.&lt;/p&gt;
&lt;p&gt;Full documentation of these options is available via &lt;a href=&quot;https://platform.openai.com/docs/guides/images/usage&quot;&gt;OpenAI’s official guide&lt;/a&gt;.&lt;/p&gt;
&lt;h3 id=&quot;style-and-use-case-alignment&quot;&gt;Style and Use Case Alignment&lt;/h3&gt;
&lt;p&gt;By supporting a wide range of stylistic templates, &lt;code&gt;gpt-image-1&lt;/code&gt; positions itself as a flexible backend for everything from marketing collateral to storyboarding tools. The output can be tailored to suit technical illustrations, concept art, or even photorealistic renderings, allowing developers to map visual outputs more directly to brand or product requirements.&lt;/p&gt;
&lt;h3 id=&quot;limitations-and-future-direction&quot;&gt;Limitations and Future Direction&lt;/h3&gt;
&lt;p&gt;As of April 2025, &lt;code&gt;gpt-image-1&lt;/code&gt; supports only one image per request and does not offer fine-grained image editing or inpainting. However, its tight coupling with GPT-4o suggests that future iterations may embrace persistent context, conversational refinement, or even integrated image-plus-text exchanges within the same session. For developers building editors or multimodal workflows, the current model lays a strong foundation for these future capabilities.&lt;/p&gt;
&lt;h2 id=&quot;api-setup-and-usage&quot;&gt;API Setup and Usage&lt;/h2&gt;
&lt;h3 id=&quot;21-get-access&quot;&gt;2.1 Get Access&lt;/h3&gt;
&lt;p&gt;To start using &lt;code&gt;gpt-image-1&lt;/code&gt;, developers must first register for access via the OpenAI platform at &lt;a href=&quot;https://platform.openai.com/&quot;&gt;platform.openai.com&lt;/a&gt;. Access requires an API key, which is tied to your OpenAI account and associated usage limits based on your billing tier. Be sure to confirm that your account is approved for image generation, as availability may differ by region and subscription level. Once authenticated, keys can be created in your dashboard and stored securely in your server or development environment.&lt;/p&gt;
&lt;h3 id=&quot;22-first-image-generation-nodejs-example&quot;&gt;2.2 First Image Generation (Node.js Example)&lt;/h3&gt;
&lt;p&gt;The image generation API for &lt;code&gt;gpt-image-1&lt;/code&gt; can be used directly via OpenAI’s official Node.js client. Below is a complete example showing how to send a prompt and receive an image URL in response:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; OpenAI &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;openai&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; fs &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;fs&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; openai&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; new&lt;/span&gt;&lt;span&gt; OpenAI&lt;/span&gt;&lt;span&gt;({&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  apiKey: process.env.&lt;/span&gt;&lt;span&gt;OPENAI_API_KEY&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;// make sure this is securely set&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;});&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;async&lt;/span&gt;&lt;span&gt; function&lt;/span&gt;&lt;span&gt; generateImage&lt;/span&gt;&lt;span&gt;() {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  try&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    const&lt;/span&gt;&lt;span&gt; prompt&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; `&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    A studio ghibli style illustration of a cyberpunk girl holding a butterfly on her finger.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    `&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    const&lt;/span&gt;&lt;span&gt; result&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; await&lt;/span&gt;&lt;span&gt; openai.images.&lt;/span&gt;&lt;span&gt;generate&lt;/span&gt;&lt;span&gt;({&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      model: &lt;/span&gt;&lt;span&gt;&apos;gpt-image-1&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      prompt,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      size: &lt;/span&gt;&lt;span&gt;&apos;1024x1024&apos;&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;// or &quot;1024x1536&quot;, &quot;1536x1024&quot;, or &quot;auto&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    });&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    const&lt;/span&gt;&lt;span&gt; image_base64&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; result.data[&lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;].b64_json;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    const&lt;/span&gt;&lt;span&gt; image_bytes&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; Buffer.&lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt;(image_base64, &lt;/span&gt;&lt;span&gt;&apos;base64&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    fs.&lt;/span&gt;&lt;span&gt;writeFileSync&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;butterfly.png&apos;&lt;/span&gt;&lt;span&gt;, image_bytes);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    console.&lt;/span&gt;&lt;span&gt;log&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;Image saved as butterfly.png&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  } &lt;/span&gt;&lt;span&gt;catch&lt;/span&gt;&lt;span&gt; (err) {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    console.&lt;/span&gt;&lt;span&gt;error&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;Error generating image:&apos;&lt;/span&gt;&lt;span&gt;, err);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  }&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;}&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;generateImage&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Remember that all outputs from &lt;code&gt;gpt-image-1&lt;/code&gt; are delivered as base64-encoded JSON. Developers should decode this data for display, storage, or further processing within their applications. For complete parameter options and examples, consult the &lt;a href=&quot;https://platform.openai.com/docs/guides/images&quot;&gt;OpenAI Images API guide&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;integrating-with-cesdk&quot;&gt;Integrating with CE.SDK&lt;/h2&gt;
&lt;p&gt;Embedding &lt;code&gt;gpt-image-1&lt;/code&gt; into a creative editor like CE.SDK is about more than just piping an image into a canvas. It reshapes how users interact with content creation, bridging manual design work and AI-driven generation within the same editing environment. Rather than operating as a standalone prompt generator, &lt;code&gt;gpt-image-1&lt;/code&gt; becomes a continuous creative partner inside your editor. For in in-depth technical guide on how to integrate &lt;code&gt;gpt-image-1&lt;/code&gt; stay tuned for our upcoming tutorial, &lt;a href=&quot;https://share.hsforms.com/1IgAOV1wASXGPnFG4ZPLejg1hk3i?ref=img.ly&quot;&gt;sign up to our newsletter&lt;/a&gt; to be notified when it goes live.&lt;/p&gt;
&lt;h3 id=&quot;embedding-image-generation-in-a-creative-editing-workflow&quot;&gt;Embedding Image Generation in a Creative Editing Workflow&lt;/h3&gt;
&lt;p&gt;The natural entry point for &lt;code&gt;gpt-image-1&lt;/code&gt; inside CE.SDK is through a dual-mode experience: offering users the option to start either from scratch or from existing context. In “from scratch” mode, a user might open a blank scene and initiate an image generation by writing a prompt for example, “Create a vibrant festival scene at sunset.” The result appears directly on the canvas, immediately editable like any other design element.&lt;/p&gt;
&lt;p&gt;Where &lt;code&gt;gpt-image-1&lt;/code&gt; shows its real potential is in “in-context editing.” Here, users interact with existing content (a background, a product shot, or a decorative element) and trigger AI enhancements based on that visual context. A user might select an image of a bird, as in the example below and ask for variants, initiate a background swap, or request a change like adding more birds in a conversational interface embedded in the editor. Because CE.SDK treats generated images as first-class canvas elements, context such as positioning, layering, and cropping is preserved throughout the process.&lt;/p&gt;
&lt;p&gt;Let’s see what this might look like in practice. We positioned an image of a single bird on our canvas, opening the AI context menu we can now manipulate that image in place using the OpenAI API:&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 940px) 940px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;940&quot; height=&quot;560&quot; src=&quot;https://img.ly/_astro/Screenshot-2025-04-25-at-16.23.20_2408iY.webp&quot; srcset=&quot;/_astro/Screenshot-2025-04-25-at-16.23.20_hoIP0.webp 640w, /_astro/Screenshot-2025-04-25-at-16.23.20_Z1lGLpv.webp 750w, /_astro/Screenshot-2025-04-25-at-16.23.20_SQHtM.webp 828w, /_astro/Screenshot-2025-04-25-at-16.23.20_2408iY.webp 940w&quot;&gt;&lt;/p&gt;
&lt;p&gt;We edit the image and prompt the API to add more birds:&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1104px) 1104px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1104&quot; height=&quot;638&quot; src=&quot;https://img.ly/_astro/Screenshot-2025-04-25-at-11.19.07_Z1nRit4.webp&quot; srcset=&quot;/_astro/Screenshot-2025-04-25-at-11.19.07_Z2uDN7i.webp 640w, /_astro/Screenshot-2025-04-25-at-11.19.07_Z1XMUJ0.webp 750w, /_astro/Screenshot-2025-04-25-at-11.19.07_9QPa7.webp 828w, /_astro/Screenshot-2025-04-25-at-11.19.07_rW40P.webp 1080w, /_astro/Screenshot-2025-04-25-at-11.19.07_Z1nRit4.webp 1104w&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1056px) 1056px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1056&quot; height=&quot;548&quot; src=&quot;https://img.ly/_astro/Screenshot-2025-04-25-at-11.19.31_Z2pGvMv.webp&quot; srcset=&quot;/_astro/Screenshot-2025-04-25-at-11.19.31_Abfg0.webp 640w, /_astro/Screenshot-2025-04-25-at-11.19.31_1OeNO8.webp 750w, /_astro/Screenshot-2025-04-25-at-11.19.31_ZFflGu.webp 828w, /_astro/Screenshot-2025-04-25-at-11.19.31_Z2pGvMv.webp 1056w&quot;&gt;&lt;/p&gt;
&lt;p&gt;We see that the model correctly identified the type of bird in the picture (seagull) and filled it in with a swarm of flying seagulls.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1130px) 1130px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1130&quot; height=&quot;622&quot; src=&quot;https://img.ly/_astro/Screenshot-2025-04-25-at-11.35.50_Z1v4WfN.webp&quot; srcset=&quot;/_astro/Screenshot-2025-04-25-at-11.35.50_Z1yln80.webp 640w, /_astro/Screenshot-2025-04-25-at-11.35.50_Z1hUCg9.webp 750w, /_astro/Screenshot-2025-04-25-at-11.35.50_Z24uBxa.webp 828w, /_astro/Screenshot-2025-04-25-at-11.35.50_Z2ufBFR.webp 1080w, /_astro/Screenshot-2025-04-25-at-11.35.50_Z1v4WfN.webp 1130w&quot;&gt;&lt;/p&gt;
&lt;p&gt;We can now continue to work with the image, overlaying filters, changing the texture, cropping etc.&lt;/p&gt;
&lt;h3 id=&quot;switching-between-manual-edits-and-ai-powered-enhancements&quot;&gt;Switching Between Manual Edits and AI-Powered Enhancements&lt;/h3&gt;
&lt;p&gt;A critical design principle when integrating &lt;code&gt;gpt-image-1&lt;/code&gt; is giving users freedom to toggle between manual edits and AI suggestions. Manual edits should always remain possible after generation, e.g. cropping, masking, compositing while users can also seamlessly prompt &lt;code&gt;gpt-image-1&lt;/code&gt; for additional changes without losing prior work. Think of variant generation as a branch: a user picks a generated image and creates “forks” by asking for alternate styles, different lighting, or new thematic elements.&lt;/p&gt;
&lt;p&gt;In this setup, the generated image serves as a stable node in the creative graph, while edits and regenerations can attach contextually. This workflow minimizes user frustration by avoiding the “start over” penalty typical of isolated generation APIs. It also opens up more complex creative behaviors, like blending user-drawn sketches with AI-augmented refinements, or iteratively developing an asset library around a consistent visual theme.&lt;/p&gt;
&lt;p&gt;An upcoming in-depth tutorial will walk through implementing this multimodal workflow step-by-step, but the key takeaway is that &lt;code&gt;gpt-image-1&lt;/code&gt; shines brightest when it is embedded into a creative loop, not treated as a black-box generator, but as an interactive, iterative design companion.&lt;/p&gt;
&lt;h2 id=&quot;prompt-engineering-tips&quot;&gt;Prompt Engineering Tips&lt;/h2&gt;
&lt;p&gt;One of the most overlooked but critical factors in successful image generation is prompt design. With &lt;code&gt;gpt-image-1&lt;/code&gt;, prompt engineering isn’t just about describing an image. It’s about steering the model toward intent, tone, composition, and usability. Because the model is capable of rendering complex scenes and a wide range of styles, thoughtful phrasing and contextual hints can dramatically affect the outcome.&lt;/p&gt;
&lt;h3 id=&quot;writing-for-visual-intent&quot;&gt;Writing for Visual Intent&lt;/h3&gt;
&lt;p&gt;Start by clarifying what the image is supposed to communicate. Are you looking for atmosphere, action, product detail, or narrative clarity? A prompt like “a city skyline at night” is a starting point, but it leaves too much to chance. Adding elements like “view from a rooftop bar, with glowing signage and overcast haze” gives the model anchors for both composition and mood.&lt;/p&gt;
&lt;h3 id=&quot;leveraging-artistic-language&quot;&gt;Leveraging Artistic Language&lt;/h3&gt;
&lt;p&gt;You can further refine outputs by referencing mediums or artistic schools. Prompts that include terms like “in watercolor style,” “oil painting,” ”80s anime aesthetic,” or “studio photography” help the model lock onto a particular visual identity. These cues not only improve stylistic fidelity but also align the output with specific brand or genre expectations, which is especially important for products with a defined look and feel.&lt;/p&gt;
&lt;h3 id=&quot;creating-consistency-in-branded-outputs&quot;&gt;Creating Consistency in Branded Outputs&lt;/h3&gt;
&lt;p&gt;When generating a set of related images, such as social media creatives, campaign assets, or UI visuals, consistency becomes more important than variety. To achieve this, structure prompts with repeatable patterns and include brand elements such as color palettes, motifs, or reference characters. While &lt;code&gt;gpt-image-1&lt;/code&gt; doesn’t yet support persistent memory across requests, consistency can be enforced by prompting with the same style terms, layout descriptions, and constraints. Teams working within CE.SDK can even pair prompt templates with locked canvas layers to preserve composition between generations.&lt;/p&gt;
&lt;p&gt;Ultimately, good prompt engineering is not about verbosity but about clarity and constraint. It’s less like writing poetry and more like drafting a product spec. The best prompts are focused, directive, and give the model just enough creative freedom within clear boundaries. However, effective prompting should not burden the user. In practice, the interface should abstract most of the complexity away. Users can be guided toward better outputs through simple UI choices (selecting predefined styles, choosing themes, or adjusting mood settings) while the system dynamically enhances and augments their input behind the scenes. By managing the technical depth invisibly, you enable a creative process that feels intuitive and powerful without ever making prompt engineering the center of the user experience.&lt;/p&gt;
&lt;h2 id=&quot;real-world-use-cases&quot;&gt;Real-World Use Cases&lt;/h2&gt;
&lt;p&gt;The versatility of &lt;code&gt;gpt-image-1&lt;/code&gt; makes it especially impactful across a variety of industries where visual content creation is either a core product feature or a major operational need. Beyond isolated image generation, the model supports workflows that demand contextual awareness, brand consistency, and iterative refinement, key ingredients for modern digital products.&lt;/p&gt;
&lt;h3 id=&quot;web-to-print&quot;&gt;Web-to-Print&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/use-cases/web-to-print-design-tool/&quot;&gt;In web-to-print applications&lt;/a&gt;, customers expect to customize marketing materials, event invitations, signage, or packaging with minimal friction. By integrating &lt;code&gt;gpt-image-1&lt;/code&gt;, developers can offer template-driven personalization where users simply select a theme or enter a few keywords, and receive ready-to-edit visual assets. Combined with CE.SDK’s layout and editing capabilities, this enables a highly interactive experience where generated backgrounds, graphical elements, or themed illustrations can be dynamically placed into editable templates.&lt;/p&gt;
&lt;h3 id=&quot;social-media-marketing-and-martech&quot;&gt;Social Media Marketing and MarTech&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/industries/marketing-tech/&quot;&gt;Marketing teams rely on high-frequency content creation&lt;/a&gt;, often needing visually consistent, campaign-specific assets. &lt;code&gt;gpt-image-1&lt;/code&gt; can assist by automating the generation of background scenes, promotional visuals, and thematic graphics based on campaign briefs. Brands can define style presets aligned with their visual identity, making it easy for marketing teams to produce “on-brand” assets without heavy design overhead. Integrating image generation directly into campaign builders or social scheduling tools amplifies speed without sacrificing quality.&lt;/p&gt;
&lt;h3 id=&quot;digital-asset-management-dam&quot;&gt;Digital Asset Management (DAM)&lt;/h3&gt;
&lt;p&gt;Asset libraries often suffer from gaps: missing variants, seasonal versions, or content tailored to different demographics. &lt;a href=&quot;https://img.ly/industries/digital-asset-management/&quot;&gt;DAM systems&lt;/a&gt; can integrate &lt;code&gt;gpt-image-1&lt;/code&gt; to extend asset catalogs dynamically. Instead of manually commissioning variations, users can generate alternative backgrounds, localize visuals with region-specific elements, or adjust brand visuals for different markets, all from a single master file. With CE.SDK handling structured editing, teams maintain asset consistency while boosting creative flexibility.&lt;/p&gt;
&lt;h3 id=&quot;e-commerce&quot;&gt;E-Commerce&lt;/h3&gt;
&lt;p&gt;Product visualization remains a huge &lt;a href=&quot;https://img.ly/industries/e-commerce/&quot;&gt;challenge in e-commerce&lt;/a&gt;, especially for smaller retailers. &lt;code&gt;gpt-image-1&lt;/code&gt; can be used to automatically create product lifestyle imagery, context backgrounds, or thematic campaigns without expensive photo shoots. For example, a single shoe photograph can be placed into a generated “urban,” “sporty,” or “luxury” background, customized according to target audiences. When tightly integrated into e-commerce platforms, this enables faster product launches, A/B tested visuals, and localized campaigns at scale.&lt;/p&gt;
&lt;h3 id=&quot;e-learning&quot;&gt;E-Learning&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/industries/e-learning/&quot;&gt;Educational platforms&lt;/a&gt; can harness &lt;code&gt;gpt-image-1&lt;/code&gt; to generate explanatory diagrams, thematic illustrations, or scene-based visual storytelling assets. Instead of relying solely on static stock imagery, teachers, course designers, or even learners themselves can prompt the generation of custom visuals aligned with the curriculum. When embedded into authoring tools, this approach accelerates content creation and enables more engaging, visually enriched learning experiences tailored to specific topics and age groups.&lt;/p&gt;
&lt;h2 id=&quot;cost-optimization&quot;&gt;Cost Optimization&lt;/h2&gt;
&lt;p&gt;While &lt;code&gt;gpt-image-1&lt;/code&gt; opens up impressive creative possibilities, it also introduces new cost considerations that developers and product teams must plan for carefully. Since image generation typically incurs higher API costs than text-based operations, structuring workflows efficiently becomes critical, especially at scale.&lt;/p&gt;
&lt;h3 id=&quot;balancing-price-quality-and-resolution&quot;&gt;Balancing Price, Quality, and Resolution&lt;/h3&gt;
&lt;p&gt;The cost of generating an image with &lt;code&gt;gpt-image-1&lt;/code&gt; depends significantly on both the requested resolution and the selected quality setting. Higher resolutions like 4096×4096 produce sharper, more detailed results, but they also consume more compute resources-and therefore cost more. For many use cases, especially for previews, lower resolutions such as 1024×1024 or 2048×2048 strike an excellent balance between visual fidelity and API efficiency. Reserving the highest quality settings for final exports or premium workflows can help manage overall spend without compromising user experience.&lt;/p&gt;
&lt;h3 id=&quot;image-reuse-and-smart-upscaling&quot;&gt;Image Reuse and Smart Upscaling&lt;/h3&gt;
&lt;p&gt;One practical cost-saving approach is to design workflows that encourage image reuse. Instead of regenerating similar images for every small variation, applications can create high-quality master images and allow users to crop, edit, or layer additional design elements dynamically. Integrating smart upscaling techniques-for instance, using specialized image enhancement libraries after initial generation-also allows teams to work with smaller base images without sacrificing end-user quality.&lt;/p&gt;
&lt;h3 id=&quot;rate-limits-and-batching-strategies&quot;&gt;Rate Limits and Batching Strategies&lt;/h3&gt;
&lt;p&gt;Every call to &lt;code&gt;gpt-image-1&lt;/code&gt; counts toward your usage quota, and OpenAI imposes rate limits depending on account tier. To optimize performance and cost, it’s helpful to batch generation requests thoughtfully where possible-for instance, combining multiple prompts into structured queues or allowing users to preview low-res draft versions before finalizing a high-res render. Building this logic into your app’s generation flow not only controls expenses but also improves perceived responsiveness, an important UX factor for creative applications.&lt;/p&gt;
&lt;p&gt;By considering cost optimization as an early design constraint rather than a late-stage patch, developers can build scalable, sustainable creative tools powered by &lt;code&gt;gpt-image-1&lt;/code&gt;.&lt;/p&gt;
&lt;h2 id=&quot;bonus-starter-kit-repo&quot;&gt;Bonus: Starter Kit Repo&lt;/h2&gt;
&lt;p&gt;We are currently in the process of integrating the new GPT-4o-powered &lt;code&gt;gpt-image-1&lt;/code&gt; model into CE.SDK. As part of this effort, we are preparing a comprehensive Starter Kit will showcase a complete with CE.SDK integration, real-time prompt input, image generation workflows, and best practices for building an AI-powered creative editor.&lt;/p&gt;
&lt;p&gt;Both a public GitHub repository and a live demo will be made available soon. If you want to be notified when the Starter Kit launches, you can subscribe to updates &lt;a href=&quot;https://share.hsforms.com/1IgAOV1wASXGPnFG4ZPLejg1hk3i&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;This Starter Kit is designed to help developers move beyond simple image generation into building full creative cycles, where users can generate, edit, refine, and remix visuals seamlessly inside the editor.&lt;/p&gt;
&lt;h2 id=&quot;faqs&quot;&gt;FAQs&lt;/h2&gt;

&lt;p&gt;Choosing to work with &lt;code&gt;gpt-image-1&lt;/code&gt; raises a number of practical and strategic questions. Below, we address the most common topics for teams evaluating the model for integration into creative workflows.&lt;/p&gt;
&lt;h3 id=&quot;how-is-gpt-image-1-different-from-dalle-3&quot;&gt;How is &lt;code&gt;gpt-image-1&lt;/code&gt; different from DALL·E 3?&lt;/h3&gt;
&lt;p&gt;While DALL·E 3 and &lt;code&gt;gpt-image-1&lt;/code&gt; both translate text prompts into images, the underlying architecture and integration paths are different. &lt;code&gt;gpt-image-1&lt;/code&gt; is built on GPT-4o’s multimodal framework, making it better suited for future conversational and iterative workflows. It also offers support for a wider range of styles, higher resolutions up to 4096×4096 pixels, and is positioned for deeper integration into dynamic user experiences rather than one-off generation tasks.&lt;/p&gt;
&lt;h3 id=&quot;can-you-fine-tune-or-train-gpt-image-1&quot;&gt;Can you fine-tune or train &lt;code&gt;gpt-image-1&lt;/code&gt;?&lt;/h3&gt;
&lt;p&gt;As of April 2025, OpenAI does not allow fine-tuning of &lt;code&gt;gpt-image-1&lt;/code&gt;. The model is optimized for broad creative use cases out of the box. Developers seeking more control typically customize the user-facing prompt engineering or combine outputs with structured editing tools like CE.SDK to achieve brand or project-specific consistency.&lt;/p&gt;
&lt;h3 id=&quot;is-offline-support-available&quot;&gt;Is offline support available?&lt;/h3&gt;
&lt;p&gt;Currently, &lt;code&gt;gpt-image-1&lt;/code&gt; requires access to OpenAI’s cloud APIs. There is no offline inference mode or local deployment option. Teams requiring strict data residency, offline workflows, or private model hosting should consider hybrid architectures where images are generated securely via backend services and then edited locally using embedded tools like CE.SDK.&lt;/p&gt;
&lt;h3 id=&quot;what-about-copyright-and-licensing&quot;&gt;What about copyright and licensing?&lt;/h3&gt;
&lt;p&gt;Images generated by &lt;code&gt;gpt-image-1&lt;/code&gt; can be used commercially according to OpenAI’s &lt;a href=&quot;https://openai.com/en-GB/policies/usage-policies/&quot;&gt;usage policies&lt;/a&gt;, but developers are encouraged to review the latest terms. Outputs are not directly copyrighted by OpenAI or the user, and responsibility for ensuring compliance with branding, likeness, or content standards typically falls on the developer or platform operator. When deploying generation features to end-users, it is good practice to provide clear terms of use and, if needed, additional moderation or review layers.&lt;/p&gt;
&lt;p&gt;By addressing these considerations early, teams can integrate &lt;code&gt;gpt-image-1&lt;/code&gt; more effectively and responsibly into creative products and workflows.&lt;/p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;&lt;code&gt;gpt-image-1&lt;/code&gt; offers developers a significant opportunity to rethink what image generation can mean inside creative applications. It is not simply a tool for producing pictures on command, but a foundation for building interactive, iterative design workflows where users stay in control of the creative process. When combined with CE.SDK, it becomes even easier to move from static outputs to living, editable canvases that support real-world design needs. As we continue to integrate GPT-4o capabilities, the next wave of creative tooling will be about more than prompting images-it will be about shaping truly collaborative creative environments. Now is the time to start experimenting, iterating, and reimagining the user experience around this new generation of multimodal AI.&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="https://blog.img.ly/2025/04/GPT-Image-1-ultimate-guide.jpg" medium="image"/><category>AI</category><category>Creative Workflows</category><category>gpt-4o</category></item><item><title>How OpenAI&apos;s Upcoming GPT-4o Image Generation API Will Change Creative Workflows</title><link>https://img.ly/blog/open-ai-gpt-4o-image-generation-api-will-change-creative-workflows/</link><guid isPermaLink="true">https://img.ly/blog/open-ai-gpt-4o-image-generation-api-will-change-creative-workflows/</guid><description>OpenAI’s GPT-4o enables real-time, interactive image generation. Instead of one-off prompts, users can refine visuals through conversation. This unlocks new UX patterns like editable outputs and character persistence. IMG.LY’s CE.SDK makes GPT-4o easy to integrate into your editor.</description><pubDate>Mon, 14 Apr 2025 10:51:53 GMT</pubDate><content:encoded>&lt;p&gt;If you’ve been working with image-generation APIs over the past year, you’ve probably gotten used to a certain flow: send a prompt, wait a few seconds, and get a flat image back. It’s a one-shot deal. Useful? Definitely. But not exactly interactive. That’s what will change with OpenAI’s upcoming GPT-4o image-generation capabilities.&lt;br&gt;
IMG.LY, which recently &lt;a href=&quot;https://img.ly/use-cases/ai-editor/&quot;&gt;released a suite of AI features for its design editor&lt;/a&gt;, is eagerly awaiting the release to expand how users can interact with AI-driven creativity even further.&lt;/p&gt;
&lt;h2 id=&quot;update-ai-first-visual-editing&quot;&gt;Update: AI-first Visual Editing&lt;/h2&gt;
&lt;p&gt;A day after the release of the &lt;code&gt;gpt-image-1&lt;/code&gt; API, we put the UX principles outlined in this post into practice and integrated it into CreativeEditor SDK. Users can now generate images, create variants and use the canvas to compose visual prompts with our design editor. See it in action:&lt;/p&gt;
&lt;div class=&quot;cta-button-wrapper&quot;&gt;&lt;a href=&quot;https://cdn.img.ly/demo/gpt-image-1/v1/&quot; target=&quot;_blank&quot; class=&quot;cta-button&quot;&gt;Open AI Editor Demo Page&lt;/a&gt;&lt;/div&gt;
&lt;h2 id=&quot;gpt-4o-beyond-the-prompt-to-image-pipeline&quot;&gt;GPT-4o: Beyond the Prompt-to-Image Pipeline&lt;/h2&gt;
&lt;p&gt;GPT-4o isn’t just another version of DALL·E. It represents a shift in how developers will integrate AI into creative applications. While DALL·E 3 is powerful it is also somewhat siloed (you send a prompt, you get an image), GPT-4o looks like it will be part of a much more dynamic, conversational model one that accepts both text and image inputs, and could soon generate visual content in context, on the fly, and as part of a back-and-forth user interaction.&lt;/p&gt;
&lt;p&gt;If you’ve used ChatGPT recently, you’ve already seen glimpses of this. You can drop an image into the chat, ask GPT to describe or edit it, and get a response that feels fluid and visual. Developers should expect the API version to follow a similar pattern. It likely won’t just be a &lt;code&gt;/generate-image&lt;/code&gt; endpoint. Instead, we may be looking at an extension of the &lt;code&gt;chat/completions&lt;/code&gt; endpoint that handles multimodal messages. That changes the way you integrate this capability into your application. Rather than simply placing an image generation step in your pipeline, you will have to build your app’s UX around this new user flow. This comes with its own set of unique challenges.&lt;/p&gt;
&lt;h3 id=&quot;rethinking-the-interface-prompting-as-a-conversation&quot;&gt;Rethinking the Interface: Prompting as a Conversation&lt;/h3&gt;
&lt;p&gt;So what does this mean if you’re planning to integrate multi-modal image generation into your own product? For starters, you’ll probably need to rethink how users initiate and refine prompts. In the DALL·E flow, you might offer a text box with a few style dropdowns and call it a day. But in a GPT-4o world, your UI needs to support image inputs, persistent context, and dynamic editing, image gen becomes more like a conversation than a command.&lt;/p&gt;
&lt;p&gt;This is where the rubber meets the road. The tools that will benefit most from GPT-4o aren’t static generators but interactive editors. Think collaborative design apps, video editors with generative overlays, or product customizers that let users sketch or upload a photo and then iterate with AI. Put differently, the model output isn’t the endpoint but rather a checkpoint in the creation process.&lt;/p&gt;
&lt;h3 id=&quot;a-typical-iteration-cycle-in-a-multimodal-workflow&quot;&gt;A Typical Iteration Cycle in a Multimodal Workflow&lt;/h3&gt;
&lt;p&gt;Here’s a rough sketch of a workflow we might be seeing more of: The user starts with a prompt and an image, maybe a rough sketch or collage created inside an editor, a product photo, or a UI frame. GPT-4o returns a generated image based on that input. The user then edits or annotates the result, maybe adds new prompt text for refinement, and resubmits that combination to further develop the output. &lt;strong&gt;This cycle might loop several times&lt;/strong&gt;: generate, tweak, refine, regenerate.&lt;/p&gt;
&lt;p&gt;That’s a fundamentally different interaction model from past AI tooling. It’s less about one-off generation and more about a guided creative journey, where the user is in dialogue with the model. The result: better alignment with the original intent, more control, and more usable creative outputs.&lt;/p&gt;
&lt;p&gt;There is an additional, more subjective benefit to this kind of workflow: it gives the user a sense of autonomy again; they are back in the driver’s seat and less at the whim of an inscrutable machine. In many contexts, that makes a difference. Most notably, as we discussed in our white paper on &lt;a href=&quot;https://img.ly/white-papers/&quot;&gt;print personalization&lt;/a&gt;, the psychological benefit of personalization lies to a large extent in the investment, the sense of ownership that comes about when you create something. “Make it yours” is the common tagline attached to personalization campaigns in e-commerce. That only works if the user exerts more control over the output than iterating over a set of prompts.&lt;/p&gt;
&lt;p&gt;The most pithy encapsulation of this paradigm that I have heard is &lt;strong&gt;Humans on top, AI on tap&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id=&quot;persistent-elements-and-visual-consistency&quot;&gt;Persistent Elements and Visual Consistency&lt;/h3&gt;
&lt;p&gt;One particularly interesting frontier here is character and object persistence. If a user defines a character early in the workflow, either via prompt, image, or a combination, they’ll increasingly expect that character to appear consistently across assets. Think of it as visual continuity, whether you’re generating scenes in a story, slides in a deck, or frames in a video.&lt;/p&gt;
&lt;p&gt;If the user of a creative marketing cloud creates a campaign avatar or mascot, that &lt;strong&gt;character needs to be&lt;/strong&gt; &lt;strong&gt;consistent&lt;/strong&gt; within and across campaigns.&lt;/p&gt;
&lt;p&gt;Being able to reference earlier outputs, prompts, or style cues gives the user control over not just individual assets but the &lt;strong&gt;whole arc of the design narrative&lt;/strong&gt;. GPT-4o’s ability to maintain that continuity is a game-changer for workflows that involve storytelling, brand identity, or serialized design work.&lt;/p&gt;
&lt;h2 id=&quot;what-to-expect-from-the-api&quot;&gt;What to Expect from the API&lt;/h2&gt;
&lt;p&gt;Technically, if GPT-4o follows OpenAI’s recent design philosophy, you can expect a JSON-based API with a &lt;code&gt;messages&lt;/code&gt; array, where content can include both &lt;code&gt;text&lt;/code&gt; and &lt;code&gt;image_url&lt;/code&gt; types. The output will likely be returned either as an image URL hosted by OpenAI or as base64-encoded image data, depending on the format you request.&lt;/p&gt;
&lt;p&gt;That structure plays nicely with modern JavaScript front-end frameworks. React, Svelte, and Vue are all well-suited to async generation flows with visual previews. If you’re already using tools like &lt;a href=&quot;https://zustand.docs.pmnd.rs/&quot;&gt;Zustand&lt;/a&gt; or &lt;a href=&quot;https://jotai.org/&quot;&gt;Jotai&lt;/a&gt; for local state or something like &lt;a href=&quot;https://trpc.io/&quot;&gt;tRPC&lt;/a&gt; or &lt;a href=&quot;https://graphql.org/&quot;&gt;GraphQL&lt;/a&gt; for structured calls, you’re in a good position to layer GPT-4o in without breaking the flow.&lt;/p&gt;
&lt;h3 id=&quot;trade-offs-and-technical-considerations&quot;&gt;Trade-offs and Technical Considerations&lt;/h3&gt;
&lt;p&gt;There are trade-offs, of course. GPT-4o will probably cost more per call than a standard DALL·E 2 or 3 generation. Its latency is still an open question, and the multimodal input support will likely require more thoughtful UX decisions. What happens when a user drops an image and wants to undo just part of the generation? Where do you store prompt context for edits? How do you communicate what’s editable and what’s not?&lt;/p&gt;
&lt;p&gt;This is where design and engineering need to work together. You’ll want to build an interface that makes AI feel like a creative partner, not just a backend service. That might mean giving users a visual prompt history or allowing partial re-generations of specific canvas elements. You’ll need sensible fallback states. What happens when generation fails or the result isn’t what the user wanted?&lt;/p&gt;
&lt;h2 id=&quot;where-imglys-cesdk-fits-in&quot;&gt;Where IMG.LY’s CE.SDK Fits In&lt;/h2&gt;
&lt;p&gt;We have already given the questions raised above some serious thought, and most of the complexities introduced by this new workflow are the table stakes for the Creative Editor. So, if you’ve already integrated IMG.LY’s CE.SDK, we have taken care of most of these problems, and you can seamlessly integrate with any AI model. We are actively working on an off-the-shelf integration of the GPT-4o image model once its public API launches.&lt;/p&gt;
&lt;p&gt;In general, you can treat GPT-4o’s image outputs as just another layer in the editing canvas, positioned, styled, cropped, and ultimately editable in the same environment as everything else. That’s the real power of multimodal workflows: not just generating but integrating. And once GPT-4o’s API goes live, you’ll want your infrastructure ready to slot it in with minimal friction.&lt;/p&gt;
&lt;h3 id=&quot;the-loop-prompt-generate-refine&quot;&gt;The Loop: Prompt, Generate, Refine&lt;/h3&gt;
&lt;p&gt;The era of single-shot generation is winding down. What’s coming next is a loop: edit, prompt, generate, refine, repeat. And this loop doesn’t just belong in the backend, it needs to live in the UI, in a way that invites user input, creativity, and correction.&lt;/p&gt;
&lt;p&gt;We’ll be publishing more on how this integrates into IMG.LY’s upcoming AI workflows soon. Expect tools that don’t just generate visuals but help teams and individuals work through ideas in real time. Because especially as AI gets more potent, it needs humans on top.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;3,000+ creative professionals gain early access to new features and updates. Don’t miss out, and&lt;/strong&gt; &lt;a href=&quot;https://share.hsforms.com/1IgAOV1wASXGPnFG4ZPLejg1hk3i?ref=img.ly&quot;&gt;&lt;strong&gt;subscribe&lt;/strong&gt;&lt;/a&gt; &lt;strong&gt;to our newsletter.&lt;/strong&gt;&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="https://blog.img.ly/2025/04/GPT-4o-API-Changes-1.jpg" medium="image"/><category>AI</category><category>gpt-4o</category><category>Creative Workflows</category></item><item><title>Top 5 Generative AI APIs for Creative Apps in 2025: A Developer’s Guide (GPT-4o, Gemini, Firefly, and More)</title><link>https://img.ly/blog/comparing-image-generation-apis-gpt-4o-gemini-firefly-and-more/</link><guid isPermaLink="true">https://img.ly/blog/comparing-image-generation-apis-gpt-4o-gemini-firefly-and-more/</guid><description>AI image generation is quickly becoming a must-have in creative apps. But with so many models and APIs to choose from, picking the right one can be tricky. This guide breaks down the top options like DALL·E, GPT-4o, Gemini, and SDXL to help you find the best fit for your product and users.</description><pubDate>Mon, 14 Apr 2025 07:57:51 GMT</pubDate><content:encoded>&lt;p&gt;If you’re working on creative tooling right now, anything from a lightweight design editor to a marketing automation suite, you’re probably already thinking about or actively working on bringing image generation into the mix. The tech is here, expectations are rising, and if your users can’t type a prompt and get a visual back in seconds, your app might feel like it’s lagging behind.&lt;/p&gt;
&lt;p&gt;But choosing &lt;em&gt;which&lt;/em&gt; AI model to integrate, and how, isn’t all that straightforward. There’s a growing ecosystem of APIs out there, and they don’t all behave the same way, some are designed for open-ended creativity, others for structured workflows. Some offer pixel-perfect fidelity with fine control, others lean toward rapid ideation. And Importantly in our content not all of them are equally accessible to developers.&lt;/p&gt;
&lt;p&gt;This is a guide to help you make sense of it all. What models are available, how do they differ, and what should you consider when embedding them into your product. This isn’t supposed to be a hype piece or a leaderboard, just a clear-eyed look at what’s out there and what’s coming.&lt;/p&gt;
&lt;h3 id=&quot;openai-gpt-4o&quot;&gt;OpenAI GPT-4o&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://openai.com/index/introducing-4o-image-generation/&quot;&gt;GPT-4o is OpenAI&lt;/a&gt;’s next-gen multimodal model, currently only available inside ChatGPT. It can take both text and images as input and is capable of generating image outputs in context.&lt;/p&gt;
&lt;p&gt;The potential upside is significant. With GPT-4o, you may soon be able to create deeply interactive creative tools where users chat, sketch, and prompt all within a single UI. It’s likely to support richer input types and more natural iteration flows.&lt;/p&gt;
&lt;p&gt;The main downside is availability. There’s no API yet, so you can’t build on it directly. It also remains to be seen how OpenAI will expose generation tools, whether through a dedicated endpoint or via the chat interface.&lt;/p&gt;
&lt;p&gt;GPT-4o is right for you if you’re planning ahead and want to design for a future where multimodal interaction is the norm. It’s not something you can use today, but it should inform how you architect your UI and prompt handling.&lt;/p&gt;
&lt;h3 id=&quot;openai-dalle-3&quot;&gt;OpenAI DALL·E 3&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://openai.com/index/dall-e-3/&quot;&gt;DALL·E 3&lt;/a&gt; is OpenAI’s current image generation API, available via both the platform and ChatGPT. It translates text prompts into images and is known for interpreting prompts accurately and producing clean, useful visuals.&lt;/p&gt;
&lt;p&gt;Its strengths are clarity, commercial readiness, and reliability. It’s easy to use and integrates well into frontend flows that involve text-to-image generation.&lt;/p&gt;
&lt;p&gt;However, it lacks features like inpainting, style tuning, or detailed layout control. You also don’t get deep iteration features: each image is a new generation.&lt;/p&gt;
&lt;p&gt;DALL·E 3 is a good fit if you want &lt;strong&gt;high-quality results from text prompts&lt;/strong&gt; with minimal complexity. It’s especially useful for marketing visuals, content automation, and simple design tools.&lt;/p&gt;
&lt;h3 id=&quot;google-gemini-imagen&quot;&gt;Google Gemini (Imagen)&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://gemini.google.com/&quot;&gt;Gemini&lt;/a&gt;, powered by Google’s Imagen models, is available via &lt;a href=&quot;https://fal.ai/&quot;&gt;fal.ai&lt;/a&gt;, Makersuite, Vertex AI. It supports not only text prompts, but also sketches and inpainting, making it one of the more flexible APIs for creative work.&lt;/p&gt;
&lt;p&gt;Its big advantage is control. You can use sketches to guide composition and make visual edits to generated outputs. That makes it ideal for &lt;strong&gt;iterative design processes&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The downside is that it can be tricky to navigate Google’s ecosystem. Access and feature sets can change quickly, and the integration overhead is higher than OpenAI.&lt;/p&gt;
&lt;p&gt;Gemini is right for you if your product needs image refinement, visual grounding, or sketch-to-image workflows. It fits e-commerce editors, mockup tools, and design collaboration features.&lt;/p&gt;
&lt;h3 id=&quot;adobe-firefly&quot;&gt;Adobe Firefly&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://www.adobe.com/de/products/firefly.html&quot;&gt;Firefly is Adobe&lt;/a&gt;’s generative image model, integrated tightly into Creative Cloud. It stands out for its licensing model: images are trained on Adobe Stock, meaning they’re cleared for commercial use.&lt;/p&gt;
&lt;p&gt;The biggest strength here is trust and integration. Designers already using Photoshop or Illustrator can use Firefly to generate content directly in their layers and work non-destructively.&lt;/p&gt;
&lt;p&gt;The drawback is API access. There is no public endpoint for Firefly yet, and its features are embedded in Adobe’s own ecosystem.&lt;/p&gt;
&lt;p&gt;Firefly is a strong option if you’re building for agencies, brand teams, or other users with high expectations around copyright and integration with existing Adobe workflows.&lt;/p&gt;
&lt;h3 id=&quot;stability-ai-sdxl&quot;&gt;Stability AI (SDXL)&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://stability.ai/&quot;&gt;Stability AI&lt;/a&gt; offers an open-source model suite, with SDXL as the flagship for high-resolution image generation. It supports both text and image inputs and can be run locally or hosted via services like Replicate.&lt;/p&gt;
&lt;p&gt;Its biggest advantage is flexibility. You can fine-tune models, build custom workflows, or even run inference offline. It’s ideal for teams that want full control.&lt;/p&gt;
&lt;p&gt;The challenge is quality consistency. Compared to closed models like DALL·E, SDXL may require more tuning, and prompt engineering matters more. Hosting and scaling also require more effort.&lt;/p&gt;
&lt;p&gt;SDXL is right for you if you need an open, customizable system that fits into a broader pipeline. It’s a solid choice for research tools, OSS projects, and privacy-conscious applications.&lt;/p&gt;
&lt;h3 id=&quot;midjourney&quot;&gt;Midjourney&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://www.midjourney.com/home&quot;&gt;Midjourney&lt;/a&gt; is a proprietary model with a focus on aesthetic, stylized image generation. It runs exclusively via Discord and is popular for its distinctive look and community-driven prompts.&lt;/p&gt;
&lt;p&gt;Its upside is the quality of its visuals, especially for stylized scenes or concept art. Designers often use it as an ideation tool.&lt;/p&gt;
&lt;p&gt;The limitation is integration. There’s no API, no SDK, and limited ways to embed it in your own product beyond scraping or bots.&lt;/p&gt;
&lt;p&gt;Midjourney is best used as an inspiration engine. If your workflow includes moodboarding or creative brainstorming, it can supplement (but not power) your product.&lt;/p&gt;
&lt;h3 id=&quot;hugging-face&quot;&gt;Hugging Face&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://huggingface.co/&quot;&gt;Hugging Face&lt;/a&gt; is a hub for open models, offering hosted APIs for SDXL variants, Playground v2, and other creative generation tools.&lt;/p&gt;
&lt;p&gt;The main benefit is diversity. You can try multiple models, experiment with variations, and deploy quickly using their hosted inference endpoints.&lt;/p&gt;
&lt;p&gt;That said, it’s not always ready for production. Some models lack documentation or support, and you may need to piece together features.&lt;/p&gt;
&lt;p&gt;Hugging Face is a great choice for experimental projects, prototyping, or if you want to stay vendor-neutral and build your own stack.&lt;/p&gt;
&lt;h3 id=&quot;runway-gen-2-and-leonardoai&quot;&gt;Runway Gen-2 and Leonardo.Ai&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://runwayml.com/&quot;&gt;Runway&lt;/a&gt; and &lt;a href=&quot;https://leonardo.ai/&quot;&gt;Leonardo&lt;/a&gt; are rising players at the edge of AI and media. Runway’s Gen-2 supports text-to-video and animated image generation, while Leonardo focuses on style-consistent 2D asset generation.&lt;/p&gt;
&lt;p&gt;These platforms bring specialization. Runway is tailored to video and cinematic scenes, while Leonardo offers structured design features for asset creators.&lt;/p&gt;
&lt;p&gt;They’re less open from a dev perspective. APIs are limited, and integration support is still maturing.&lt;/p&gt;
&lt;p&gt;Use these tools if your use case leans into video, motion, or asset generation for games and content libraries. They’re best when you’re not looking to build your own editor, but to enhance creative capacity.&lt;/p&gt;
&lt;h3 id=&quot;quick-comparison&quot;&gt;Quick Comparison&lt;/h3&gt;





















































































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model/API&lt;/th&gt;&lt;th&gt;Input&lt;/th&gt;&lt;th&gt;Output&lt;/th&gt;&lt;th&gt;Control&lt;/th&gt;&lt;th&gt;API Access&lt;/th&gt;&lt;th&gt;Best For&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;GPT-4o (OpenAI)&lt;/td&gt;&lt;td&gt;text, image (chat)&lt;/td&gt;&lt;td&gt;image (likely)&lt;/td&gt;&lt;td&gt;medium-high&lt;/td&gt;&lt;td&gt;not yet&lt;/td&gt;&lt;td&gt;assistants, multimodal UIs&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;DALL·E 3&lt;/td&gt;&lt;td&gt;text&lt;/td&gt;&lt;td&gt;image&lt;/td&gt;&lt;td&gt;medium&lt;/td&gt;&lt;td&gt;yes&lt;/td&gt;&lt;td&gt;content tools, illustrations&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Gemini (Google)&lt;/td&gt;&lt;td&gt;text, sketch&lt;/td&gt;&lt;td&gt;image&lt;/td&gt;&lt;td&gt;high&lt;/td&gt;&lt;td&gt;yes&lt;/td&gt;&lt;td&gt;e-commerce, product mockups&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Firefly (Adobe)&lt;/td&gt;&lt;td&gt;text&lt;/td&gt;&lt;td&gt;image, layers&lt;/td&gt;&lt;td&gt;very high&lt;/td&gt;&lt;td&gt;no&lt;/td&gt;&lt;td&gt;professional design tools&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;SDXL&lt;/td&gt;&lt;td&gt;text, image&lt;/td&gt;&lt;td&gt;image&lt;/td&gt;&lt;td&gt;high&lt;/td&gt;&lt;td&gt;yes&lt;/td&gt;&lt;td&gt;custom tools, OSS projects&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Midjourney&lt;/td&gt;&lt;td&gt;text&lt;/td&gt;&lt;td&gt;image&lt;/td&gt;&lt;td&gt;very high&lt;/td&gt;&lt;td&gt;no&lt;/td&gt;&lt;td&gt;stylized inspiration&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Hugging Face&lt;/td&gt;&lt;td&gt;text, image&lt;/td&gt;&lt;td&gt;image&lt;/td&gt;&lt;td&gt;medium-high&lt;/td&gt;&lt;td&gt;yes&lt;/td&gt;&lt;td&gt;experimentation, open models&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Runway Gen-2&lt;/td&gt;&lt;td&gt;text&lt;/td&gt;&lt;td&gt;video/image&lt;/td&gt;&lt;td&gt;medium&lt;/td&gt;&lt;td&gt;yes&lt;/td&gt;&lt;td&gt;motion design, AI video&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Leonardo.Ai&lt;/td&gt;&lt;td&gt;text&lt;/td&gt;&lt;td&gt;image&lt;/td&gt;&lt;td&gt;high&lt;/td&gt;&lt;td&gt;limited&lt;/td&gt;&lt;td&gt;game assets, style templates&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;h3 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h3&gt;
&lt;p&gt;If you’re building for creative users, especially those used to real-time feedback and control, then how you wrap these APIs into your workflow matters more than which model you use. It’s not just about generating images. It’s about how you let users prompt, refine, iterate, and remix inside your canvas.&lt;/p&gt;
&lt;p&gt;That’s the opportunity here. Not just plugging in a model, but designing a loop where generation feels native to creation. The APIs are improving fast. The real challenge, and the real product value, is in how you build around them.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;3,000+ creative professionals gain early access to new features and updates. Don’t miss out, and &lt;a href=&quot;https://share.hsforms.com/1IgAOV1wASXGPnFG4ZPLejg1hk3i?ref=img.ly&quot;&gt;subscribe&lt;/a&gt; to our newsletter.&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="https://blog.img.ly/2025/04/AI-Comparison.jpg" medium="image"/><category>AI</category><category>Image Gen</category><category>Creative Workflows</category><category>gpt-4o</category></item><item><title>How To: Video Generation With Javascript</title><link>https://img.ly/blog/how-to-video-generation-with-javascript/</link><guid isPermaLink="true">https://img.ly/blog/how-to-video-generation-with-javascript/</guid><description>Learn how to programmatically create videos at scale with Javascript and CreativeEditor SDK.</description><pubDate>Wed, 22 Jan 2025 09:53:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;Follow this tutorial to learn how to programmatically create videos in JavaScript directly in the browser.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;In 2025 video is firmly established as powerhouse for engagement, gaining even more traction with the rise of short-form videos. After all, as humans, we naturally resonate with video content more than any other type of media.&lt;/p&gt;
&lt;p&gt;The problem is that video generation is challenging, and doing it within the browser through JavaScript makes things even more complex. Fortunately, production-grade solutions like &lt;a href=&quot;https://img.ly/docs/cesdk/js/starterkits/video-editor-e1nlor/&quot;&gt;CreativeEditor SDK (CE.SDK)&lt;/a&gt; make the process not only possible but also easy to implement.&lt;/p&gt;
&lt;p&gt;In this guide, you will learn how to programmatically generate videos in the browser via JavaScript using CreativeEditor SDK. We will cover various applications, including merging videos, adding audio tracks, and even integrating AI features.&lt;/p&gt;
&lt;p&gt;Let’s dive in!&lt;/p&gt;
&lt;h2 id=&quot;why-programmatic-video-generation-matters&quot;&gt;Why Programmatic Video Generation Matters&lt;/h2&gt;
&lt;p&gt;Videos are dominating social media, marketing, and online platforms such as e-commerce platforms (e.g., to showcase products). Numerous studies have proven that &lt;a href=&quot;https://thesocialshepherd.com/blog/video-marketing-statistics&quot;&gt;video is the most effective form of media&lt;/a&gt; for marketing, as it resonates deeply with us.&lt;/p&gt;
&lt;p&gt;As the demand for personalized and dynamic content increases, traditional video production methods are becoming inefficient. The proliferation of channels with different demands on size and video quality as well as the need to create personalized videos at scale means that automated video creation is becoming indispensable. Most existing solutions run batch processing of videos on the server after gathering input data from various source. However, while server-side processing quickly becomes costly and introduces significant communication overhead client devices have become more powerful than ever.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://img.ly/products/creative-sdk/&quot;&gt;CreativeEditor SDK&lt;/a&gt; bridges that gap by simplifying programmatic video generation in JavaScript within modern browsers. With CE.SDK, you can automate repetitive tasks and define scalable, personalized, and engaging video creation workflows, ensuring each generated video resonates with its target audience.&lt;/p&gt;
&lt;h2 id=&quot;get-started-with-cesdk-for-javascript&quot;&gt;Get Started with CE.SDK for JavaScript&lt;/h2&gt;
&lt;p&gt;Now that we established why programmatic video generation is important, you are ready to get started with the tool designed for it: CreativeEditor SDK.&lt;/p&gt;
&lt;p&gt;Learn how to integrate CE.SDK for video editing into a vanilla JavaScript application!&lt;/p&gt;
&lt;p&gt;You can access the code for this tutorial on &lt;a href=&quot;https://github.com/imgly/video-generation-js&quot;&gt;Github&lt;/a&gt;.&lt;/p&gt;
&lt;h3 id=&quot;requirements&quot;&gt;Requirements&lt;/h3&gt;
&lt;p&gt;The only prerequisites to follow this tutorial are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;A modern browser&lt;/strong&gt;: CreativeEditor SDK runs directly in your browser’s JavaScript engine, so no additional setup is required. Any &lt;a href=&quot;https://caniuse.com/wasm&quot;&gt;browser supporting WASM&lt;/a&gt; is enough.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A CreativeEditor SDK license&lt;/strong&gt;: If you do not have one yet, &lt;a href=&quot;https://img.ly/docs/cesdk/&quot;&gt;sign up for a free trial&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If your web application uses Node.js, ensure you have the &lt;a href=&quot;https://nodejs.org/en/download&quot;&gt;latest stable versions of both Node.js and NPM installed&lt;/a&gt;.&lt;/p&gt;
&lt;h3 id=&quot;import-the-library&quot;&gt;Import the library&lt;/h3&gt;
&lt;p&gt;In a vanilla JavaScript application, create a JavaScript module file (e.g., video-editor.js) and import the CreativeEditor SDK engine library inside it:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;jsx&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; CreativeEngine &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;&amp;#x3C;https://cdn.img.ly/packages/imgly/cesdk-engine/1.42.0/index.js&gt;&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: In this example, the SDK is served from our CDN for convenience. In a production environment, it is recommended to host all &lt;a href=&quot;https://img.ly/docs/cesdk/js/serve-assets-b0827c/&quot;&gt;assets and libraries on your own servers&lt;/a&gt; for improved control and performance.&lt;/p&gt;
&lt;p&gt;To use the video-editor.js module in an HTML page, import it with this line:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;html&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;#x3C;&lt;/span&gt;&lt;span&gt;script&lt;/span&gt;&lt;span&gt; type&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&quot;module&quot;&lt;/span&gt;&lt;span&gt; src&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&quot;video-editor.js&quot;&lt;/span&gt;&lt;span&gt;&gt;&amp;#x3C;/&lt;/span&gt;&lt;span&gt;script&lt;/span&gt;&lt;span&gt;&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Note the &lt;a href=&quot;https://developer.mozilla.org/en-US/docs/Web/JavaScript/Guide/Modules&quot;&gt;type=“module”&lt;/a&gt; attribute tells the browser to treat the script as an ES6 module, so that you can use import and export statements.&lt;/p&gt;
&lt;p&gt;Alternatively, if you are using a module bundler like Webpack, Rollup, Parcel, or Vite, add the &lt;a href=&quot;https://www.npmjs.com/package/@cesdk/engine&quot;&gt;@cesdk/engine&lt;/a&gt; npm package to your project’s dependencies:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;npm&lt;/span&gt;&lt;span&gt; install&lt;/span&gt;&lt;span&gt; @cesdk/cesdk-js&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;In this case, you can then import the headless SDK into your JavaScript code as follows:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;jsx&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; CreativeEngine &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;@cesdk/engine&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Well done!&lt;/p&gt;
&lt;h2 id=&quot;setting-up-your-environment&quot;&gt;&lt;strong&gt;Setting Up Your Environment&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;In your JavaScript module, right below the library import, initialize the &lt;a href=&quot;https://img.ly/docs/cesdk/node/get-started/overview-e18f40/&quot;&gt;CreativeEditor SDK headless engine&lt;/a&gt; using the following code:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;jsx&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// src/video-editor.js&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; CreativeEngine &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;&amp;#x3C;https://cdn.img.ly/packages/imgly/cesdk-engine/1.42.0/index.js&gt;&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// your CE.SDK license and user configs&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; config&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  license: &lt;/span&gt;&lt;span&gt;&apos;&amp;#x3C;YOUR_LICENSE&gt;&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;};&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;CreativeEngine.&lt;/span&gt;&lt;span&gt;init&lt;/span&gt;&lt;span&gt;(config).&lt;/span&gt;&lt;span&gt;then&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;async&lt;/span&gt;&lt;span&gt; (&lt;/span&gt;&lt;span&gt;engine&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;=&gt;&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  // attach the engine canvas to the DOM&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  document.&lt;/span&gt;&lt;span&gt;getElementById&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;cesdk_container&apos;&lt;/span&gt;&lt;span&gt;).&lt;/span&gt;&lt;span&gt;append&lt;/span&gt;&lt;span&gt;(engine.element);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  // do something with your instance of CreativeEditor SDK...&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  // detach the engine and clean up its resource&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  engine.element.&lt;/span&gt;&lt;span&gt;remove&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  engine.&lt;/span&gt;&lt;span&gt;dispose&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;});&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Replace &lt;code&gt;&amp;#x3C;YOUR_LICENSE&gt;&lt;/code&gt; with the license key provided in your CreativeEditor SDK subscription.&lt;/p&gt;
&lt;p&gt;Ensure that the HTML page importing &lt;code&gt;video-editor.js&lt;/code&gt; contains the following &lt;code&gt;&amp;#x3C;div&gt;&lt;/code&gt; element with the &lt;code&gt;cesdk_container&lt;/code&gt; ID:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;html&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;#x3C;&lt;/span&gt;&lt;span&gt;div&lt;/span&gt;&lt;span&gt; id&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&quot;cesdk_container&quot;&lt;/span&gt;&lt;span&gt; style&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&quot;width: 100%; height: 100vh;&quot;&lt;/span&gt;&lt;span&gt;&gt;&amp;#x3C;/&lt;/span&gt;&lt;span&gt;div&lt;/span&gt;&lt;span&gt;&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The &lt;a href=&quot;https://img.ly/docs/cesdk/js/get-started/overview-e18f40/&quot;&gt;&lt;code&gt;CreativeEngine.init()&lt;/code&gt;&lt;/a&gt; method will mount the editor engine component within this div.&lt;/p&gt;
&lt;p&gt;Congratulations! You have successfully integrated the video editing engine.&lt;/p&gt;
&lt;h2 id=&quot;load-a-video&quot;&gt;&lt;strong&gt;Load a Video&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Use the logic below inside the body of the CreativeEndine.init() callback function to load a video in CE.SDK:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;jsx&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// initialize a new video scene&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; scene&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.scene.&lt;/span&gt;&lt;span&gt;createVideo&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// create a page block and attach it to the scene&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; page&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;create&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;page&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;appendChild&lt;/span&gt;&lt;span&gt;(scene, page);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// set the dimensions of the video page&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setWidth&lt;/span&gt;&lt;span&gt;(page, &lt;/span&gt;&lt;span&gt;720&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setHeight&lt;/span&gt;&lt;span&gt;(page, &lt;/span&gt;&lt;span&gt;1280&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// create a graphic block to hold the video content&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// and set its shape to a rectangle&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; video&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;create&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;graphic&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setShape&lt;/span&gt;&lt;span&gt;(video, engine.block.&lt;/span&gt;&lt;span&gt;createShape&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;rect&apos;&lt;/span&gt;&lt;span&gt;));&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// create a fill type for the video and set the video file URI&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; videoFill&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;createFill&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;video&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setString&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  videoFill,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &apos;fill/video/fileURI&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &apos;&amp;#x3C;https://cdn.img.ly/assets/demo/v2/ly.img.video/videos/pexels-drone-footage-of-a-surfer-barrelling-a-wave-12715991.mp4&gt;&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// apply the video fill to the video graphic&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setFill&lt;/span&gt;&lt;span&gt;(video, videoFill);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// create a track block to manage the video timeline&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// and attache the track to the page and the video graphic&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// to the track&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; track&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;create&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;track&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;appendChild&lt;/span&gt;&lt;span&gt;(page, track);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;appendChild&lt;/span&gt;&lt;span&gt;(track, video);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// make the track cover the parent block&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;fillParent&lt;/span&gt;&lt;span&gt;(track);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This initializes a video scene using &lt;a href=&quot;https://img.ly/docs/cesdk/js/concepts/scenes-e8596d/&quot;&gt;&lt;code&gt;createVideo()&lt;/code&gt;&lt;/a&gt; and creates a &lt;code&gt;&quot;graphic&quot;&lt;/code&gt; block to display the video in a rectangle, adding it to the page. Next, it loads a video from a remote resource and appends it as a track to the parent block.&lt;/p&gt;
&lt;p&gt;For a complete explanation of how the above code works, &lt;a href=&quot;https://img.ly/docs/cesdk/js/create-video-c41a08/&quot;&gt;refer to the documentation&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;If you comment out the cleanup part, this is what you should see in your browser application:&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 608px) 608px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;608&quot; height=&quot;697&quot; src=&quot;https://img.ly/_astro/JS-video-generation-video-fill_Z2oN5x6.webp&quot; srcset=&quot;/_astro/JS-video-generation-video-fill_Z2oN5x6.webp 608w&quot;&gt;&lt;/p&gt;
&lt;p&gt;Notice how the video from the remote URI has been loaded into the &lt;/p&gt;&lt;div&gt; element where CreativeEditor SDK is attached to the page’s DOM.Time to explore more programmatic video generation features!&lt;p&gt;&lt;/p&gt;
&lt;h2 id=&quot;programmatic-video-generation-features&quot;&gt;Programmatic Video Generation Features&lt;/h2&gt;
&lt;p&gt;Follow the use cases below to see how CE.SDK makes automated video generation in JavaScript easier.&lt;strong&gt;Generate Video in Different Formats&lt;/strong&gt; CE.SDK supports video export in &lt;a href=&quot;https://img.ly/docs/cesdk/js/file-format-support-3c4b2a/&quot;&gt;MP4 format&lt;/a&gt; &lt;a href=&quot;https://img.ly/docs/cesdk/faq/video-support/?platform=node&quot;&gt;&lt;/a&gt;(with MP3 audio) with different resolutions and aspect ratios.To export a video with a specific aspect ratio, first adjust the width and height of the parent block to which the video block is attached:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// specify an aspect ratio of 9:16 (ideal for vertical videos)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setWidth&lt;/span&gt;&lt;span&gt;(page, &lt;/span&gt;&lt;span&gt;720&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setHeight&lt;/span&gt;&lt;span&gt;(page, &lt;/span&gt;&lt;span&gt;1280&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;In this case, the video will be automatically adjusted to fit the specified dimensions. Horizontal videos will be cropped to match the defined vertical outline of the block, while vertical videos will be rendered as they are within the specified block.&lt;/p&gt;
&lt;p&gt;Next, you can export the video with the configured 9:16 aspect ratio and a given resolution with:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// define the MIME type for the video export (MP4 format)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; mimeType&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; &apos;video/mp4&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// define a callback function to track the progress of video rendering and encoding&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; progressCallback&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; (&lt;/span&gt;&lt;span&gt;renderedFrames&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;encodedFrames&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;totalFrames&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;=&gt;&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  console.&lt;/span&gt;&lt;span&gt;log&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    &apos;Rendered&apos;&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;// log the number of rendered frames&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    renderedFrames,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    &apos;frames and encoded&apos;&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;// log the number of encoded frames&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    encodedFrames,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    &apos;frames out of&apos;&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;// log the total frames to be processed&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    totalFrames&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  );&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;};&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// define the video export options&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; videoOptions&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  duration: &lt;/span&gt;&lt;span&gt;5&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;// video duration in seconds&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  framerate: &lt;/span&gt;&lt;span&gt;30&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;// video framerate (frames per second)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  targetWidth: &lt;/span&gt;&lt;span&gt;480&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;// target video width in pixels (for the resolution)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  targetHeight: &lt;/span&gt;&lt;span&gt;852&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;// target video height in pixels (for the resolution)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;};&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// export the page as an MP4 video&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; blob&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; await&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;exportVideo&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  page, &lt;/span&gt;&lt;span&gt;// the page containing the design video block&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  mimeType, &lt;/span&gt;&lt;span&gt;// the MIME type for the video (MP4)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  progressCallback, &lt;/span&gt;&lt;span&gt;// the callback to track video rendering progress&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  videoOptions &lt;/span&gt;&lt;span&gt;// the options for the video export&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// create an anchor element to trigger the video download&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; anchor&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; document.&lt;/span&gt;&lt;span&gt;createElement&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;a&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;anchor.href &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; URL&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;createObjectURL&lt;/span&gt;&lt;span&gt;(blob);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;anchor.download &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; &apos;video.mp4&apos;&lt;/span&gt;&lt;span&gt;; &lt;/span&gt;&lt;span&gt;// output video name&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;anchor.&lt;/span&gt;&lt;span&gt;click&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The &lt;a href=&quot;https://img.ly/docs/cesdk/js/export-save-publish/export/overview-9ed3a8/&quot;&gt;&lt;code&gt;exportVideo()&lt;/code&gt;&lt;/a&gt; function produces a &lt;a href=&quot;https://developer.mozilla.org/en-US/docs/Web/API/Blob&quot;&gt;&lt;code&gt;blob video file&lt;/code&gt;&lt;/a&gt; of the specified MIME type. Note that the export occurs over multiple iterations of the update loop, with a frame being encoded in each iteration. &lt;/p&gt;
&lt;p&gt;The &lt;code&gt;targetWidth&lt;/code&gt; and &lt;code&gt;targetHeight&lt;/code&gt; options are optional. If used, the video will be resized to fit the target dimensions while maintaining its aspect ratio. This is useful for setting a specific resolution for the output video while preserving the parent block’s width and height on the output video. If omitted, the produced video will match the resolution of the width and height specified with &lt;code&gt;setWidth()&lt;/code&gt; and &lt;code&gt;setHeight()&lt;/code&gt;, respectively.&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;progressCallback()&lt;/code&gt; function will track and display the video generation progress in the browser’s console:&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1066px) 1066px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1066&quot; height=&quot;295&quot; src=&quot;https://img.ly/_astro/video-generation-output_Z2wCvPW.webp&quot; srcset=&quot;/_astro/video-generation-output_Z1WN1Sx.webp 640w, /_astro/video-generation-output_Z16737N.webp 750w, /_astro/video-generation-output_Z2wIbfk.webp 828w, /_astro/video-generation-output_Z2wCvPW.webp 1066w&quot;&gt;&lt;/p&gt;
&lt;p&gt;The final lines of the above code snippet trigger the download operation, mimicking the action of clicking a link to download a resource. This is a common JavaScript technique to programmatically trigger a file download.&lt;/p&gt;
&lt;p&gt;Once the export process completes, the output video file (video.mp4) will be automatically downloaded to the browser’s download directory.  Open its properties and this is what you will see:&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 404px) 404px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;404&quot; height=&quot;512&quot; src=&quot;https://img.ly/_astro/js-video-generation-file_2s55xB.webp&quot; srcset=&quot;/_astro/js-video-generation-file_2s55xB.webp 404w&quot;&gt;&lt;/p&gt;
&lt;p&gt;Note the video file’s resolution and the fact that its duration is set to 5 seconds, as defined by the &lt;code&gt;duration&lt;/code&gt; attribute.&lt;/p&gt;
&lt;p&gt;These options enable you to easily create clips with different lengths, aspect ratios, and resolutions, optimized for various audiences and &lt;a href=&quot;https://img.ly/industries/social-media/&quot;&gt;social media platforms like LinkedIn, Instagram, Facebook, X, TikTok, YouTube, and more&lt;/a&gt;.&lt;/p&gt;
&lt;h3 id=&quot;add-audio-tracks&quot;&gt;Add Audio Tracks&lt;/h3&gt;
&lt;p&gt;Adding an audio track for narration or background music to your clip is straightforward. First, create an &lt;a href=&quot;https://img.ly/docs/cesdk/js/create-video-c41a08/&quot;&gt;“audio” block&lt;/a&gt; and add your audio resource to it:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; audio&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;create&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;audio&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;appendChild&lt;/span&gt;&lt;span&gt;(page, audio);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setString&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  audio,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &apos;audio/fileURI&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &apos;https://cdn.img.ly/assets/demo/v1/ly.img.audio/audios/far_from_home.m4a&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Next, you can easily adjust the volume, and apply fade-in and fade-out effects as follows:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// set the volume level to 80%&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setVolume&lt;/span&gt;&lt;span&gt;(audio, &lt;/span&gt;&lt;span&gt;0.8&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// start the audio after 1 second of playback&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setTimeOffset&lt;/span&gt;&lt;span&gt;(audio, &lt;/span&gt;&lt;span&gt;1&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// set the audio block&apos;s duration to 5 seconds&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setDuration&lt;/span&gt;&lt;span&gt;(audio, &lt;/span&gt;&lt;span&gt;5&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Here is breakdown of the functions used above:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://img.ly/docs/cesdk/js/create-video/control-daba54/&quot;&gt;&lt;strong&gt;&lt;code&gt;setVolume()&lt;/code&gt;&lt;/strong&gt;&lt;/a&gt;: Adjusts the audio volume of the block, with a range from 0 (0%) to 1 (100%).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;setTimeOffset()&lt;/code&gt;&lt;/strong&gt;: Sets the time offset of the block relative to its parent. This determines when the block starts playing in the timeline.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;setDuration()&lt;/code&gt;&lt;/strong&gt;: Sets the playback duration of the audio block in seconds, defining how long the block will play during the scene.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In this example, the audio starts at the 1-second mark and plays for 5 seconds (from 1s to 6s), with the volume set to 80% of the original.&lt;/p&gt;
&lt;h3 id=&quot;merge-videos&quot;&gt;&lt;strong&gt;Merge Videos&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;To stitch multiple video clips together in CE.SDK, you first need to import and create separate video blocks:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// create the first video block&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; video1&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;create&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;graphic&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setShape&lt;/span&gt;&lt;span&gt;(video1, engine.block.&lt;/span&gt;&lt;span&gt;createShape&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;rect&apos;&lt;/span&gt;&lt;span&gt;));&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; videoFill1&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;createFill&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;video&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setString&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  videoFill1,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &apos;fill/video/fileURI&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &apos;https://cdn.img.ly/assets/demo/v2/ly.img.video/videos/pexels-drone-footage-of-a-surfer-barrelling-a-wave-12715991.mp4&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setFill&lt;/span&gt;&lt;span&gt;(video1, videoFill1);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// create the second video block&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; video2&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;create&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;graphic&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setShape&lt;/span&gt;&lt;span&gt;(video2, engine.block.&lt;/span&gt;&lt;span&gt;createShape&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;rect&apos;&lt;/span&gt;&lt;span&gt;));&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; videoFill2&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; engine.block.&lt;/span&gt;&lt;span&gt;createFill&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&apos;video&apos;&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setString&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  videoFill2,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &apos;fill/video/fileURI&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &apos;https://cdn.img.ly/assets/demo/v2/ly.img.video/videos/pexels-kampus-production-8154913.mp4&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setFill&lt;/span&gt;&lt;span&gt;(video2, videoFill2);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;At this point, you have two different video clips loaded from separate video files.&lt;/p&gt;
&lt;p&gt;Now, assume you want to produce a 10-second video that includes both clips. Start by setting the duration of the page block to 10 seconds with the &lt;a href=&quot;https://img.ly/docs/cesdk/js/create-video/control-daba54/&quot;&gt;&lt;code&gt;setDuration()&lt;/code&gt;&lt;/a&gt; method:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.&lt;/span&gt;&lt;span&gt;setDuration&lt;/span&gt;&lt;span&gt;(page, &lt;/span&gt;&lt;span&gt;10&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Then, suppose you want the first clip (video1) to play for 7 seconds, and the second clip (video2) to fade in after that and play for 3 seconds (from second 1 to second 4 of the original second clip). This is how you can set that up:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// set the duration of the first and second video clips&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.setDuration(video1, 7); // video1 plays for 7 seconds&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.setDuration(video2, 3); // video2 plays for 3 seconds&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// wait for the second video to be fully loaded before applying the trim options&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;await engine.block.forceLoadAVResource(videoFill2);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// trim the second video to start at 1 second and play up to second 4&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.setTrimOffset(videoFill2, 1);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.setTrimLength(videoFill2, 4);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// add a fade-in animation to the second video&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const fadeInAnimation = engine.block.createAnimation(&quot;fade&quot;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.setInAnimation(video2, fadeInAnimation);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;These are the methods used above:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;setDuration()&lt;/code&gt;&lt;/strong&gt;: Defines the duration (in seconds) that a video block will be active during playback.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;**forceLoadAVResource()**&lt;/code&gt;: Ensures that the audio or video resource (in this case, &lt;code&gt;video2&lt;/code&gt;) is fully loaded before applying trim operations.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;setTrimOffset()&lt;/code&gt;&lt;/strong&gt;: Sets the start point (in seconds) within the video where playback begins.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;setTrimLength()&lt;/code&gt;&lt;/strong&gt;: Defines the length of the video to be played from the trim offset.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://img.ly/docs/cesdk/js/animation/overview-6a2ef2/&quot;&gt;&lt;strong&gt;&lt;code&gt;createAnimation()&lt;/code&gt;&lt;/strong&gt;&lt;/a&gt;: Creates a fade-in animation for the second video (&lt;code&gt;video2&lt;/code&gt;) to smoothly transition into the clip.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Finally, add both video blocks to the track block for playback:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const track = engine.block.create(&quot;track&quot;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.appendChild(page, track);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.appendChild(track, video1); // add the first video to the track&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.appendChild(track, video2); // add the second video to the track&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.fillParent(track);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The result will be:&lt;/p&gt;
&lt;figure class=&quot;kg-card kg-embed-card&quot;&gt;&lt;iframe width=&quot;225&quot; height=&quot;400&quot; src=&quot;https://blog.img.ly/2025/01/video-1.mp4&quot; frameborder=&quot;0&quot; allowfullscreen&gt;&lt;/iframe&gt;&lt;/figure&gt;
&lt;p&gt;Note how the second clip fades in after 7 seconds of the first clip, lasting for 3 seconds as specified.&lt;/p&gt;
&lt;h3 id=&quot;incorporate-text-variables&quot;&gt;&lt;a href=&quot;https://img.ly/blog/how-to-video-generation-with-javascript//#Incorporate-Text-Variables&quot;&gt;&lt;/a&gt;Incorporate Text Variables&lt;/h3&gt;
&lt;p&gt;Now, suppose you want to display cool words or messages in your videos. You can easily achieve this by adding a &lt;a href=&quot;https://img.ly/docs/cesdk/js/text/edit-c5106b/&quot;&gt;&lt;code&gt;&quot;text&quot;&lt;/code&gt; block&lt;/a&gt; to your page, as shown below:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const text = engine.block.create(&quot;text&quot;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// set the vluae of the text variable&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.replaceText(text, &quot;Surfing\nis\nCOOL!&quot;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// set the text color to white and semi-transparent&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.setTextColor(text, { r: 255.0, g: 255.0, b: 255.0, a: 0.8 });&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// set the font size for the text&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.setTextFontSize(text, 30);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// positioning the text on an absolute position on the video&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.setWidthMode(text, &quot;Auto&quot;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.setHeightMode(text, &quot;Auto&quot;);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.setPositionX(text, 130);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.setPositionY(text, 800);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;// append the text block to the page&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;engine.block.appendChild(page, text);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The video canvas will now display a white “Surfing is COOL!” message as follows:&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 514px) 514px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;514&quot; height=&quot;812&quot; src=&quot;https://img.ly/_astro/js-video-generation-surfing_1WpaM3.webp&quot; srcset=&quot;/_astro/js-video-generation-surfing_1WpaM3.webp 514w&quot;&gt;&lt;/p&gt;
&lt;p&gt;In this example, the content of the text block is static. However, it can easily be retrieved from a database or any other source for programmatic video generation with different messages, including personalized ones for tailored outreach (e.g., for custom greetings or offers).&lt;/p&gt;
&lt;h2 id=&quot;advanced-use-case-prompt-based-video-generation-using-ai&quot;&gt;&lt;a href=&quot;https://img.ly/blog/how-to-video-generation-with-javascript//#Advanced-Use-Case-Prompt-Based-Video-Generation-Using-AI&quot;&gt;&lt;/a&gt;Advanced Use Case: Prompt-Based Video Generation Using AI&lt;/h2&gt;
&lt;p&gt;As highlighted in our &lt;a href=&quot;https://www.linkedin.com/posts/eray-basar-57684711_imagine-building-a-full-video-app-in-a-single-activity-7262806525797089280-S_L8?utm_source=share&amp;#x26;utm_medium=member_desktop&quot;&gt;recent LinkedIn post&lt;/a&gt;, &lt;a href=&quot;https://IMG.LY&quot;&gt;IMG.LY&lt;/a&gt; has just experimented with an MVP that turns keywords into fully edited short videos, complete with a script, voiceover, and visuals, all automated.&lt;/p&gt;
&lt;p&gt;That is made possible by integrating CE.SDK with an LLM for coding and scriptwriting, &lt;a href=&quot;https://elevenlabs.io/&quot;&gt;ElevenLabs&lt;/a&gt; for voice synthesis, and Flux (via &lt;a href=&quot;https://fal.ai/&quot;&gt;fal.ai&lt;/a&gt;) for visuals. Specifically, the LLM uses CE.SDK for video composition, animation, and rendering.&lt;/p&gt;
&lt;p&gt;This is just the beginning, showing how AI can be combined with CE.SDK to programmatically generate video scripts, or even templates. A possible scenario would be to use the &lt;a href=&quot;https://img.ly/docs/cesdk/js/user-interface/ui-extensions-d194d1/&quot;&gt;Creative Editor SDK plugin API&lt;/a&gt; to extend CE.SDK with custom plugins that adds AI directly to the design editor engine.&lt;/p&gt;
&lt;p&gt;For example, you could develop a plugin with an LLM integration that allows users to input a prompt like: &lt;em&gt;“Create a 15-second video with the text ‘Welcome to our App!’ and upbeat background music.”&lt;/em&gt; The LLM will process the request, using Creative Editor SDK to automate video creation by leveraging its powerful features.&lt;/p&gt;
&lt;p&gt;Alternatively, the AI could produce a video that you can then automatically edit and enhance using CE.SDK capabilities. One thing is certain: the possibilities are endless with CE.SDK + AI!&lt;/p&gt;
&lt;h2 id=&quot;conclusion-the-future-of-programmatic-video-generation&quot;&gt;&lt;a href=&quot;https://img.ly/blog/how-to-video-generation-with-javascript//#Conclusion-The-Future-of-Programmatic-Video-Generation&quot;&gt;&lt;/a&gt;Conclusion: The Future of Programmatic Video Generation&lt;/h2&gt;
&lt;p&gt;CE.SDK for JavaScript offers exceptional versatility in automating video workflows, allowing you to easily generate, edit, and render videos with just a few lines of code. You can stitch multiple video files, add background audio tracks, and incorporate dynamic text variables to finally produce videos in the browser, all thanks to the CE.SDK headless engine.&lt;/p&gt;
&lt;p&gt;Looking ahead, AI integration with CE.SDK is unlocking even more powerful possibilities, such as prompt-based video content creation and automatic editing directly in the browser.&lt;/p&gt;
&lt;p&gt;With the headless engine provided by CreativeEditor SDK, programmatic video generation in JavaScript can be implemented in just a few minutes. Explore the video capabilities of CE.SDK and &lt;a href=&quot;https://img.ly/docs/cesdk/js/get-started/overview-e18f40/&quot;&gt;dive into the docs&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Stay tuned for more updates, and please &lt;a href=&quot;https://img.ly/forms/contact-sales/&quot;&gt;reach out&lt;/a&gt; if you have any questions. Thank you for reading.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Over 3,000 creative professionals gain early access to our new features, demos, and updates. Don’t miss out, and&lt;/strong&gt; &lt;a href=&quot;https://share.hsforms.com/1IgAOV1wASXGPnFG4ZPLejg1hk3i?ref=img.ly&quot;&gt;&lt;strong&gt;subscribe&lt;/strong&gt;&lt;/a&gt; &lt;strong&gt;to our newsletter.&lt;/strong&gt;&lt;/p&gt;&lt;/div&gt;</content:encoded><dc:creator>Antonello</dc:creator><dc:creator>Jan</dc:creator><media:content url="https://blog.img.ly/2025/01/javascript-generate-video.jpg" medium="image"/><category>Video Editing</category><category>Creative Automation</category><category>CE.SDK</category></item><item><title>A Modern Video Editor SDK for Your React Native App</title><link>https://img.ly/blog/a-modern-video-editor-sdk-for-your-react-native-app/</link><guid isPermaLink="true">https://img.ly/blog/a-modern-video-editor-sdk-for-your-react-native-app/</guid><description>Learn how to integrate IMG.LY&apos;s video editor for React Native into your app and how to best leverage its capabilities.</description><pubDate>Mon, 06 Jan 2025 10:40:35 GMT</pubDate><content:encoded>&lt;p&gt;Learn how to integrate &lt;a href=&quot;https://img.ly/docs/cesdk/react-native/prebuilt-solutions/video-editor-9e533a/&quot;&gt;IMG.LY’s video editor for React Native&lt;/a&gt; into your app and how to leverage its capabilities best.&lt;/p&gt;
&lt;h2 id=&quot;why-add-a-video-editor-to-your-react-native-app&quot;&gt;Why Add a Video Editor to Your React Native App?&lt;/h2&gt;
&lt;p&gt;Video content is solidifying its status as the most engaging form of digital media. Platforms like TikTok, Instagram Reels, and YouTube Shorts have made video consumption and creation an engrained habit with billions of users. You can harness this trend by adding a video editor to your app significantly enhancing user engagement and retention.&lt;/p&gt;
&lt;p&gt;React Native, with its ability to create cross-platform applications from a single codebase, is a perfect match for IMG.LY’s &lt;strong&gt;CreativeEditor SDK Video Editor&lt;/strong&gt;. It ensures seamless performance on iOS and Android, powered by the same unified graphics engine across platforms.&lt;/p&gt;
&lt;p&gt;Whether your app focuses on social media, marketing, or eCommerce, integrating a video editor empowers users with a creative tool set while elevating the overall experience.&lt;/p&gt;
&lt;h2 id=&quot;getting-started-integrating-the-video-editor-in-react-native&quot;&gt;Getting Started: Integrating the Video Editor in React Native&lt;/h2&gt;
&lt;h3 id=&quot;requirements&quot;&gt;Requirements&lt;/h3&gt;
&lt;p&gt;Before diving into the integration, ensure your environment meets these requirements:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;React Native&lt;/strong&gt;: 0.73+&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;iOS&lt;/strong&gt;: 16+&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Swift&lt;/strong&gt;: 5.10 (Xcode 15.4)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Android&lt;/strong&gt;: 7+&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;To get started, add the &lt;strong&gt;@imgly/editor-react-native&lt;/strong&gt; package to your project:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;npm&lt;/span&gt;&lt;span&gt; install&lt;/span&gt;&lt;span&gt; @imgly/editor-react-native&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;setting-up-the-editor&quot;&gt;Setting Up the Editor&lt;/h3&gt;
&lt;p&gt;Once the package is installed, initialize the editor by importing the necessary modules and creating an instance of &lt;code&gt;EditorSettingsModel&lt;/code&gt;. You’ll need a license key, which you can obtain from the IMG.LY dashboard.&lt;/p&gt;
&lt;p&gt;Here’s how to set up and launch the video editor:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;tsx&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; IMGLYEditor, { EditorSettingsModel } &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;@imgly/editor-react-native&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;export&lt;/span&gt;&lt;span&gt; const&lt;/span&gt;&lt;span&gt; openVideoEditor&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; async&lt;/span&gt;&lt;span&gt; ()&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; Promise&lt;/span&gt;&lt;span&gt;&amp;#x3C;&lt;/span&gt;&lt;span&gt;void&lt;/span&gt;&lt;span&gt;&gt; &lt;/span&gt;&lt;span&gt;=&gt;&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  const&lt;/span&gt;&lt;span&gt; settings&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; new&lt;/span&gt;&lt;span&gt; EditorSettingsModel&lt;/span&gt;&lt;span&gt;({&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    license: &lt;/span&gt;&lt;span&gt;&apos;YOUR_LICENSE_KEY&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  });&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  const&lt;/span&gt;&lt;span&gt; result&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; await&lt;/span&gt;&lt;span&gt; IMGLYEditor.&lt;/span&gt;&lt;span&gt;openEditor&lt;/span&gt;&lt;span&gt;(settings);&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;};&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This launches the editor with the &lt;strong&gt;Video Editor&lt;/strong&gt; preset, enabling users to trim, cut, and enhance their videos with filters, text overlays, stickers, and music.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1080px) 1080px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1080&quot; height=&quot;1080&quot; src=&quot;https://img.ly/_astro/Video-UI_Z1lXuUU.webp&quot; srcset=&quot;/_astro/Video-UI_2qSkEH.webp 640w, /_astro/Video-UI_ZfvyXY.webp 750w, /_astro/Video-UI_EFnqb.webp 828w, /_astro/Video-UI_Z1lXuUU.webp 1080w&quot;&gt;&lt;/p&gt;
&lt;h2 id=&quot;use-cases-building-a-tiktok-like-experience&quot;&gt;Use Cases: Building a TikTok-Like Experience&lt;/h2&gt;
&lt;p&gt;Now that the video editor is integrated into your React Native app, let’s explore some key use cases and how to configure the editor to support them.&lt;/p&gt;
&lt;h3 id=&quot;short-form-video-creation&quot;&gt;Short-Form Video Creation&lt;/h3&gt;
&lt;p&gt;For apps targeting TikTok-style short-form video content, prioritize features that simplify the editing process:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Timeline Control&lt;/strong&gt;: Allow users to trim clips and sync overlays like stickers or music.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Filters &amp;#x26; Effects&lt;/strong&gt;: Offer filters to let users add certain moods to their videos with themes like retro, high contrast, or pastel tones.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Text &amp;#x26; Stickers&lt;/strong&gt;: Add captions and playful stickers for videos that are often consumed on mute.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Music &amp;#x26; Audio&lt;/strong&gt;: Let users add soundtracks or sound effects, with a library of trending music for inspirational background music.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Video Templates&lt;/strong&gt;: Provide ready-made templates that users can customize for specific occasions or themes.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This setup allows influencers and brands to create professional-looking videos that they can share across social media platforms with ease.&lt;/p&gt;
&lt;h3 id=&quot;influencers-and-marketing&quot;&gt;Influencers and Marketing&lt;/h3&gt;
&lt;p&gt;If your app serves influencers or businesses, the ability to create polished, on-brand videos is key. Features to consider:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Branded Templates&lt;/strong&gt;: Pre-designed templates aligned with specific brand aesthetics, simplifying content creation for marketing campaigns.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Watermarks&lt;/strong&gt;: Add brand logos or watermarks to videos, ensuring creators and brands maintain visibility.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These features empower influencers and businesses to quickly generate content that’s ready for distribution across platforms.&lt;/p&gt;
&lt;h3 id=&quot;e-commerce-and-user-generated-content&quot;&gt;E-commerce and User-Generated Content&lt;/h3&gt;
&lt;p&gt;Video editing is a powerful tool for e-commerce apps, enabling sellers and customers to create engaging product demos, unboxing videos, reviews and tutorials. Features like trimming, filters, and music tools make it easy to produce compelling content that enhances the shopping experience.&lt;/p&gt;
&lt;p&gt;For example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Vendors can create promotional videos for products.&lt;/li&gt;
&lt;li&gt;Customers can share authentic reviews or tutorials, increasing trust and boosting sales.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;customization-options-with-creativeeditor-sdk&quot;&gt;Customization Options with CreativeEditor SDK&lt;/h2&gt;
&lt;p&gt;The React Native plugin offers a range of customization options to adapt the editor to your app’s needs:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;UI Customization&lt;/strong&gt;: Change themes, colors, and layouts to match your app’s branding.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Video Presets&lt;/strong&gt;: Configure presets for specific use cases, such as limiting video length or optimizing for a particular resolution.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Custom Assets&lt;/strong&gt;: Add unique filters, stickers, and fonts to create a personalized experience for your audience.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Templates&lt;/strong&gt;: Use the CreativeEditor SDK Web UI to create custom templates, enabling users to start with professional-grade designs.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;Integrating a video editor into your React Native appwill improve your UX, and help boost engagement, retention, and potential distribution of your product, whether you’re building a social media platform, an influencer tool, or an e-commerce app.&lt;br&gt;
With CreativeEditor SDK, you can create a TikTok-like video editing experience or offer specialized tools for businesses and creators.&lt;/p&gt;
&lt;p&gt;By following the steps in this guide, you can empower your users to create professional-quality video content directly within your app. Explore the full capabilities of the SDK by visiting the &lt;a href=&quot;https://img.ly/docs/cesdk/react-native/prebuilt-solutions/video-editor-9e533a/&quot;&gt;React Native documentation&lt;/a&gt; and getting started today.&lt;/p&gt;
&lt;p&gt;Stay tuned for updates, and don’t hesitate to &lt;a href=&quot;https://img.ly/forms/contact-sales/&quot;&gt;reach out&lt;/a&gt; with any questions.&lt;/p&gt;
&lt;p&gt;Thanks for reading!&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;3,000+ creative professionals gain early access to new features and updates. Don’t miss out, and&lt;/strong&gt; &lt;a href=&quot;https://share.hsforms.com/1IgAOV1wASXGPnFG4ZPLejg1hk3i?ref=img.ly&quot;&gt;&lt;strong&gt;subscribe&lt;/strong&gt;&lt;/a&gt; &lt;strong&gt;to our newsletter.&lt;/strong&gt;&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="https://blog.img.ly/2024/12/how-to-react-native-video-editor-sdk.jpg" medium="image"/><category>React Native</category><category>How-To</category><category>Video Editor</category><category>Expo</category></item><item><title>Javascript Video Editing: Ultimate Guide for Developers and PMs</title><link>https://img.ly/blog/javascript-video-editing-guide/</link><guid isPermaLink="true">https://img.ly/blog/javascript-video-editing-guide/</guid><description>A comprehensive guide to JavaScript video editing for developers and PMs. Learn about essential features, modern web technologies, and common tools like FFmpeg.js and IMG.LY&apos;s CE.SDK to create powerful web-based video editors.</description><pubDate>Fri, 20 Dec 2024 14:06:55 GMT</pubDate><content:encoded>&lt;p&gt;When developing or integrating a JavaScript-based video editor for the web, you must consider a number of factors to ensure the solution is both efficient and robust. This post is the definitive guide for anyone embarking on such a project. It explores the key technologies involved, their strengths and weaknesses, and how different use cases influence the choice of tech stack. We’ll examine diverse use cases, from lightweight, browser-based editors for quick edits to more advanced tools requiring complex processing and rendering. We’ll then discuss how these scenarios drive the selection of features and technology.&lt;/p&gt;
&lt;p&gt;Additionally, we will show you the best open-source solutions that can accelerate development. This technical analysis will help you make an informed “build vs. buy” decision, ensuring you select the right approach for your project. Throughout this post, we will use the example of the IMG.LY’s JavaScript video editor (see &lt;a href=&quot;https://img.ly/demos/video-ui/web/&quot;&gt;here for a demo&lt;/a&gt;), showing you how these considerations shaped its architecture and feature set. You will learn practical insights into the decision-making process behind a successful web-based video editor.&lt;/p&gt;
&lt;p&gt;Modern web technologies like WebGL, WebCodecs and WebAssembly have enabled browsers to efficiently perform resource-intensive tasks such as video editing that had hitherto been the sole domain of desktop applications. As a result Javascript based video editing tools such as &lt;a href=&quot;https://www.veed.io/&quot;&gt;veed.io&lt;/a&gt; have increased in popularity and users expects ever more sophisticated video editing capabilities inside modern video based web apps.&lt;/p&gt;
&lt;h2 id=&quot;video-editing-use-cases&quot;&gt;Video Editing Use Cases&lt;/h2&gt;
&lt;p&gt;Let’s first outline use cases and requirements of video editing applications on the web. These will inform our discussion of technical features of respective solutions below and grant a conceptual framework for evaluating potential benefits and drawbacks of each solution.&lt;/p&gt;
&lt;h3 id=&quot;simple-video-editing&quot;&gt;Simple Video Editing&lt;/h3&gt;
&lt;p&gt;Let’s start with the base case, &lt;strong&gt;simple video edits&lt;/strong&gt; including trimming, cropping and resizing and simple effects such as adjusting brightness or saturation. This is usually sufficient feature set to support video editing for the following use cases:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sales Outreach Videos:&lt;/strong&gt; Sales teams often need to quickly edit customer-specific videos to personalize their outreach. This may involve trimming irrelevant portions, adding company logos, or adjusting brightness to ensure visual clarity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Messaging Applications:&lt;/strong&gt; These often require basic editing tools to allow users to crop, trim, or apply simple filters to shared videos, ensuring they’re concise and visually appealing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Screencasting:&lt;/strong&gt; Screencasting tools benefit from trimming and resizing capabilities to focus on key parts of recorded screens. Adding effects like brightness adjustment can make tutorials clearer and more professional.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;CMS Systems:&lt;/strong&gt; Content management systems may offer built-in video editing to help users optimize media assets for specific platform requirements, such as resizing for web embeds or adding subtle branding.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Screen Recording Applications:&lt;/strong&gt; Screen recording applications often include simple editing options for cleaning up recorded content by trimming extraneous sections or cropping to highlight the most relevant parts.&lt;/p&gt;
&lt;h3 id=&quot;video-annotation&quot;&gt;Video Annotation&lt;/h3&gt;
&lt;p&gt;Next, users might want to add an additional layer of information to videos and overlay other assets, such as stickers, shapes, overlays, or text. Crucially these assets need to be time-based. That is, users need control over when the asset is shown so that particular video sequences can be referenced. This introduces the need for a timeline to arrange different video components relative to each other in time, as well as features like voice-over and audio track support.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;E-commerce Reviews:&lt;/strong&gt; Sellers and reviewers can annotate product demo videos with callouts, price tags, and feature highlights to make the content more engaging and informative.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Claims Management / Insurance:&lt;/strong&gt; Insurance companies can use annotation to highlight key details in submitted video claims, such as timestamps of damage or explanatory text overlaying critical sections of footage.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Real Estate:&lt;/strong&gt; Realtors might annotate property walkthroughs by adding labels, dimensions, or descriptive text overlays to highlight key features of the home or property.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Educational Applications:&lt;/strong&gt; Instructors can use annotation tools to emphasize key moments in lectures or tutorials, such as overlaying text with formulas or concepts, or adding visual shapes to guide attention.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Productivity Tools:&lt;/strong&gt; Users can annotate meeting recordings with timestamps, text notes, or overlay diagrams to summarize key decisions or action points.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Healthcare and Telemedicine:&lt;/strong&gt; Medical professionals might annotate diagnostic videos or procedure recordings to explain findings or highlight areas of interest for training purposes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Customer Support &amp;#x26; Onboarding Tools:&lt;/strong&gt; Companies can add annotations to video tutorials or troubleshooting guides to direct users through specific steps or highlight important information.&lt;/p&gt;
&lt;h3 id=&quot;video-composition&quot;&gt;Video Composition&lt;/h3&gt;
&lt;p&gt;As we’re moving away from single media editing to creating video compositions of several types of media such as audio tracks, text, images, animations, and effects to create appealing visual designs, a well-designed timeline becomes even more important. Users need to manage many different video components in time, requiring user-friendly ways to browse and integrate external assets. This editor variant is most relevant for the following use cases:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Marketing Tech (Promotional Videos):&lt;/strong&gt; Marketers can create visually stunning promotional videos by combining custom animations, background music, and text overlays for branding.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Social Media (Stories and Reels):&lt;/strong&gt; Social media creators can quickly craft short, engaging videos that combine multiple assets like stickers, animations, and dynamic transitions tailored to platform-specific formats.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Event Highlight Reels:&lt;/strong&gt; Event organizers can compile videos, photos, and music into cohesive highlight reels that encapsulate the essence of the occasion.&lt;/p&gt;
&lt;h3 id=&quot;template-based-video-creation&quot;&gt;Template-based Video Creation&lt;/h3&gt;
&lt;p&gt;For many applications, users need starting points and examples for their designs. Starting from a blank canvas is rarely necessary for most use cases. Take a simple product video that includes promotional text, a brand logo, and animations. There is no need to reinvent the wheel for these types of videos. Instead, design applications should offer template libraries to accelerate user workflows.&lt;/p&gt;
&lt;p&gt;Once we introduce workflows involving several stakeholders, such as designers setting up a certain brand framework and marketers working with and adapting these preconfigured designs, templates need to be able to constrain what edit operations adaptors can perform. These requirements are particularly important for these use cases:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Digital Asset Management:&lt;/strong&gt; Organizations can manage and deploy branded video templates across teams, ensuring consistent design and messaging.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Social Media Publishing:&lt;/strong&gt; Social media managers can utilize templates for quick turnaround on platform-specific video formats, such as Instagram Stories or LinkedIn posts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Marketing Tech:&lt;/strong&gt; Marketing teams can rely on pre-designed templates to churn out campaign videos at scale while maintaining brand consistency.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Training Videos:&lt;/strong&gt; HR or L&amp;#x26;D departments can adapt existing templates to quickly produce training materials customized for different teams or scenarios.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;IMG.LY&amp;#39;s Video Editor enables template based constraints&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1380px) 1380px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1380&quot; height=&quot;789&quot; src=&quot;https://img.ly/_astro/Video-Templating_Z2kkIrS.webp&quot; srcset=&quot;/_astro/Video-Templating_Z1JYSMc.webp 640w, /_astro/Video-Templating_jgTfX.webp 750w, /_astro/Video-Templating_1F1X1m.webp 828w, /_astro/Video-Templating_Z21vhAn.webp 1080w, /_astro/Video-Templating_Z22RxuK.webp 1280w, /_astro/Video-Templating_Z2kkIrS.webp 1380w&quot;&gt;&lt;/p&gt;
&lt;h3 id=&quot;creative-automation-for-video&quot;&gt;Creative Automation for Video&lt;/h3&gt;
&lt;p&gt;Finally, you can leverage data and automation to generate variations of your templates at scale, significantly boosting the productivity of use cases where users need to either test a large number of designs, publish to different channels with different requirements, or adapt designs to many slightly different instances, for example, menu designs for a franchise.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Marketing Tech:&lt;/strong&gt; Marketers can generate hundreds of ad variations for A/B testing or localization by dynamically adjusting text, visuals, or calls to action.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Social Media Publishing:&lt;/strong&gt; Social media teams can automate the creation of videos tailored to platform specifications, such as aspect ratios or resolution, while retaining core branding.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;E-commerce:&lt;/strong&gt; Retailers can produce personalized product showcase videos for different customer segments, featuring tailored offers, pricing, or recommendations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Hospitality Industry:&lt;/strong&gt; Hotels and restaurants can dynamically generate location-specific promotional videos showcasing seasonal offers, menus, or events.&lt;/p&gt;
&lt;h2 id=&quot;essential-video-editing-features&quot;&gt;Essential Video Editing Features&lt;/h2&gt;
&lt;p&gt;The above use cases are enabled by a set of features from simple transforms to complex compositions. Some video editors are focused more on manipulating individual video clips while others are more oriented towards video creation providing fully-fledged multi-media composition tool. This concise overview serves as a reference for evaluating video editing solutions:&lt;br&gt;
Transforms&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Cut, Trim, and Split&lt;/strong&gt;: Basic editing tools for segmenting video clips.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Resize and Scale&lt;/strong&gt;: Adjusts clip dimensions, ensuring proper fit or aspect ratio.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Crop and Rotate&lt;/strong&gt;: Controls framing and orientation.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Zooming Capabilities&lt;/strong&gt;: A UX feature enabling detailed editing and close-up views for precision.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;adjustments&quot;&gt;Adjustments&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Basic Adjustments&lt;/strong&gt;: Controls brightness, contrast, and other visual enhancements.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Filters and Effects&lt;/strong&gt;: Sets the visual appearance and atmosphere of videos.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Audio Tracks and Mixing&lt;/strong&gt;: Supports multi-track audio adjustments and volume balancing.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Text and Overlay Options&lt;/strong&gt;: Allows visual enhancements with text, captions, and emojis, providing opacity and layering options.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;composition&quot;&gt;Composition&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Multi-media Composition&lt;/strong&gt;: Overlaying images, stickers, text as well as audio tracks are essential for creating videos through composition.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Multi-Track Editing&lt;/strong&gt;: Elements and layers need to be position relative to each other in time allowing the creation of complex composition.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Canvas-Based Editing&lt;/strong&gt;: Provides a flexible workspace for positioning and layering media elements.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Timeline Management&lt;/strong&gt;: Split, join and arrange clips on a timeline. Manages the positioning of clips and timing for seamless transitions.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Animations&lt;/strong&gt;: Make static elements dynamic and control their behavior in time. Can be use to create videos from scratch and enhance the storytelling workflow of existing ones.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;img alt=&quot;IMG.LY&amp;#39;s Video Editor Timeline Feature&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1376px) 1376px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1376&quot; height=&quot;960&quot; src=&quot;https://img.ly/_astro/Video-Editor-Timeline_1BcwDq.webp&quot; srcset=&quot;/_astro/Video-Editor-Timeline_rx9BU.webp 640w, /_astro/Video-Editor-Timeline_V8i7J.webp 750w, /_astro/Video-Editor-Timeline_Z1kAtyl.webp 828w, /_astro/Video-Editor-Timeline_Z1r7Sqi.webp 1080w, /_astro/Video-Editor-Timeline_11ktN6.webp 1280w, /_astro/Video-Editor-Timeline_1BcwDq.webp 1376w&quot;&gt;&lt;/p&gt;
&lt;h2 id=&quot;the-technology-landscape&quot;&gt;The Technology Landscape&lt;/h2&gt;
&lt;p&gt;Some time in the early 2000s the web started to be viewed as a platform to run applications not just a distributed database to deliver static html documents. While Javascript development was not yet at a point to serve as foundation for the kind of application development we see today, Adobe Flash filled the role of development platform. Video editing in the early days of the web was mainly provided by Flash-based tools involving simple transform operations such as trimming and basic transitions. Java applets were another way to deliver processing intensive programs via the browser while executing them outside of the browser inside the JVM. These early attempts were either too clunky, since they could not run natively in the browser or lacked the performance to be useful for high quality editing.&lt;/p&gt;
&lt;p&gt;When HTML5 arrived, reaching a W3C recommendation in 2014, it became possible to handle multimedia content natively in the browser. Coupled with significant performance advances of Javascript and the introduction of WebRTC this opened the door for video editing on the web.&lt;/p&gt;
&lt;p&gt;Some early cloud-based editors, such as WeVideo and Magisto, while more convenient and user-friendly than its predecessors still suffered from performance issues due to latency and lacked the depth of features found in desktop software.&lt;/p&gt;
&lt;p&gt;In the past few years advances such as WebAssembly and Javascript libraries built upon it e.g. ffmpeg.js have transformed browser-based video editing. User can now perform complex editing tasks like multi-track timelines, video effects and real-time in-browser rendering. Modern web-based solution have even come to rival desktop apps, because they can take advantage of GPU acceleration, modern APIs like WebGL and UI frameworks such as React.&lt;/p&gt;
&lt;p&gt;Let’s have a look at the state of the art web technologies that allow us to create performant video editing experiences inside the browsers. We are going to explore their relative strength, their complexities and conclude with when to best use each of these technologies:&lt;/p&gt;
&lt;h3 id=&quot;canvas-api&quot;&gt;Canvas API&lt;/h3&gt;
&lt;p&gt;The &lt;a href=&quot;https://developer.mozilla.org/en-US/docs/Web/API/Canvas_API&quot;&gt;HTML5 Canvas API&lt;/a&gt; should be most familiar to most modern web developers, since it’s a familiar and popular way to render 2D graphics in the browser. It’s use for video editing is limited and restricted to simple overlays, trimming, composition and animation:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Complexity:&lt;/strong&gt; Canvas is the least complex of the technologies introduced below and overlaps with the WebGL feature set to some extend, it’s API is well-known to most web developers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Advantages:&lt;/strong&gt; Canvas is lightweight, compatible across all browsers, and easy to implement, making it a great choice for simple use cases that require basic effects and no timeline.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Downsides:&lt;/strong&gt; While well suited for simpler tasks Canvas lacks the GPU-acceleration of WebGL, hence it is not recommended for resource intensive operation or performance sensitive use cases.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;When to Use:&lt;/strong&gt; Choose Canvas for simple editing tasks, like trimming or basic overlays, or if you need a fallback technology that works well on legacy browsers and devices.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;webgl&quot;&gt;WebGL&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://developer.mozilla.org/en-US/docs/Web/API/WebGL_API&quot;&gt;WebGL&lt;/a&gt; opens up the power of GPU-accelerated graphics to the browser, making it the default choice for any use case requiring high performance rendering. As such modern Javascript video editing solution are well served by webGL when looking to render visual effects in real time, add transitions or facilitate multi-layer compositions by leveraging the GPU.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Complexity:&lt;/strong&gt; WebGL comes with a steep learning curve especially if you are unfamiliar with 3D graphics processing, shaders and GPU programming.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Advantages:&lt;/strong&gt; WebGL delivers high performance for processing complex video effects and can handle significant volume of data without latency. That makes it ideal for advanced, highly responsive video editing features.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Downsides:&lt;/strong&gt; It can be challenging to engineer modular and maintainable WebGL code, debugging shaders and ensuring consistent performance across devices is an art in itself. WebGL might also be overpowered for simple editing use cases that do not justify the added complexity.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;When to Use:&lt;/strong&gt; Use WebGL if you’re creating a video editor with high-performance demands, particularly if real-time effects, transitions, or multi-layer compositing are needed.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;webcodecs-api&quot;&gt;WebCodecs API&lt;/h3&gt;
&lt;p&gt;Next up in our arsenal of formidable web technologies is a recent addition to modern browsers, WebCodecs, provide direct access to hardware accelerated video encoding and decoding.&lt;/p&gt;
&lt;p&gt;The WebCodecs API is a recent addition to the browser, allowing direct access to hardware-accelerated video decoding and encoding. This API makes it easier to handle high-resolution video without delays, thanks to efficient decoding directly in the browser.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Complexity:&lt;/strong&gt; WebCodecs is relatively easy to use for those familiar with media formats. However, you may need other technologies for rendering and complex editing.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Advantages:&lt;/strong&gt; This API offers quick access to hardware acceleration, ideal for managing large video files. With it, you can efficiently handle high-definition video playback and streaming.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Downsides:&lt;/strong&gt; Browser support for WebCodecs is still expanding, so compatibility might be a concern. Additionally, WebCodecs alone doesn’t provide a full editing suite; it needs to be combined with rendering and compositing technologies.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;When to Use:&lt;/strong&gt; WebCodecs is perfect if you need efficient decoding and encoding, especially for handling high-definition playback or live streaming.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;webassembly-wasm&quot;&gt;WebAssembly (Wasm)&lt;/h3&gt;
&lt;p&gt;The next performance frontier browser pushed towards was enabling near-native performance by compiling languages like C++ or Rust into a format that efficiently runs in Javascript. This format is &lt;a href=&quot;https://webassembly.org/&quot;&gt;WebAssembly&lt;/a&gt;, which makes it possible to handle computationally expensive tasks like encoding, decoding or frame manipulation inside the browser.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Complexity:&lt;/strong&gt; Working with WebAssembly is significantly more complex that the technologies discussed above. Developers need knoweldge of C++ or Rust as well as experience compiling to Wasm.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Advantages:&lt;/strong&gt; WebAssembly enables browsers to handle complex video processing tasks, such as encoding and decoding, with near-native performance. Libraries such as FFmpeg build upon WebAssembly to make familiar video tools available to web developers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Downsides:&lt;/strong&gt; Development with Wasm is complex, and debugging can be a challenge. Additionally, Wasm modules may increase load times.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;When to Use:&lt;/strong&gt; Use WebAssembly when building native libraries for complex operation and expensive computations, such as encoding and decoding.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Our article o&lt;a href=&quot;https://img.ly/blog/how-to-build-a-video-editor-with-wasm-in-react/&quot;&gt;n building a video editor with React and Wasm&lt;/a&gt; gives you a real-world starting point for your own apps.&lt;/p&gt;
&lt;h3 id=&quot;mediastream-api&quot;&gt;MediaStream API&lt;/h3&gt;
&lt;p&gt;The &lt;a href=&quot;https://developer.mozilla.org/en-US/docs/Web/API/MediaStream&quot;&gt;MediaStream API&lt;/a&gt; is essential for gaining easy access to live video sources from within the browser, such as webcams. If your video editing app includes a feature for real-time recording or streaming the MediaStream API is an essential tool.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Complexity:&lt;/strong&gt; MediaStream exposes a fairly simple API and is relatively easy to learn for experienced web developers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Advantages:&lt;/strong&gt; This API is well-suited for video feeds with low latency such as recording or streaming and does require addition encoding or decoding.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Downsides:&lt;/strong&gt; Any use case requiring post-production video has to rely on additional technologies or frameworks, as the MediaStream API is limited to live feeds. The quality of the stream depends on the input source and can vary.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;When to Use:&lt;/strong&gt; Use MediaStream for applications with live recording or streaming functionality.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;web-audio-api&quot;&gt;Web Audio API&lt;/h3&gt;
&lt;p&gt;The final web technology essential for video editing is the &lt;a href=&quot;https://developer.mozilla.org/en-US/docs/Web/API/Web_Audio_API&quot;&gt;Web Audio API&lt;/a&gt;. Whether you want to add effects or mix multiple audio tracks, this API offers sophisticated tools for audio manipulation.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Complexity:&lt;/strong&gt; For developers unfamiliar with audio processing the learning curve can be stepp, however the API is well designed and there is a host of educational resources.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Advantages:&lt;/strong&gt; Web Audio API allows precise control over audio enabling your app to offer audio adjustments such as effects or reverb and multi-track editing.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Downsides:&lt;/strong&gt; It can be challenging to accurately synchronize audio with video, especially if your app includes a real-time component.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;When to Use:&lt;/strong&gt; Opt for Web Audio API when audio editing is a core feature of your video editor.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;using-these-technologies-together&quot;&gt;Using These Technologies Together&lt;/h3&gt;
&lt;p&gt;Together, these technologies can be use to build a comprehensive browser based editor that is virtually indistinguishable from its desktop based counterparts. Different parts of the video editing workflow can be handled by each technology, while some choices have to made depending on the complexity and requirements of your projects, e.g. whether you’ll need resource intensive processing or an integration will real-time sources.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Video Decoding&lt;/strong&gt; - WebCodecs API (for efficient video loading and playback).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Rendering&lt;/strong&gt; - WebGL (for high-performance effects) and Canvas API (for simpler editing features).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Complex Processing&lt;/strong&gt; - WebAssembly (for encoding/decoding and using native libraries like FFmpeg).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Audio Processing&lt;/strong&gt; - Web Audio API (for mixing and adding effects to soundtracks).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Live Recording&lt;/strong&gt; - MediaStream API (to capture live video from devices).&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&quot;open-source-javascript-video-editing-libraries&quot;&gt;Open Source Javascript Video Editing Libraries&lt;/h2&gt;
&lt;p&gt;Fortunately in many cases someone has already done the heavy lifting and abstracted those browser APIs into simple to use libraries and frameworks. We’ll explore the most prominent ones distinguishing between projects that are for processing raw video and those that already provide an API for you to use. It is important to note that the libraries introduced below do not provide a video editing UI, unfortunately there are no open source video editors offering an end user interface that meet the standards for inclusion in a production grade application. None of the projects we evaluated met the demands we place on third-party libraries with respect to robustness, stability, maintenance and future development. However many can provide good starting points and references on how to build your own UI such as &lt;a href=&quot;https://github.com/mifi/reactive-video&quot;&gt;Reactive Video&lt;/a&gt; for React UIs.&lt;/p&gt;
&lt;h3 id=&quot;ffmpegjs&quot;&gt;FFmpeg.js&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/ffmpegwasm/ffmpeg.wasm&quot;&gt;FFmpeg.js&lt;/a&gt; is the JavaScript port of the popular FFmpeg library compiled into WebAssembly. It enables developers to manipulate and process raw video files entirely in the browser. You can trim, convert formats, extract audio, and apply filters to videos without needing external software. However, FFmpeg just gives you the raw toolchain for building video editing features, you are still tasked with building the entire UX and editing workflows. We have an extensive guide on FFmpeg that you can consult to explore its capabilities and get started learning FFmpeg syntax.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Advantages:&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;Most comprehensive library for web based video and audio processing.&lt;/li&gt;
&lt;li&gt;Fully client-side, no need for server-side processing.&lt;/li&gt;
&lt;li&gt;Highly flexible, allowing fine-grained control over video workflows.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Downsides:&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;Fairly complex for beginners with a steep learning curve, requires familiarity with FFmpeg’s command-line syntax.&lt;/li&gt;
&lt;li&gt;FFmpeg is licensed under LGPL which makes it ill-suited for most commercial projects.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;**When to Use:**FFmpeg.js is best suited for projects with custom requirements and UI making it necessary to exert fine-grained control over video processing tasks, especially for use cases like format conversion, advanced editing pipelines, or serverless video workflows.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The following two libraries allow you to programatically create and edit videos, we focus on the most stable, popular libraries here, while there might be promising up and comers we want to ensure a degree of dependability of the third-party code to be used in production systems.&lt;/p&gt;
&lt;h3 id=&quot;remotion&quot;&gt;Remotion&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://www.remotion.dev/&quot;&gt;Remotion&lt;/a&gt; is a React-based framework that enables developers to edit or create videos programmatically in a declarative fashion using React components. React’s declarative syntax makes it comparatively simple to manage video content and Remotion provides a convenient UI for previewing and editing videos. While its studio UI cannot be embedded directly, Remotion can serve as the graphics processing engine for custom UIs. Although it is a commercial project, it offers a generous free tier for individuals and small businesses.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Advantages:&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;Leverages React’s declarative model for seamless video creation.&lt;/li&gt;
&lt;li&gt;Provides robust video preview capabilities for iterative development.&lt;/li&gt;
&lt;li&gt;Flexible integration with custom-built UIs.&lt;/li&gt;
&lt;li&gt;Generous free tier for non-commercial and small-scale projects.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Downsides:&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;The built-in studio UI cannot be embedded, requiring developers to create their own UI for embedding.&lt;/li&gt;
&lt;li&gt;React-specific, limiting use in non-React environments.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;**When to Use:**Use Remotion when building React-based video applications that require programmatic video generation or dynamic video content. It’s ideal for teams already familiar with React and for applications where embedding is not a core requirement.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;etrojs&quot;&gt;&lt;strong&gt;Etro.js&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;Similarly to Remotion, Etro.js is a JavaScript framework for programmatic video editing, it is framework agnostic and provides an API for simple composition of layers, effects and exporting videos.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Advantages:&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;Includes support for advanced effects using GLSL shaders.&lt;/li&gt;
&lt;li&gt;Designed for in-browser video workflows.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Downsides:&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;Minimal community support compared to larger projects.&lt;/li&gt;
&lt;li&gt;Requires additional effort to create a full UI.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;When to Use:&lt;/strong&gt; Suitable for applications where programmatic control and customization are key.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;extensibility&quot;&gt;Extensibility&lt;/h2&gt;
&lt;p&gt;An important consideration when deciding to buy a prebuilt video editing solution is the degree of customization and extensibility required by your use case versus how critical speed to market is. While most SDKs provide enough flexibility by exposing cor API to build a solution that best suits your needs, this level of technical control comes at the expense of higher development and maintenance efforts. White-label solutions on the other end of the spectrum offer limited customization often relying on iFrames for embedding while providing maximum speed to market. If you are still in the phase where you need to make the business case for video editing features these might be key to allowing to quickly iterate. To learn more about weighing the pros and cons of SDK vs. White-label explore our &lt;a href=&quot;https://img.ly/blog/sdk-vs-white-label-consider-these-differences-when-you-choose-a-solution/&quot;&gt;extensive guide on the topic&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;IMG.LY’s CE.SDK achieves extensibility by providing an &lt;a href=&quot;https://img.ly/blog/img-ly-sdk-plugin-system/&quot;&gt;interface for plugins&lt;/a&gt; that allow developers to add custom features and UI elements.&lt;/p&gt;
&lt;h2 id=&quot;cross-platform-video-editing&quot;&gt;Cross Platform Video Editing&lt;/h2&gt;
&lt;p&gt;While this post is dedicated specifically to Javascript video editing an important consideration for choosing a video editing tech stack is whether there exists an intermediary output format that preserves editing operations so that videos can be edited on other devices and platforms. This is important for collaborative video editing use cases as well as use cases where videos are captured and basic edits performed on mobile and the finishing touches performed on the desktop.&lt;br&gt;
These consideration are particular pertinent for the Real Estate, Telemedicine and Productivity use cases described above, since they involve either a multi-device or collaborative component.&lt;/p&gt;
&lt;p&gt;At most common libraries or SDKs can achieve that by offering support for some form of serialization, that is exporting the edit operations along with the unedited video to then reapply edits upon import.&lt;/p&gt;
&lt;p&gt;IMG.LY’s CE.SDK, on the other hand, is the only video editing SDK that is truly cross-platform. That means it is built atop a single creative engine that is portable to any platform. Whether iOS, Android, Desktop, or the Web, every platform uses the same underlying tech and uses the same custom format for scenes which contain all assets and edits making error-prone serializations obsolete.&lt;/p&gt;
&lt;p&gt;Furthermore:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;iOS and Android make use of the same underlying API.&lt;/li&gt;
&lt;li&gt;Cross platform feature parity is baked into the cake. Since core functionality is implemented at the engine level, features are guaranteed to be available on both platform, although the timeline might differ somewhat.&lt;/li&gt;
&lt;li&gt;Designs are 100% interoperable and consistent between platforms. Exporting and importing design files across platforms works seamlessly and final renderings are guaranteed to be consistent.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;javascript-video-editing--generative-ai&quot;&gt;Javascript Video Editing &amp;#x26; Generative AI&lt;/h2&gt;
&lt;p&gt;Generative AI affects video editing and generation across the board. Text-to-video is the obvious case, but AI is also changing how videos get composed and edited in less visible ways: automated transitions, color grading, voice enhancement, voiceover generation, background removal, and inpainting.&lt;/p&gt;
&lt;p&gt;These capabilities come from a shifting set of specialist providers, and which one is best changes on a scale of months. Two design consequences follow from that.&lt;/p&gt;
&lt;p&gt;First, a JavaScript video editor has to be open enough to act as an integrator. It needs to combine voice tracks, clips, text, and other media programmatically before a user ever sees the timeline, and it needs to let you swap the model behind any one of those steps without rebuilding the editor.&lt;/p&gt;
&lt;p&gt;Second, the human has to stay in the loop. AI output is uneven, and refining an unchangeable artifact by re-prompting is slow and imprecise. Generated material should land in the scene as editable blocks, so users can trim, replace, and regenerate individual parts rather than rolling the dice on the whole clip again.&lt;/p&gt;
&lt;p&gt;We built &lt;a href=&quot;https://img.ly/blog/build-in-a-day-ai-video-clipping-with-ce-sdk/&quot;&gt;an AI video clipping tool in a day&lt;/a&gt; on that principle, and &lt;a href=&quot;https://img.ly/blog/animate-between-images-ai-native-video-workflows-with-ce-sdk-and-veo-3/&quot;&gt;animating between images with Veo 3&lt;/a&gt; shows the generated-but-still-editable workflow end to end.&lt;/p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;When building or integrating a JavaScript video editor, developers and PMs have to navigate a complex technology landscape and make myriad decisions along the way. This guide has covered everything from essential use cases and feature sets to the cutting-edge technologies that underly browser-based video editing. Whether your project requires basic trimming tools or advanced multimedia composition, understanding the benefits and downsides of the various solutions like Canvas, WebGL, WebCodecs, WebAssembly, and more is crucial for choosing the right stack.&lt;/p&gt;
&lt;p&gt;We’ve also explored how open-source tools like FFmpeg.js, Remotion, and Etro.js can accelerate development, while SDKs like IMG.LY’s CE.SDK offer unmatched extensibility and cross-platform compatibility. For those navigating the impact of generative AI on video editing, the need for seamless integration and user-centric design remains important.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;If you’re looking for a robust, modern video editing solution for your web application, explore our showcase and start a free trial.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;For the framework-specific version of this, we have setup guides for &lt;a href=&quot;https://img.ly/blog/a-modern-react-video-editor/&quot;&gt;React&lt;/a&gt;, &lt;a href=&quot;https://img.ly/blog/a-modern-vue-js-video-editor/&quot;&gt;Vue.js&lt;/a&gt; and &lt;a href=&quot;https://img.ly/blog/a-modern-angular-video-editor/&quot;&gt;Angular&lt;/a&gt;. If you are still comparing vendors rather than writing code, start with our breakdown of the &lt;a href=&quot;https://img.ly/blog/top-7-video-editing-sdks-in-2025/&quot;&gt;leading video editing SDKs&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;One practical note if you build with an AI assistant. The APIs in this space move quickly, and models tend to produce integrations against whichever version was in their training data. &lt;a href=&quot;https://img.ly/docs/cesdk/js/get-started/agent-skills-f7g8h9/&quot;&gt;Agent Skills for CE.SDK&lt;/a&gt; and the &lt;a href=&quot;https://img.ly/docs/cesdk/js/get-started/mcp-server-fde71c/&quot;&gt;MCP server&lt;/a&gt; load the current documentation into Claude Code, Cursor, and similar tools so the generated code matches the API that actually ships. We wrote about why in &lt;a href=&quot;https://img.ly/blog/img-ly-agent-skills-web/&quot;&gt;Introducing IMG.LY Agent Skills&lt;/a&gt;.&lt;/p&gt;</content:encoded><dc:creator>Jan</dc:creator><media:content url="https://blog.img.ly/2024/12/Javascript-Editor--1-.jpg" medium="image"/><category>Video Editing</category><category>Video Editor</category><category>CE.SDK</category><category>Automation</category><category>Cross-Platform</category><category>JavaScript</category></item></channel></rss>