<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>Mirko – IMG.LY Blog</title><description>Posts by Mirko on the IMG.LY blog.</description><link>https://img.ly/blog/author/mirko/</link><language>en-us</language><image><url>https://img.ly/apple-touch-icon.png</url><title>Mirko – IMG.LY Blog</title><link>https://img.ly/blog/author/mirko/</link></image><atom:link href="https://img.ly/blog/author/mirko/rss.xml" rel="self" type="application/rss+xml"/><generator>Astro</generator><lastBuildDate>Tue, 01 Sep 2026 09:42:35 GMT</lastBuildDate><ttl>60</ttl><item><title>Evaluating CoDesign MCP performance in resizing</title><link>https://img.ly/blog/evaluating-codesign-mcp-performance-in-resizing/</link><guid isPermaLink="true">https://img.ly/blog/evaluating-codesign-mcp-performance-in-resizing/</guid><description>One approved design, five ad formats, six AI models, the same brief. All six came back with ads you could ship, and Gemini 3.7 Flash did it for thirty cents. Here is what our automated checks can prove about that work, and what they cannot.</description><pubDate>Fri, 21 Aug 2026 09:00:00 GMT</pubDate><content:encoded>&lt;p&gt;You designed a banner. The same day, you need it in five sizes, and none of them can be wrong.&lt;/p&gt;
&lt;p&gt;We handed that job to six AI models through &lt;a href=&quot;https://img.ly/codesign/&quot;&gt;CoDesign&lt;/a&gt; MCP: take this approved design and rebuild it for five ad formats. All six came back with ads you could ship. Gemini 3.7 Flash did it for thirty cents and passed every check we wrote.&lt;/p&gt;
&lt;p&gt;That is what CoDesign MCP is for. You point an AI assistant such as Claude at your approved design, name the formats you need, and it opens the file, rebuilds the layout for each one and exports production files. What comes back is an editable design file, not a picture of one.&lt;/p&gt;
&lt;p&gt;Speed and price are the easy things to measure. We wanted to measure quality too, so we built an eval that runs the job automatically, opens every result and reports what passed. This is the first scenario out of that suite.&lt;/p&gt;
&lt;h2 id=&quot;three-banners-one-brief&quot;&gt;Three banners, one brief&lt;/h2&gt;
&lt;p&gt;Three different models were handed the same 320×50 slot, cut from the same master creative against the same brief, and all three came back with an ad you could traffic tomorrow. Every one of them keeps the wordmark and the button, nothing is clipped, and no type is set below 11px.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;Three 320×50 mobile banners built from the same ATLAS master by three different models&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1200px) 1200px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1200&quot; height=&quot;207&quot; src=&quot;https://img.ly/_astro/hook-three-mobile-strips.Q8QXs3k1_2XYzx.webp&quot; srcset=&quot;/_astro/hook-three-mobile-strips.Q8QXs3k1_Z1pTCV1.webp 640w, /_astro/hook-three-mobile-strips.Q8QXs3k1_2oKvSC.webp 750w, /_astro/hook-three-mobile-strips.Q8QXs3k1_27uKLv.webp 828w, /_astro/hook-three-mobile-strips.Q8QXs3k1_ZNKOe4.webp 1080w, /_astro/hook-three-mobile-strips.Q8QXs3k1_2XYzx.webp 1200w&quot;&gt;&lt;/p&gt;
&lt;p&gt;The 300×250 they came from carried five things: the wordmark, the headline, a retention chart, the offer line and the button. At 320×50 the wordmark and the button survive and the chart cannot, which leaves room for exactly one line of copy, so each model had to choose between the headline and the offer.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;One dropped the headline and kept the offer, “14-day free trial”.&lt;/li&gt;
&lt;li&gt;One kept the headline but cut “your users” down to “users” to buy itself the line.&lt;/li&gt;
&lt;li&gt;One kept the headline whole across two lines and dropped the offer.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;We have opinions about which of those reads best, the way any design team would. What we do not have is a way to prove one of them correct, because the choice depends on whether this campaign is selling the product or the trial, and that lives in the media brief rather than in the design. Our checks can confirm that all three are valid and that none of them are broken; they cannot tell you which one the campaign actually needed. Grading the half that has a right answer and leaving the other half to the people who own the brief is the distinction our evals are built around.&lt;/p&gt;
&lt;h2 id=&quot;what-is-gradeable-and-what-is-not&quot;&gt;What is gradeable and what is not&lt;/h2&gt;
&lt;p&gt;The obvious thing to evaluate is the thing that demos well, prompt in and poster out, but we were more interested in the work that actually fills a designer’s week, which is mostly derivative. The design decisions were made once, at the master, and everything after that is a matter of carrying them faithfully into thirty other shapes. That is more technical work than creative work, which is exactly why it is worth handing to an agent.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;Three valid 300×600 versions of the same ATLAS banner, differing in type size and spacing&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1200px) 1200px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1200&quot; height=&quot;702&quot; src=&quot;https://img.ly/_astro/gradeable-three-halfpage-variants.D3Bbuya5_2lpIzm.webp&quot; srcset=&quot;/_astro/gradeable-three-halfpage-variants.D3Bbuya5_Z1FdAzV.webp 640w, /_astro/gradeable-three-halfpage-variants.D3Bbuya5_Z1ihu0z.webp 750w, /_astro/gradeable-three-halfpage-variants.D3Bbuya5_FqKnt.webp 828w, /_astro/gradeable-three-halfpage-variants.D3Bbuya5_ZWQJb7.webp 1080w, /_astro/gradeable-three-halfpage-variants.D3Bbuya5_2lpIzm.webp 1200w&quot;&gt;&lt;/p&gt;
&lt;p&gt;Ask which of the finished designs is better and you have a question with no ground truth to appeal to, since human raters disagree with each other and LLM judges reliably prefer their own output. You could ask a lot of humans and rank the results, but that is not something you can run on every build. Wrong designs, though, are wrong in ways you can query: text sitting behind another block, content off the page, type under the legibility floor. Part of design quality is judgement and part of it is fact, and what we have built grades the second part, which means everything below is an automated result rather than a design review. A model can pass every check we wrote and still make an ugly ad.&lt;/p&gt;
&lt;h2 id=&quot;what-we-are-doing-right-now&quot;&gt;What we are doing right now&lt;/h2&gt;
&lt;p&gt;When a model finishes a scenario, programmatic checks open the result and ask whether it is plausibly good. That starts with whether the run produced every design artifact and .imgly file at the right sizes, and the rest depends on the scenario: in a localization run we also check that the words were actually replaced.&lt;/p&gt;
&lt;p&gt;This is a technical floor rather than a grade. LLM judges are still unreliable enough that we do not use one on output quality, because if it is hard for us as designers to say which design is more correct, it is harder for a model, and the result would be noise.&lt;/p&gt;
&lt;p&gt;In the future we plan to experiment more with such methodologies but we also do not want to rely on the false security of an extremely noisy grading process.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;The eval suite&amp;#39;s comparison view for the display-banner matrix&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 2396px) 2396px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;2396&quot; height=&quot;1344&quot; src=&quot;https://img.ly/_astro/evals-suite-comparison-resize-matrix.CV8e1T_u_8PXp1.webp&quot; srcset=&quot;/_astro/evals-suite-comparison-resize-matrix.CV8e1T_u_Znaqv1.webp 640w, /_astro/evals-suite-comparison-resize-matrix.CV8e1T_u_Z20zsc2.webp 750w, /_astro/evals-suite-comparison-resize-matrix.CV8e1T_u_Z26h51n.webp 828w, /_astro/evals-suite-comparison-resize-matrix.CV8e1T_u_Z1Eg8Vb.webp 1080w, /_astro/evals-suite-comparison-resize-matrix.CV8e1T_u_CxK9F.webp 1280w, /_astro/evals-suite-comparison-resize-matrix.CV8e1T_u_Z1lqFrF.webp 1668w, /_astro/evals-suite-comparison-resize-matrix.CV8e1T_u_ZgDr5i.webp 2048w, /_astro/evals-suite-comparison-resize-matrix.CV8e1T_u_8PXp1.webp 2396w&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The comparison view for the display-banner matrix: cost and time sitting directly above what each model actually produced. The 1/1 counts runs, one repeat per cell, not check scores.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&quot;one-master-five-ad-sizes&quot;&gt;One master, five ad sizes&lt;/h2&gt;
&lt;p&gt;The scenario is a display campaign, because that is the version of this problem customers actually have: the same size matrix, every campaign. The master is a 300×250 MPU for ATLAS, an imaginary product-analytics SaaS running a free-trial campaign, carrying a wordmark, the headline “See what your users &lt;strong&gt;actually&lt;/strong&gt; do”, a retention bar chart labelled “43% retained”, the line “14-day free trial. No card.”, and a “Start free trial” button.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;The ATLAS 300×250 master at the top, with an arrow fanning down to the same ad rebuilt at 300×600, 728×90, 320×50 and 160×600&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1200px) 1200px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1200&quot; height=&quot;675&quot; src=&quot;https://img.ly/_astro/master-and-five-sizes.CzF5wnY5_21RxX1.webp&quot; srcset=&quot;/_astro/master-and-five-sizes.CzF5wnY5_FUIkG.webp 640w, /_astro/master-and-five-sizes.CzF5wnY5_Z7tPBU.webp 750w, /_astro/master-and-five-sizes.CzF5wnY5_mzCbS.webp 828w, /_astro/master-and-five-sizes.CzF5wnY5_Zpv1pN.webp 1080w, /_astro/master-and-five-sizes.CzF5wnY5_21RxX1.webp 1200w&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The 300×250 master at the top, and the four other sizes on the media plan cut from it. ATLAS is our own fixture, not a real customer.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The brief asks for the full matrix: 300×250, 728×90, 160×600, 300×600 and 320×50. Those sizes span a 16:1 swing in aspect ratio, so the right answer at each one is a genuinely different composition, and at some of them it means taking content out rather than shrinking it.&lt;/p&gt;
&lt;p&gt;Ten checks grade the result and two of them do the real work. &lt;strong&gt;&lt;code&gt;728×90: the chart was dropped, not shrunk&lt;/code&gt;&lt;/strong&gt; kills the model that squashes a bar chart into a 90-pixel strip, which satisfies every geometric constraint in the brief while being completely wrong. &lt;strong&gt;&lt;code&gt;the five sizes are different layouts&lt;/code&gt;&lt;/strong&gt; kills the model that scales one composition five ways and calls it a matrix.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;Contact sheet of six 728×90 leaderboards, master first&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1280px) 1280px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1280&quot; height=&quot;676&quot; src=&quot;https://img.ly/_astro/sheet-display-leaderboard.BStENqrh_1X5bgP.webp&quot; srcset=&quot;/_astro/sheet-display-leaderboard.BStENqrh_2uIgrs.webp 640w, /_astro/sheet-display-leaderboard.BStENqrh_Z1sXCsn.webp 750w, /_astro/sheet-display-leaderboard.BStENqrh_Z1JQH3t.webp 828w, /_astro/sheet-display-leaderboard.BStENqrh_Z2gFqpY.webp 1080w, /_astro/sheet-display-leaderboard.BStENqrh_1X5bgP.webp 1280w&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The master first, then the six 728×90 leaderboards. Every model that delivered deleted the retention chart.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Every model that delivered arrived at the same composition, wordmark left, headline across the middle, offer underneath it, button on the right and the retention chart deleted, which is six models from five different vendors reaching the same answer independently. We did not expect that, and it is still the result we find hardest to stop thinking about.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&quot;Contact sheet of six 160×600 skyscrapers, master first&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 1280px) 1280px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;1280&quot; height=&quot;676&quot; src=&quot;https://img.ly/_astro/sheet-display-skyscraper.C88KXRcA_Z1lHK5K.webp&quot; srcset=&quot;/_astro/sheet-display-skyscraper.C88KXRcA_Z2cPTfD.webp 640w, /_astro/sheet-display-skyscraper.C88KXRcA_1krJOz.webp 750w, /_astro/sheet-display-skyscraper.C88KXRcA_ZmDw0F.webp 828w, /_astro/sheet-display-skyscraper.C88KXRcA_Z1S3Ozm.webp 1080w, /_astro/sheet-display-skyscraper.C88KXRcA_Z1lHK5K.webp 1280w&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The same six at 160×600, master first. With vertical room they all kept the chart. The flat format is where the decision to drop it had to be made.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&quot;which-model-is-best-for-using-with-codesign-mcp&quot;&gt;Which model is best for using with CoDesign MCP?&lt;/h2&gt;





























































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align=&quot;left&quot;&gt;#&lt;/th&gt;&lt;th align=&quot;left&quot;&gt;model&lt;/th&gt;&lt;th align=&quot;left&quot;&gt;checks&lt;/th&gt;&lt;th align=&quot;left&quot;&gt;model spend&lt;/th&gt;&lt;th align=&quot;left&quot;&gt;time&lt;/th&gt;&lt;th align=&quot;left&quot;&gt;tool calls&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align=&quot;left&quot;&gt;1&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;gemini-3.7-flash&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;10 / 10&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;$0.30&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;6.7 min&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;37&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align=&quot;left&quot;&gt;2&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;glm-5.2&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;10 / 10&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;$0.40&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;8.2 min&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;25&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align=&quot;left&quot;&gt;3&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;gpt-5.6-sol&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;10 / 10&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;$0.52&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;4.0 min&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;35&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align=&quot;left&quot;&gt;4&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;grok-4.6&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;10 / 10&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;$1.20&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;20.9 min&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;39&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align=&quot;left&quot;&gt;5&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;claude-fable-5&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;10 / 10&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;$3.95&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;10.3 min&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;26&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align=&quot;left&quot;&gt;6&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;claude-opus-5&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;10 / 10&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;$3.97&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;18.8 min&lt;/td&gt;&lt;td align=&quot;left&quot;&gt;40&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Running this scenario illustrates general numbers across our eval suite. The 6 models above are consistently able to follow the prompts to produce different output designs as requested in the scenario. These numbers also highlight quite the spread in model costs and time spend. What you do not see in these numbers is that glm 5.2, gpt-5.6-sol and grok 4.6 are producing suboptimal designs in our opinion. They are valid, but subjectively fable, opus and gemini 3.7 consistently across runs and across scenarios perform simply better.&lt;/p&gt;
&lt;p&gt;The biggest surprise for us is &lt;strong&gt;Gemini 3.7 Flash&lt;/strong&gt;, which produced the whole five-size matrix for $0.30 in model spend, in under seven minutes. At thirty cents a designer opens six composed variants instead of a blank artboard and takes the best one further, which changes what the tool is for.&lt;/p&gt;
&lt;p&gt;Meanwhile there were no surprises regarding &lt;strong&gt;Opus 5.0 and Fable 5.0&lt;/strong&gt;, both models perform really well - but are slow and expensive.&lt;/p&gt;
&lt;h2 id=&quot;what-does-it-mean-for-codesign-mcp&quot;&gt;What does it mean for CoDesign MCP&lt;/h2&gt;
&lt;p&gt;You can now resize a master design into 5 different sizes reliably for $0.30 using Gemini 3.7 Flash and the CoDesign MCP.&lt;/p&gt;
&lt;p&gt;If you are on a Claude Subscription plan, you might just want to continue using your Claude Code with the CoDesign MCP since it offers the best overall design quality - even if it’s slow and somewhat expensive.&lt;/p&gt;
&lt;p&gt;We will continue to work on upgrading and improving measurements of how well CoDesign MCP works with different models, looking into other scenarios like localization (translating and changing a design for a different locality and language) as well as rebranding existing design assets.&lt;/p&gt;
&lt;p&gt;You can install CoDesign MCP at &lt;a href=&quot;https://img.ly/codesign/&quot;&gt;&lt;strong&gt;img.ly/codesign&lt;/strong&gt;&lt;/a&gt;.&lt;/p&gt;</content:encoded><dc:creator>Dustin</dc:creator><dc:creator>Mirko</dc:creator><media:content url="/_astro/key-visual-resize-fanout.qeZDEyVs.png" medium="image"/><category>AI</category><category>Insights</category><category>CE.SDK</category></item><item><title>Make Your Docs Agent-Ready: Compiling MDX into Markdown</title><link>https://img.ly/blog/making-docs-machine-readable-why-we-native-compile-markdown-for-ai-agents/</link><guid isPermaLink="true">https://img.ly/blog/making-docs-machine-readable-why-we-native-compile-markdown-for-ai-agents/</guid><description>We rebuilt our docs pipeline to serve both HTML for humans and fully resolved Markdown for AI agents: 7× lighter and easier for tools to ingest.</description><pubDate>Wed, 11 Mar 2026 14:38:44 GMT</pubDate><content:encoded>&lt;h2 id=&quot;tldr-the-7x-efficiency-gain&quot;&gt;TL;DR: The 7x Efficiency Gain&lt;/h2&gt;
&lt;p&gt;We rebuilt our documentation pipeline to treat AI agents as a first-class audience. By natively compiling MDX into clean, fully-resolved Markdown (rather than heavy HTML or unresolved source code), agents and LLMs can now ingest our docs &lt;strong&gt;7x faster&lt;/strong&gt; and with far higher accuracy.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Test it yourself:&lt;/strong&gt;&lt;br&gt;
Request our docs with the &lt;code&gt;Markdown&lt;/code&gt; accept header to see the agent-optimized view:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;curl&lt;/span&gt;&lt;span&gt; -s&lt;/span&gt;&lt;span&gt; -H&lt;/span&gt;&lt;span&gt; &quot;Accept: text/markdown&quot;&lt;/span&gt;&lt;span&gt; https://img.ly/docs/cesdk/js/settings-970c98/&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;the-documentation-mismatch-humans-vs-agents&quot;&gt;The Documentation Mismatch: Humans vs. Agents&lt;/h2&gt;
&lt;p&gt;If you build developer tools today, you have two distinct audiences: human engineers and the AI agents they use to write code. Recently, we realized we were treating them exactly the same.&lt;/p&gt;
&lt;p&gt;Our docs pipeline is built on MDX and optimized for the browser. But when an LLM tries to ingest a page, it has to wade through navigation chrome, layout wrappers, and raw JSX. We didn’t just need a “clean text” mode; we needed a fundamentally different architecture.&lt;/p&gt;
&lt;h3 id=&quot;humans-vs-agents-a-comparative-overview&quot;&gt;Humans vs. Agents: A Comparative Overview&lt;/h3&gt;



































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Feature&lt;/th&gt;&lt;th&gt;Humans (Browser)&lt;/th&gt;&lt;th&gt;Agents (Context Window)&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Primary Format&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;HTML + CSS + JS&lt;/td&gt;&lt;td&gt;Clean Markdown&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Navigation&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Visual UI (Sidebar/Tabs)&lt;/td&gt;&lt;td&gt;Explicit links + hierarchy&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Context&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Implicit (Site layout)&lt;/td&gt;&lt;td&gt;Explicit (Frontmatter + headers)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Constraints&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;Performance / Core Web Vitals&lt;/td&gt;&lt;td&gt;Token / Context Window budgets&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Payload Size&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;~222 KB&lt;/td&gt;&lt;td&gt;&lt;strong&gt;~31 KB (7x smaller)&lt;/strong&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;h2 id=&quot;the-problem-why-html-and-raw-mdx-fail-agents&quot;&gt;The Problem: Why HTML and Raw MDX Fail Agents&lt;/h2&gt;
&lt;h3 id=&quot;1-html-is-expensive-chrome&quot;&gt;1. HTML is expensive “Chrome”&lt;/h3&gt;
&lt;p&gt;A documentation site is an application. Even static sites ship navigation chrome, scripts, and styling hooks. For AI, this is “token noise.” You are paying for bytes that provide zero value to the LLM and forcing the agent to reconstruct a hierarchy that you already had at authoring time.&lt;/p&gt;
&lt;h3 id=&quot;2-raw-mdx-is-unresolved-source-code&quot;&gt;2. Raw MDX is unresolved source code&lt;/h3&gt;
&lt;p&gt;Serving raw MDX files doesn’t solve the problem either. MDX is for maintainers. It is full of imports and unresolved dependencies. Our documentation often pulls code from external, tested repositories:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;## Initialize the Engine&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;#x3C;CodeBlock file=&quot;examples/getting-started/src/index.ts&quot; lines=&quot;12-24&quot; /&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;In the browser, this renders beautifully. In raw MDX, it’s a pointer to a file the agent cannot see. The actual code isn’t there.&lt;/p&gt;
&lt;h2 id=&quot;the-solution-treat-markdown-as-a-compilation-target&quot;&gt;The Solution: Treat Markdown as a Compilation Target&lt;/h2&gt;
&lt;p&gt;We stopped thinking about this as “exporting text” and started treating it as what it really is: a &lt;strong&gt;second build target&lt;/strong&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Old Pipeline:&lt;/strong&gt; &lt;code&gt;MDX (Source) → HTML (Browser View)&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;New Pipeline:&lt;/strong&gt; &lt;code&gt;MDX (Source) → HTML&lt;/code&gt; &lt;strong&gt;AND&lt;/strong&gt; &lt;code&gt;MDX (Source) → Markdown (Agent View)&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;why-we-didnt-just-convert-html&quot;&gt;Why we didn’t just “convert” HTML&lt;/h3&gt;
&lt;p&gt;We tried rendering to HTML and then running a converter. It failed at &lt;strong&gt;code blocks&lt;/strong&gt;. Syntax highlighting adds spans and wrappers around tokens; reversing that into clean, trustworthy code is nearly impossible.&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;html&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;#x3C;&lt;/span&gt;&lt;span&gt;pre&lt;/span&gt;&lt;span&gt;&gt;&amp;#x3C;&lt;/span&gt;&lt;span&gt;code&lt;/span&gt;&lt;span&gt; class&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&quot;language-js&quot;&lt;/span&gt;&lt;span&gt;&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#x3C;&lt;/span&gt;&lt;span&gt;span&lt;/span&gt;&lt;span&gt; class&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&quot;token keyword&quot;&lt;/span&gt;&lt;span&gt;&gt;const&amp;#x3C;/&lt;/span&gt;&lt;span&gt;span&lt;/span&gt;&lt;span&gt;&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;#x3C;&lt;/span&gt;&lt;span&gt;span&lt;/span&gt;&lt;span&gt; class&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&quot;token variable&quot;&lt;/span&gt;&lt;span&gt;&gt;engine&amp;#x3C;/&lt;/span&gt;&lt;span&gt;span&lt;/span&gt;&lt;span&gt;&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  ...&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;#x3C;/&lt;/span&gt;&lt;span&gt;code&lt;/span&gt;&lt;span&gt;&gt;&amp;#x3C;/&lt;/span&gt;&lt;span&gt;pre&lt;/span&gt;&lt;span&gt;&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;By working at the AST (Abstract Syntax Tree) level &lt;em&gt;before&lt;/em&gt; rendering, we avoid this entirely.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;implementation-transforming-mdx-at-the-ast-level&quot;&gt;Implementation: Transforming MDX at the AST Level&lt;/h2&gt;
&lt;p&gt;Rather than trying to reconstruct meaning from rendered markup, we preserve it directly. We built a &lt;code&gt;remark&lt;/code&gt; plugin that transforms MDX at the MDAST level.&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;javascript&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; { remark } &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;remark&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; remarkMdx &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;remark-mdx&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; remarkStringify &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;remark-stringify&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; { remarkTransformForExport } &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; &apos;./remarkTransformForExport&apos;&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; processor&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; remark&lt;/span&gt;&lt;span&gt;()&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  .&lt;/span&gt;&lt;span&gt;use&lt;/span&gt;&lt;span&gt;(remarkMdx) &lt;/span&gt;&lt;span&gt;// Parse MDX/JSX nodes&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  .&lt;/span&gt;&lt;span&gt;use&lt;/span&gt;&lt;span&gt;(remarkTransformForExport, {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    baseUrl: &lt;/span&gt;&lt;span&gt;&apos;[https://img.ly/docs/cesdk/](https://img.ly/docs/cesdk/)&apos;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    paths: resolvedPathMap,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  })&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  .&lt;/span&gt;&lt;span&gt;use&lt;/span&gt;&lt;span&gt;(remarkStringify); &lt;/span&gt;&lt;span&gt;// Back to markdown&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; markdown&lt;/span&gt;&lt;span&gt; =&lt;/span&gt;&lt;span&gt; await&lt;/span&gt;&lt;span&gt; processor.&lt;/span&gt;&lt;span&gt;process&lt;/span&gt;&lt;span&gt;(mdxContent);&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;the-component-contract&quot;&gt;The Component Contract&lt;/h3&gt;
&lt;p&gt;Every MDX component must define how it exports to the agent view. We colocate the transform directly with the component: &lt;code&gt;Aside.astro&lt;/code&gt; (Human UI) ↔ &lt;code&gt;Aside.toMarkdown.ts&lt;/code&gt; (Agent logic).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Example: Aside Component → Blockquote&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Input:&lt;/strong&gt; &lt;code&gt;&amp;#x3C;Aside title=&quot;Pro Tip&quot;&gt;Use the basePath...&amp;#x3C;/Aside&gt;&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Output:&lt;/strong&gt; &lt;code&gt;&gt; **Pro Tip:** Use the basePath...&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Example: Resolving Code References&lt;/strong&gt; This is the most critical transform. Instead of a file pointer, the agent gets the actual, inlined code block fetched during the build process.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;translating-navigation-ui-into-text&quot;&gt;Translating Navigation UI into Text&lt;/h2&gt;
&lt;p&gt;When you strip the sidebar and footer, you lose context. We reintroduce “navigational chrome” as text-native primitives at the top and bottom of every Markdown file:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; tabindex=&quot;0&quot; data-language=&quot;markdown&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;---&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;title&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;&apos;Working with Filters&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;platform&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;&apos;react&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;url&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;&apos;[https://docs.example.com/react/guides/filters/](https://docs.example.com/react/guides/filters/)&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;---&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&gt; You’re reading the React docs. For the full corpus, see llms-full.txt.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;## &lt;/span&gt;&lt;span&gt;**Path:**&lt;/span&gt;&lt;span&gt; [&lt;/span&gt;&lt;span&gt;Home&lt;/span&gt;&lt;span&gt;](&lt;/span&gt;&lt;span&gt;https://docs.example.com/&lt;/span&gt;&lt;span&gt;) &gt; [&lt;/span&gt;&lt;span&gt;Guides&lt;/span&gt;&lt;span&gt;](&lt;/span&gt;&lt;span&gt;https://docs.example.com/guides/&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;[Self-contained Page Content]&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;---&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;## Continue Reading&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;-&lt;/span&gt;&lt;span&gt; [&lt;/span&gt;&lt;span&gt;Color Adjustments&lt;/span&gt;&lt;span&gt;](&lt;/span&gt;&lt;span&gt;https://docs.example.com/react/guides/color-adjustments/&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;-&lt;/span&gt;&lt;span&gt; [&lt;/span&gt;&lt;span&gt;API Reference&lt;/span&gt;&lt;span&gt;](&lt;/span&gt;&lt;span&gt;https://docs.example.com/api/&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;lessons-for-scaling-agentic-dx&quot;&gt;Lessons for Scaling Agentic DX&lt;/h2&gt;
&lt;p&gt;If you maintain docs at scale, “optimizing for bots” is no longer optional. It is the new standard for Developer Experience (DX).&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Don’t serve HTML and hope:&lt;/strong&gt; Reverse-engineering structure is hard for AI. Give it the structure directly.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Every component needs an identity:&lt;/strong&gt; If a component carries meaning, it needs a Markdown equivalent. If it’s just layout, unwrap it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Resolve everything:&lt;/strong&gt; Assume zero ambient context. Links must be absolute, and code must be inlined.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Content Negotiation:&lt;/strong&gt; Use the &lt;code&gt;Accept: text/markdown&lt;/code&gt; header. It is becoming the industry standard for tools like Claude Code and OpenCode.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The Single-File Corpus:&lt;/strong&gt; In addition to per-page files, generate a &lt;code&gt;llms-full.txt&lt;/code&gt; that concatenates everything. Some agents prefer one large fetch over a crawl.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;By acknowledging that AI agents are a primary consumer of our documentation, we’ve made our SDK significantly easier to integrate. The future of docs isn’t just “readable”. It’s “storable and traversable.”&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;Note: We built this for CE.SDK. The implementation uses Astro and Vercel, but the approach is framework-agnostic.&lt;/em&gt;&lt;/p&gt;</content:encoded><dc:creator>Mirko</dc:creator><media:content url="https://blog.img.ly/2026/03/optimize-docs-for-agents-ai-llm-1.jpg" medium="image"/></item><item><title>IMG.LY Research: AI-based Generative Design Editing</title><link>https://img.ly/blog/img-ly-research-ai-based-generative-editing/</link><guid isPermaLink="true">https://img.ly/blog/img-ly-research-ai-based-generative-editing/</guid><description>Combine Large Language Models (LLMs) with CE.SDK for design edits and automated creative workflows.</description><pubDate>Tue, 16 Jul 2024 10:44:05 GMT</pubDate><content:encoded>&lt;p&gt;Generative AI is transforming the tech landscape, finding applications in virtually every field. At &lt;a href=&quot;https://img.ly/?utm_source=imgly&amp;#x26;utm_medium=blog&amp;#x26;utm_campaign=generative-editing&quot;&gt;IMG.LY&lt;/a&gt;, we’re exploring how these advancements can revolutionize creative workflows. This article presents a research project, where we integrate Large Language Models (LLMs) with our flagship product, CreativeEditor SDK (CE.SDK), to enable natural language-driven design edits.&lt;/p&gt;
&lt;p&gt;Our flagship product, CreativeEditor SDK (CE.SDK) allows for advanced creative workflows for countless use cases in industries ranging from &lt;a href=&quot;https://img.ly/industries/print/&quot;&gt;print&lt;/a&gt; to &lt;a href=&quot;https://img.ly/industries/marketing-tech/&quot;&gt;marketing tech&lt;/a&gt;. Most use cases can be realized with the out-of-the-box feature set, but it also exposes a best-in-class API, called Engine API, to build complex custom workflows with designs and videos.&lt;/p&gt;
&lt;p&gt;In this article, we will showcase how to combine our CE.SDK Engine API with LLMs to edit designs with natural language.&lt;/p&gt;
&lt;p&gt;&lt;video src=&quot;https://storage.googleapis.com/imgly-static-assets/static/blog/videos/imgly-research/imgly-ai-template-editor-demo.mp4&quot; controls autoplay muted loop playsinline&gt;&lt;/video&gt;&lt;/p&gt;
&lt;h2 id=&quot;introduction-to-llms&quot;&gt;Introduction to LLMs&lt;/h2&gt;
&lt;p&gt;Generative AI is often associated with chatbots, but its capabilities stretch further. It is a versatile text processor that can transform any textual input into various structured outputs. This adaptability is due to its training on diverse textual patterns, allowing it to support a wide range of text-to-text applications beyond just generating conversational prose.&lt;/p&gt;
&lt;p&gt;Carefully crafting an input text (prompt) in a way that instructs the LLM to output a specific structured output format allows us to use LLMs to solve almost any arbitrary text-based task.&lt;/p&gt;
&lt;p&gt;Crafting prompts is an art in itself. When ensuring that we receive the required output we need to adhere to the following steps:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Consider what type of data our model was trained on to ensure the correct formatting of our input text. The most basic example is using English as our “main prompt language” since most LLMs are mainly trained on English text samples.&lt;/li&gt;
&lt;li&gt;All necessary problem-specific information to solve the task needs to be included in the prompt. While LLMs often possess an inherent understanding of the world inferred from the vast amount of text they are trained on, they may not know much about our specific problem. Furthermore, LLMs have the notorious tendency to hallucinate, that is fill in missing context with incoherent or incorrect information. To ensure the best performance, the input to the LLM must provide as much context as possible.&lt;/li&gt;
&lt;li&gt;Finally, we need to instruct the LLM well enough to output text-based data in a format we can then parse and process.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;human-vs-ai-workflows-for-executing-design-tasks&quot;&gt;Human vs. AI Workflows for Executing Design Tasks&lt;/h2&gt;
&lt;p&gt;We started this project with the vision to use generative AI to magically handle requests like these (in increasing order of complexity):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;Make the logo bigger&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Translate this design into German&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Adopt this template to our brand colors and brand assets&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Transform this Instagram story portrait design into a landscape YouTube thumbnail&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;When trying to delegate a task to an AI, it’s best to start by thinking about how these tasks are currently solved by humans.&lt;/p&gt;
&lt;p&gt;Let’s walk through how a human would complete a task such as&lt;br&gt;
&lt;code&gt;Make the logo bigger&lt;/code&gt;:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Humans would visually scan the design and automatically segment it by its elements such as objects, backgrounds, or text.&lt;/li&gt;
&lt;li&gt;Humans would then read and comprehend the task “Make the logo bigger”&lt;/li&gt;
&lt;li&gt;Finally, users would use their existing knowledge of how to move and interact with design software to fulfill the task by manipulating the individual elements in the design.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Based on these considerations we can extract the implicit knowledge necessary to fulfill a task and make it explicit for the benefit of our LLM.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Design Representation:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;To enable the LLM to understand and manipulate a design, it is essential to provide a representation of the design. This can be either textual or a mix of textual and visual (if the LLM has vision capabilities) data. While supplying the current design as a raster image to the LLM is trivial, serializing a CE.SDK design into a textual format requires a custom serialization process. The textual representation is important since it allows the LLM to identify, address, and comprehend the different components of the design effectively.&lt;/p&gt;
&lt;p&gt;Refer to &lt;code&gt;Appendix: Use-Case dependent serialization of CE.SDK Designs&lt;/code&gt; for a more in-depth explanation of how to accomplish this.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Editing Protocol:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;LLMs do not interact with design software using traditional human interfaces like a mouse, keyboard, or visual feedback. Therefore, we need a specific protocol for the LLM to propose changes to the design. We have developed a method where we pass a textual representation of the design to the LLM as part of the prompt so that the LLM can indicate changes to the design by returning a modification of this representation.&lt;/p&gt;
&lt;p&gt;Practically, this means that if we pass in an element such as &lt;code&gt;&amp;#x3C;Image id=&quot;1337&quot; x=”100” y=&quot;100&quot; .../&gt;&lt;/code&gt;, the LLM can change those x and y attributes by simply returning &lt;code&gt;&amp;#x3C;Image id=&quot;1337&quot; x=&quot;0&quot; y=&quot;0&quot; .. /&gt;&lt;/code&gt; inside its output text. Since we can identify the design element that was changed using the ID attribute, we can then calculate the programmatic changes that need to be applied to the design, like in this case &lt;code&gt;engine.block.setPositionX(1337, 0)&lt;/code&gt; and &lt;code&gt;engine.block.setPositionY(1337, 0)&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;Refer to &lt;code&gt;Appendix: Parsing and transforming LLM response&lt;/code&gt; for a deeper look into this topic.&lt;/p&gt;
&lt;h2 id=&quot;using-generative-ai-to-execute-design-related-tasks&quot;&gt;Using Generative AI to Execute Design-Related Tasks&lt;/h2&gt;
&lt;p&gt;&lt;img alt=&quot;Employ LLMs to modify and improve CE.SDK design templates.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; sizes=&quot;(min-width: 2000px) 2000px, 100vw&quot; data-astro-image=&quot;constrained&quot; data-astro-image-pos=&quot;center&quot; width=&quot;2000&quot; height=&quot;879&quot; src=&quot;https://img.ly/_astro/ai-generated-design-templates-LLM_ZqAA29.webp&quot; srcset=&quot;/_astro/ai-generated-design-templates-LLM_Z1SedY4.webp 640w, /_astro/ai-generated-design-templates-LLM_1OUcAS.webp 750w, /_astro/ai-generated-design-templates-LLM_2l7vpa.webp 828w, /_astro/ai-generated-design-templates-LLM_mgL61.webp 1080w, /_astro/ai-generated-design-templates-LLM_Zw8us0.webp 1280w, /_astro/ai-generated-design-templates-LLM_16YhWD.webp 1668w, /_astro/ai-generated-design-templates-LLM_ZqAA29.webp 2000w&quot;&gt;&lt;/p&gt;
&lt;p&gt;Based on what we have learned, we assemble a workflow, with the LLM as the center point that allows the LLM to execute design-related tasks on any of our CE.SDK designs. This workflow can be divided into two sub-tasks: Composing an input text (prompt) with all necessary context and parsing and applying the output of the LLM to the CE.SDK design.&lt;/p&gt;
&lt;h3 id=&quot;composing-the-input-text&quot;&gt;Composing the Input Text&lt;/h3&gt;
&lt;p&gt;As seen in the graphic above we first compose the input text based on different components to provide the model with all necessary context to fulfill the user’s editing request: This includes a general text to “instruct” the model, the actual request that the user entered, an exemplary design representation to explain our output format to the model and a representation of the currently edited design.&lt;/p&gt;
&lt;p&gt;The general, static text to “instruct” the model is composed of the following three parts.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;What we are trying to achieve in general:&lt;br&gt;
&lt;code&gt;&quot;You are an AI with expertise in design, specifically focused on XML representations of designs&quot;&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Output format instructions:&lt;br&gt;
&lt;em&gt;Your responses should only contain one XML document. Ensure that you do not introduce new attributes to any XML elements. You can change image elements by setting the alt attribute, which then will be used to search Unsplash for a fitting image. A sample alt text is “A mechanic changing tires with a pair of beautiful work gloves on”.&lt;br&gt;
Additionally, pay close attention to the layout: verify that no elements in the XML document extend beyond the page boundaries. This constraint is critical for maintaining consistency and accuracy in XML formatting. Always double-check your XML output for these requirements.&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;The actual user request:&lt;br&gt;
e.g., &lt;code&gt;&quot;Make the logo bigger&quot;&lt;/code&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;By including a textual representation of a comprehensive example design we can show the model which layer types are available as well as which properties of those layers can be manipulated.&lt;/p&gt;
&lt;p&gt;The model now has all the necessary context and building blocks to respond to user requests in a format that we can process downstream.&lt;/p&gt;
&lt;h3 id=&quot;applying-the-output-text&quot;&gt;Applying the Output Text&lt;/h3&gt;
&lt;p&gt;In the first step, we scan the output text for an XML-like document and if we find one, attempt to parse it.&lt;/p&gt;
&lt;p&gt;This will yield a structured data object we can compare to the one we passed in and calculate which elements have been modified, added, or removed.&lt;/p&gt;
&lt;p&gt;The resulting change set can then be translated into specific calls to the &lt;a href=&quot;https://img.ly/docs/cesdk/js/user-interface/ui-extensions-d194d1/&quot;&gt;CE.SDK Engine API&lt;/a&gt; to change the current design.&lt;/p&gt;
&lt;h2 id=&quot;issues-faced&quot;&gt;Issues faced&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Latency&lt;/strong&gt;: One issue with state-of-the-art models is their big latency: Each LLM response contains approximately 1000 tokens. That means each request takes 7-45 seconds (depending on the model) to complete. This long delay may be unacceptable for some user experiences. However, we see this issue as transitory and expect upcoming models to have much smaller latency while maintaining their capabilities.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pricing&lt;/strong&gt;: Each request/response with GPT-4 turbo as a backing model costs around 5 cents and restricts some use cases. We also expect the pricing to drop significantly. The new GPT-4o model for example reduces the price by half.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Hallucinations&lt;/strong&gt;: LLMs do not always follow the instructions properly and, e.g., produce output that is not parsable. Hallucinations directly correlate with the capabilities of the model and this issue is not apparent at current state-of-the-art LLMs like e.g GPT-4/GPT-4o.&lt;/p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;We present a novel and adaptable approach to use Generative AI and LLMs specifically to interact with IMG.LY’s CreativeEditor SDK. We showcase how this technology can be used to execute common design requests on arbitrary CE.SDK Designs. The proposition that LLMs can understand textual representations of visual elements was by no means obvious. This research project has revealed that it is very well within the scope of LLMs to translate instructions from a visual semantic context to its textual representation and back. This invites more inquiries into LLMs as assistants for tasks with a heavy visual component such as design.&lt;/p&gt;
&lt;p&gt;While further research is needed to make this technology available in production environments, we are confident that Generative AI-based editing will play a big role in the future of Graphics and Video editing.&lt;/p&gt;
&lt;h2 id=&quot;appendix-use-case-dependent-serialization-of-cesdk-designs&quot;&gt;Appendix: Use-case Dependent Serialization of CE.SDK Designs&lt;/h2&gt;
&lt;p&gt;LLMs work based on “tokens” which are equal to words. However, a design, like e.g a poster design or a social media graphic, is highly visual. That means that we need a way to convert a design into text, a way to serialize it. Our CE.SDK engine can serialize an existing scene using our &lt;code&gt;engine.block.saveToString&lt;/code&gt; method. However, this serialization contains a huge pile of information that is not necessary to do edits inside the file. LLMs are priced by token and their speed is also relative to the number of tokens the input and output have. Thus, the number of tokens should be reduced.&lt;/p&gt;
&lt;p&gt;We looked at several ways to convert the current state of the design into a textual representation. Since GenAI is trained on a lot of (X)HTML which has an XML-like format, we decided to serialize any designs into a tag-based XML-like format.&lt;/p&gt;
&lt;p&gt;The IMG.LY editor internally refers to design elements like images or texts as “blocks”. These blocks are uniquely identifiable and addressable using a numeric ID. We use this ID to be able to identify a serialized design block in the input and output of the LLM. Example: &lt;code&gt;&amp;#x3C;Image id=&quot;12582927&quot; x=&quot;0&quot; y=&quot;0&quot; width=&quot;800&quot; height=&quot;399&quot; /&gt;&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;For each of the CE.SDK block types like e.g “Text” or “Graphics” (representing images or vector shapes), we will the CE.SDK Engine API to query very specific data from the block. That means that for example we only have a text attribute for Text blocks.&lt;/p&gt;
&lt;p&gt;This rather specific mapping of only certain properties from the CE.SDK design into the text serialization allows us to optimize the design serialization for different use cases. A use-case where we e.g. want to automatically name each layer does maybe not require fine-grained information about e.g the font size.&lt;/p&gt;
&lt;h2 id=&quot;appendix-parsing-and-transforming-llm-response&quot;&gt;Appendix: Parsing and transforming LLM response&lt;/h2&gt;
&lt;p&gt;The LLM answers with arbitrary tokens. It’s not possible to restrict the response to a certain syntax. By settling on a well-defined and widely used format we instruct the model to also reply with an XML-like document, similar to the one we passed in as “current state”.&lt;/p&gt;
&lt;p&gt;After receiving the LLM’s Response we first make sure that only a single XML document is present inside the response. We then compare the retrieved XML document with the state of the Design that we passed into the LLM and generate a change set. This change set contains entries like “Color of block with ID=123 has changed”. These change set entries are then converted into programmatic commands, like e.g &lt;code&gt;engine.block.setColor(123)&lt;/code&gt; and executed on the current design.&lt;/p&gt;
&lt;p&gt;One challenges when working with an LLM is the inability to restrict the output space. Thus, we are never guaranteed that the LLM did not add e.g new XML node names or that it even replies with a proper, valid XML-like document. The only lever to influence the probability of a proper XML-like document is to use strong prompting and LLM that are good at following those instructions.&lt;/p&gt;
&lt;p&gt;In our tests, state-of-the-art models like GPT-4 can follow those instructions without any further tooling.&lt;/p&gt;
&lt;h2 id=&quot;further-research-topics&quot;&gt;Further Research Topics&lt;/h2&gt;
&lt;p&gt;It’s also worth exploring fine-tuning an LLM specifically for this task which could improve the performance of the LLM for the specific tasks.&lt;/p&gt;
&lt;p&gt;It would also be possible to use more advanced libraries like &lt;a href=&quot;https://github.com/guidance-ai/guidance&quot;&gt;Guidance&lt;/a&gt;, which allows to define a grammar for the LLM response thus making sure that the output of the LLM is always parseable.&lt;/p&gt;
&lt;p&gt;Another way to improve the performance would be to methodically test different prompt templates and find a way to measure and compare the output quality.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thank you for reading!&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;3,000+ creative professionals gain exclusive access and hear of our releases first.&lt;/strong&gt; &lt;a href=&quot;https://share.hsforms.com/1IgAOV1wASXGPnFG4ZPLejg1hk3i&quot;&gt;&lt;strong&gt;Subscribe&lt;/strong&gt;&lt;/a&gt; &lt;strong&gt;to our newsletter and never miss out.&lt;/strong&gt;&lt;/p&gt;</content:encoded><dc:creator>Mirko</dc:creator><media:content url="https://blog.img.ly/2024/07/0_AI-Template-generator.jpg" medium="image"/><category>AI</category><category>Automation</category><category>Design Editor</category><category>Machine Learning</category></item></channel></rss>