In Instagram advertising, a proposal system using generative AI outperformed human designers. In the final comparison—where 10 images each were run under the same week, same budget, and same targeting conditions—the image set chosen by the proposal system achieved an average link click-through rate (CTR) of 0.98%, compared to 0.65% for the human designer's image set. What was being compared was how background images were selected for insertion into a fixed template; the logo and text remained unchanged. What was tested was a system that combines generation, brand-fit judgment, and learning from actual delivery data into a single loop.

On August 4, 2026, INFORMS published the results of a paper titled "Leveraging Generative Artificial Intelligence to Create Visual Content in Digital Advertising," published in Marketing Science. The paper was received on October 14, 2024, accepted on March 31, 2026, and published online on May 5, 2026, as peer-reviewed research. The authors are Remi Daviet and Yohei Nishimura, and the research funding was fully covered by the authors' affiliated institutions. These results demonstrate the performance of an ad-creation process that does not separate what to generate from what to test next based on actual delivery outcomes.

AD

0.98% vs. 0.65%: The Gap in Average Background Image CTR

The experiment was conducted using Instagram ads for an outdoor experience company. The target audience was parents aged 18-55 with children aged 3-12, and ads were placed in the Feed, Explore, and search results. Learning took place during the fourth quarter of 2023. The metric evaluated was link CTR—not clicks on the Instagram post itself, nor purchases, bookings, or sales, which were not treated as primary metrics.

The final comparison involved 10 images each from the proposal system, an AI optimized solely for aesthetic score, and human designers. The proposal system's average CTR was 0.98% (standard deviation 0.26%, range 0.67%1.52%), the aesthetic AI's was 0.78% (standard deviation 0.41%, range 0.00%1.30%), and the human designer's was 0.65% (standard deviation 0.31%, range 0.25%1.26%). Across 10,000 bootstrap iterations, the proposal system's mean exceeded the human designer's in 99.59% of cases, yielding p=0.0041.

Meanwhile, it exceeded the aesthetic AI's mean in 90.76% of cases, with p=0.0924—not statistically significant at the conventional 5% threshold. The human comparison relied on just one approved designer, and each group consisted of a small sample of only 10 images. The authors themselves note that caution is warranted in generalizing the results, citing the small sample size and the single human benchmark.

How to Find Candidates That Fit the Brand and Get Clicks

The system generates candidates within a 304-dimensional latent space—split into 48 dimensions for layout and 256 for style—and uses SDXL in the final stage to enhance quality. Conventional A/B testing and multi-armed bandit approaches are well-suited for selecting from a predetermined, finite set of candidates, but they don't determine what image to generate next. This design instead integrates generation and live-delivery testing into the same learning loop.

AcceptAI is a Bayesian neural network that predicts brand-fit conditions—photorealism, natural scenery appropriate to the target region, and absence of AI-derived artifacts. Company staff evaluated 60 batches of 12 images each, with manual work totaling under two hours. Including additional rejected images, the model was ultimately trained on 1,417 images.

The other component, PerfAI, is a separate Bayesian neural network that predicts expected CTR from an image's latent representation. It was trained on 125 images across 9 trials; each trial combined 12 highly informative candidate images with the top-predicted candidate and a control candidate to check for seasonal variation. Within the range AcceptAI deems brand-appropriate, the system selects the batch that most reduces prediction uncertainty about which images are likely to achieve high CTR. Whereas A/B testing selects a winner from a completed set of candidates, this proposed approach feeds test results back into image generation and explores untested regions of the space.

However, these results do not measure the pure causal effect of the image alone on an individual. The research aims to predict outcomes within a real operational environment, including Instagram's own delivery optimization. The average Spearman correlation when the same batch was redelivered was 0.43, indicating that variation in the delivery environment is embedded in the results.

AD

Beauty and Brand Fit Are Separate Axes

Raising general aesthetic scores did not directly translate into better ad performance. The aesthetic AI, optimized using TANet—which had a 75.8% Spearman correlation with human aesthetic ratings—achieved only a 0.78% average CTR, falling short of the proposal system's 0.98%. A model that measures visual appeal and a model that selects images driving link clicks are doing different jobs.

Within the training sample, only 5.3% of images ranked in the top quartile for both brand fit and predicted performance. Just 0.6% ranked in the top 10% on both dimensions. Candidates that satisfy a company's visual requirements while also generating engagement upon delivery occupy a narrow region of the candidate space. This is precisely why the procedure—maintaining fit via AcceptAI while learning performance through PerfAI and testing high-uncertainty regions—matters.

In the final comparison, the 10 images selected by the proposal method were evaluated as a single set. The proposal system's standard deviation of 0.26% was smaller than both the human designer's 0.31% and the aesthetic AI's 0.41%. Whether a candidate set avoided having performance concentrated in just a few images was assessed alongside average CTR.

Beyond CTR, the Evaluation Diverges

Without retraining, the authors re-tested the proposal system's original 10 images against 10 images the company had produced at the time, running both under identical conditions during a later peak booking period. The average CTR was 3.38% for the proposal system (standard deviation 0.76%, maximum 4.32%) and 3.24% for the company's images (standard deviation 0.83%, maximum 5.39%). However, across 10,000 bootstrap iterations, the proposal system outperformed in only 67.7% of cases, with a one-sided p=0.323—meaning the average difference was not statistically significant.

The scope of validation remains limited. The study covers a single company's single campaign, on a single platform, targeting a specific parent demographic. It did not address product images or people, nor did it test video or copy. Precise text rendering and complex object interactions also fell outside the study's scope. Generalization across products, customers, and media remains unverified. CTR is a metric for capturing attention and driving link clicks—not for measuring purchases, bookings, sales, or advertising return on investment. It cannot be concluded that AI-generated ads outsell human-made ads in general.

This boundary echoes a 2026 peer-reviewed conference paper from HKBU. That study, a randomized field experiment using over 150 video ads, reported that AI-generated ads outperformed on upstream metrics like CTR and viewer retention, but underperformed human-made ads on conversion. Future experiments will need to shift the evaluation metric from background-image link CTR to bookings, sales, and advertising return on investment.