Seven months after OpenAI began selling ads inside ChatGPT, media buyers are still wrestling with a basic question: how do you measure them? Conversations with agencies suggest the platform has made progress, but not enough to move most budgets beyond experimentation.
The measurement mismatch
Some campaigns are delivering leads that never show up in the platform. Accuracast ran a U.K. lead-generation campaign on ChatGPT and watched form fills arrive, yet the dashboard still showed zero conversions two weeks later. The same agency tracked about 20 clicks where ChatGPT reported around 100, and OpenAI support could not explain the gap.
Not every buyer sees the same issue. AdRoll, Jellyfish and one U.S. independent agency said they had not experienced click discrepancies. But Web Guide Partner sees conversions that do register arrive 24 to 36 hours late, compared with a few hours on Google, and there is no demographic or prompt-level data in reports.
Trust is the bottleneck
Part of the problem is structural. Advertisers who want better measurement must sign OpenAI’s tracking terms and switch on conversion tracking. Many legal teams are refusing, citing liability concerns and unease about what could happen with customer data—even when few can point to the exact clauses they dislike. One independent U.S. agency said nearly all of its clients, across every vertical, are holding back.
OpenAI has also signaled it will not negotiate red lines, leaving advertisers with less room to shape terms than they have with Google or Meta, where years of familiarity have built comfort. The same hesitation is slowing custom audience uploads, which require customer lists.
- Conversion data can arrive 24-36 hours late, or not at all
- Sales tracking often shows zero despite confirmed leads
- Reports lack demographic and prompt-level detail
- Legal teams remain wary of OpenAI’s data terms
The catch-22
The result is a measurement loop: OpenAI needs reliable conversion data to improve tracking and optimization, but the advertisers with the most data are the least willing to share it. Until that changes, many campaigns are still optimized toward clicks or impressions rather than cost per acquisition or return on ad spend.
Buyers also describe unpredictable delivery. Accuracast saw one campaign serve thousands of impressions one week and fewer than 100 the next, with “low bid” warnings that told it to pay more without saying how much—and then disappeared.
“The biggest issue we’re seeing is around inventory — it isn’t predictable, and there’s definitely room to serve more,” said Farhad Divecha, group CEO of Accuracast.
What buyers can do now
Mile Marker’s Monica Shukla, VP of biddable centre of excellence, suggests building an independent measurement framework before testing, then using it to define KPIs rather than accepting platform-first metrics at face value. Mobile measurement partners such as AppsFlyer and Adjust are a welcome step, though verification still mainly confirms ads were seen, not necessarily that they worked.
The practical outcome shows in budgets. One executive described clients spending $10 million a month on Google and less than $100,000 a month on ChatGPT, despite wanting to invest more. Accuracast keeps clients at test budgets of £10,000 to £20,000 instead of scaling to £100,000 until the numbers are verifiable.
OpenAI declined Digiday’s request for comment.
Why this matters for marketers
ChatGPT ads matter because the platform sits at the center of a shift from search results to AI-assisted answers. But for performance marketers, a new channel only becomes real when it can be optimized. Until OpenAI closes the gap between what advertisers see and what the platform reports, it will remain a small, experimental line in media plans.
A useful mental model is “trust before scale.” Marketers should treat ChatGPT ads like an emerging channel: run small, independently tracked tests, lock KPIs in advance, and increase spend only when third-party verification and conversion data align.
Source: Digiday




