Dynamic Learning and Optimal Advertising Mechanism for LLM Platforms

Ali Makhdoumi
Azarakhsh Malekian
2026
Google Scholar

Abstract

We study advertising in conversational LLM platforms, where the platform gradually learns a user's preferences before deciding when and what ad to offer. We model this as a stochastic control problem in which the platform balances the value of information acquisition against the risk of user departure. Although the optimal policy that maximizes welfare admits a threshold structure, computing it exactly is infeasible without strong assumptions. We develop an approximate threshold policy based on upper and lower bounds on the continuation value and prove explicit performance guarantees that scale with uncertainty, learning dynamics, and heterogeneity of the advertisers. Furthermore, we extend this framework to revenue maximization, where advertisers hold private valuations. We establish a "virtual welfare equivalence" in this dynamic setting, demonstrating that the revenue-optimal incentive-compatible mechanism is implemented by applying the welfare-maximizing policy to advertisers' properly defined virtual valuations. This mechanism generalizes classical optimal auction theory to environments where the information structure evolves endogenously through user interaction, enabling us to derive approximately revenue-optimal mechanisms.

×