Saeed Alaei

Saeed Alaei

I am a research scientist in the market algorithms team in Mountain View. I received my Ph.D in Computer Science from University of Maryland - College Park. I was a post doctoral research associate at Cornell University prior to joining Google. My research interests include mechanism design/algorithmic game theory, combinatorial and convex optimization, and online algorithms. The focus of my research is on developing general algorithms and techniques for optimization problems involving strategic agents arising in online markets.
Authored Publications
Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
Optimal Mechanism Design with Post-Allocation Signal
Azarakhsh Malekian
Ali Makhdoumi
Ali Daei Naby
2026
Preview abstract We study the design of the optimal mechanism in a setting with one buyer and one item in which the seller receives a binary signal correlated with the buyer's type after the allocation decision and can condition the payment on the signal realization. The optimal mechanism depends on the correlation between the signal and the buyer's type through the conditional probability of the signal given the type, which we call the \emph{buyer's belief function}. Under the standard monotone likelihood property (MLRP), we establish that the curvature of the buyer's belief function plays a central role: full surplus extraction is achievable if and only if the belief function is convex in the buyer's type. In contrast, when beliefs are concave, it is optimal to fully allocate to the buyer if and only if her type is outside an intermediary range of types (corresponding to types with a higher correlation with the signal). We further establish that, in the case of concave beliefs, the optimal mechanism can be implemented as a \emph{state-contingent posted payment} offered on a take-it-or-leave-it basis. We further provide a structural characterization of the optimal deterministic mechanism under general belief functions: the state space partitions into several regions, depending on the curvature of the belief function, and in each region, we either have full surplus extraction or a state-contingent posted payment. Finally, we show how our results can be extended to a setting with multiple buyers. View details
Incentivizing Data Collaboration: A Mechanism Design Approach
Ali Makhdoumi
Azarakhsh Malekian
Ali Daei Naby
2026
Preview abstract We study the problem of incentivizing strategic agents to truthfully contribute high-quality data in collaborative learning settings, where each agent benefits from improved estimation based on others’ data. Each agent privately observes the quality of their data, and agents may misreport it if not incentivized properly. We cast this problem with a Bayesian mechanism design framework in which the platform aims to find the optimal data-sharing mechanism that jointly determines allocations and payments to maximize both estimation accuracy and platform revenue. We prove that the optimal mechanism that incentivizes truthful reporting takes the form of a \emph{personalized threshold and pricing} mechanism, in which each agent is allocated the learned estimator if their reported quality exceeds a (personalized) threshold and is charged a price based on the relevance of other agents' data in the learning task. We analyze this mechanism in a canonical Gaussian mean estimation task, derive a closed-form solution to the optimal mechanism, and highlight how data correlation affects the mechanism. We further extend the model to allow agents to exert costly efforts to improve their data quality before collaboration. We show that ''free-riding'' is mitigated as the optimal data-sharing mechanism induces a supermodular game: each agent is incentivized to exert more effort when others exert more. Finally, we show that equilibrium efforts form a complete lattice, and in the highest-effort equilibrium, each agent increases effort as others' data becomes more relevant in the learning task. View details
Preview abstract We study advertising in conversational LLM platforms, where the platform gradually learns a user's preferences before deciding when and what ad to offer. We model this as a stochastic control problem in which the platform balances the value of information acquisition against the risk of user departure. Although the optimal policy that maximizes welfare admits a threshold structure, computing it exactly is infeasible without strong assumptions. We develop an approximate threshold policy based on upper and lower bounds on the continuation value and prove explicit performance guarantees that scale with uncertainty, learning dynamics, and heterogeneity of the advertisers. Furthermore, we extend this framework to revenue maximization, where advertisers hold private valuations. We establish a "virtual welfare equivalence" in this dynamic setting, demonstrating that the revenue-optimal incentive-compatible mechanism is implemented by applying the welfare-maximizing policy to advertisers' properly defined virtual valuations. This mechanism generalizes classical optimal auction theory to environments where the information structure evolves endogenously through user interaction, enabling us to derive approximately revenue-optimal mechanisms. View details
Inspect or Guess? Mechanism Design with Unobservable Inspection
Azarakhsh Malekian
Ali Daei Naby
The 21st Conference on Web and Internet Economics (WINE) (2025) (to appear)
Preview abstract We study the problem of selling $k$ units of an item to $n$ unit-demand buyers to maximize revenue, where buyers' values are independently (and not necessarily identically) distributed. The buyers' values are initially unknown but can be learned at a cost through inspection sources. Motivated by applications in e-commerce, where the inspection is unobservable by the seller (i.e., buyers can externally inspect their values without informing the seller), we introduce a framework to find the optimal selling strategy when the inspection is unobservable by the seller. We fully characterize the optimal mechanism for selling to a single buyer, subject to an upper bound on the allocation probability. Building on this characterization and leveraging connections to the \emph{Prophet Inequality}, we design an approximation mechanism for selling $k$ items to $n$ buyers that achieves $1-1/\sqrt{k+3}$ of the optimal revenue. Our mechanism is simple and sequential and achieves the same approximation bound in an online setting, remaining robust to the order of buyer arrivals. Additionally, in a setting with observable inspection, we leverage connections to index-based \emph{committing policies} in \emph{Weitzman's Pandora's problem with non-obligatory inspection} and propose a new sequential mechanism for selling an item to $n$ buyers that significantly improves the existing approximation factor to the optimal revenue from $0.5$ to $0.8$. View details
Response Prediction for Low-Regret Agents
Ashwinkumar Badanidiyuru Varadaraja
Sadra Yazdanbod
Web and Internet Economics 2019
Preview abstract Companies like Google and Microsoft run billions of auctions every day to sell advertising opportunities. Any change to the rules of these auctions can have a tremendous effect on the revenue of the company and the welfare of the advertisers and the users. Therefore, any change requires careful evaluation of its potential impacts. Currently, such impacts are often evaluated by running simulations or small controlled experiments. This, however, misses the important factor that the advertisers respond to changes. Our goal is to build a theoretical framework for predicting the actions of an agent (the advertiser) that is optimizing her actions in an uncertain environment. We model this problem using a variant of the multi armed bandit setting where playing an arm is costly. The cost of each arm changes over time and is publicly observable. The value of playing an arm is drawn stochastically from a static distribution and is observed by the agent and not by us. We, however, observe the actions of the agent. Our main result is that assuming the agent is playing a strategy with a regret of at most f(T) within the first T rounds, we can learn to play the multi-armed bandits game without observing the rewards) in such a way that the regret of our selected actions is at most O(k^4 (f(T) + 1) log(T)). View details
×