Amazon Payments applied a multi-objective contextual bandit to a product acquisition funnel and recorded a high single-digit percentage relative lift in final-funnel conversion for one customer population during a seven-week online A/B test. Another population showed no improvement over the existing experience, which the team attributed to content rather than the model. The work matters for business because generative AI has made content production cheap, shifting the bottleneck to selecting the right variation for each visitor.
How Amazon Payments built funnel-wide personalization
The project ran on Amazon SageMaker AI and extended earlier work on generating personalized content with Amazon Bedrock within brand guardrails. The customer journey covered three steps: application start, submission, and approval. The team optimized all three stages together instead of a single click metric. A code repository with a Jupyter notebook, command-line demo and unit tests was published for testing the method on synthetic data.
The system uses Linear UCB, or LinUCB, introduced by Li et al. in 2010, in a contextual form. Each arm maintains running tallies of experience and reward, producing an estimate plus an uncertainty bonus that is large for rarely seen visitor types and small for familiar ones. The team set the exploration parameter alpha to 1.0 as a default, with a practical range of 0.1 to 2.0. Higher values suit a large arm space with limited history, while lower values favor the current best option once evidence accumulates.
The funnel design addresses what the authors call the seesaw problem. Content tuned only for starts attracts a broad audience but can reduce approvals, while tuning only for approvals leaves too little signal because approvals are rare and delayed by days. The solution runs one LinUCB model per stage and combines the three scores with a linear combination using approximately equal weights. Starts and submissions update immediately, while approval outcomes wait for a later batch cycle inside an attribution window aligned to weekly processing.
What this means for AI-driven conversion
For companies operating acquisition funnels, the approach offers a way to personalize without waiting for a test to finish. The bandit serves live traffic while learning, shifts impressions toward stronger arms, and keeps a fraction of traffic for continued testing. Small firms with limited traffic can benefit because patterns learned from one context transfer to similar visitors through a feature vector of behavioral signals. Large firms gain a reproducible serving rule, since UCB selection is deterministic and auditable for every impression.
The limits concern content supply, measurement discipline, and privacy handling. The authors state that a bandit is only as good as its arm pool, and their arms were Cartesian pairings of industry-themed images and benefit-focused taglines vetted as individual blocks rather than every combination. Teams should verify stage weights, alpha settings, and attribution windows before copying the setup. The entity identifier was used only for routing a decision back to the visitor and never as a model input, a constraint worth confirming with any vendor.
A useful confirmation signal will be whether the repository and notebook lead to wider deployments that report funnel-wide gains beyond single-stage clicks. Watch for follow-up results with longer attribution periods and larger pools of reviewed blocks expanded with generative tools. If additional funnels show sustained approval growth without loss in starts, the combination of generation plus bandit selection will have moved from experiment to operating practice.
