Evolving LLM Prompts to Generate Customer Shopping Narratives: AI-Guided Evolution with Cohesive Specificity

In a previous article, we explored how combining large language models (LLMs) with reinforcement learning can create smarter business rules. Building on the AlphaEvolve framework, we extended this approach to tackle a long-standing challenge: generating personalized yet cohesive customer shopping narratives — natural language stories that capture customer habits, preferences, and lifestyle.

These narratives have long been used in luxury industry to nurture 1:1 relationships, but scaling them has typically required manual effort. Our approach uses reinforcement learning to evolve prompts automatically, enabling high-quality, personalized narratives at scale.

By generating narratives at scale, we can apply them not only in segmentation and personalization, these narratives also provide a transparent, human-readable view into how a business understands its customers — enhancing value for both the brands and the customers.

Evolve and Optimize Prompt to Create Balanced Customer Narratives

At the heart of this solution is the LLM-evolved customer narrative prompt generator — a system that creates generative summaries of behavior based on longitudinal purchase activity. To ensure narratives are both unique and cohesive, the prompt evolution framework iteratively refines how prompts steer the LLM. The process starts with the model generating a diverse set of prompt candidates, which are evaluated against an optimization objective that balances individual distinctiveness and segment cohesion. Powered by MAP-Elites and Q-Learning, the system selects the next best prompts, creating a feedback loop that continuously improves both the quality and strategic value of customer narratives.

From Shopping Journey to Customer Narratives

We start with sequences of shopping activities filled with products over time. Each sequence reflects evolving patterns, preferences, and trip missions — whether it’s a quick convenience run, a family meal prep, or a weekend restock.

An illustrated example of a customer shopping journey:

Sep 05 — Dole Organic Gala Apples, Earthbound Baby Spinach Clamshell, Fresh Cravings Pico de Gallo Mild, Horizon Organic 2% Milk, Chobani Vanilla Greek Yogurt 4ct, Blue Bell Mini Ice Cream Sandwiches
Sep 12 — Cuties Clementines 3lb Bag, Fresca Sparkling Citrus 12pk, La Mexicana Fire-Roasted Salsa, Mission Whole Wheat Tortillas, Fresh Jalapeños (Loose), Olivia’s Organics Baby Arugula
Sep19 — Organic Fuji Apples (Bulk), Nestlé Drumstick Cones Variety Pack, Sargento Shredded Mild Cheddar Cheese, Bush’s Spicy Black Beans, Dave’s Killer Bread — 21 Grain, Sabra Roasted Garlic Hummus
Sep 26 — Marie Callender’s Pumpkin Pie (Frozen), Kind Cinnamon Oat Granola, Diet Coke Mini Cans 8pk, Fresh Rosemary Bunch, Green Giant Brussels Sprouts, Eggland’s Best Cage-Free Large Eggs
Oct 03 — Taylor Farms Baby Spinach, Halos Clementines, Drumstick Fudge Cones, Bubly Peach Citrus Sparkling Water, Herdez Jalapeño Lime Hot Sauce, Market Inspirations Hatch Chile Queso
Oct 10 — FAGE Total Greek Yogurt, Organic Girl Spring Mix, Diet Coke Caffeine-Free 12pk, Barilla Rustic Basil Pesto, Mission Flour Tortillas Soft Taco Size, Little Debbie Pumpkin Muffins
Oct 17 — Blue Bell Cookies & Cream Pint, Organic Cilantro Bunch, Kraft Shredded Mozzarella, Old El Paso Red Enchilada Sauce, Uncle Ben’s Mexican Style Rice, Cacique Cilantro Lime Crema
Oct 24 — Noosa Pumpkin Spice Yogurt 4ct, Halos Mandarins, NatureSweet Organic Roma Tomatoes, Cascadian Farm Maple Brown Sugar Granola, Diet Coke 12pk, Colavita Rosemary Infused Olive Oil

Using LLM-generated prompts, we translate these list into natural-language narratives. These narratives describe behavioral traits, such as health-consciousness, culinary interest, pet ownership, indulgent buying patterns, or seasonal behaviors. These semantic summaries become the backbone of downstream applications like:

  • Behavioral segmentation
  • Hyper-personalization
  • Store clustering
  • Journey-aware search and recommendation

Illustrated customer shopping narrative generated by LLM

This shopper maintains a health-conscious lifestyle with frequent purchases of organic produce like apples, spinach, and clementines, alongside staple wellness items such as Greek yogurt, multigrain bread, and cage-free eggs. They balance nutrition with indulgence, regularly adding treats like Drumstick cones, Blue Bell ice cream, and pumpkin desserts. Beverage choices like Diet Coke, Fresca, and sparkling water reflect a preference for low-calorie options, while spicy condiments and seasonal ingredients — jalapeños, hatch chile queso, rosemary — reveal a bold, flavor-driven cooking style. Their seasonal shifts, especially toward fall-themed items, suggest high responsiveness to holidays and promotional cues.

Why Do We Need to Evolve and Optimize Prompts?

While using LLMs to generate customer narratives shows great potential, a one-size-fits-all prompt doesn’t work across different applications. If a narrative is too detailed and granular, it becomes difficult to compare customers or identify common patterns. On the other hand, if it’s too broad and abstract, we lose the ability to capture what makes each customer unique. That’s why prompt optimization is critical: to strike the right balance of cohesive specificity tailoring each narrative to serve its specific purpose — capturing enough detail to preserve individuality, while maintaining enough consistency to support segmentation, personalization, or strategic planning.

Optimized Prompt: Customer Narrative with Cohesive Specificity

Evolving Better Prompts with MAP-Elites + Q-Learning

The goal of a well-crafted prompt is to generate customer narratives that are:

  • Cohesive within clusters — Customers with similar shopping behaviors should be described in a way that highlights shared traits using consistent, interpretable language.
  • Unique across individuals — At the same time, each customer’s narrative should reflect their distinct preferences, routines, and lifestyle.

To strike this balance, we use a hybrid optimization approach based on MAP-Elites + Q-Learning framework in a previous article

🧬 Evolution Objective: Optimizing Spectral Entropy in Narrative Embedding Space

To evaluate the quality of a prompt, we compute an spectral entropy score that reflects the structure of the entire narrative space. This is done by embedding each customer narrative and analyzing how they relate to each other in a graph structure. The entropy tells us how well-structured or fragmented the space is:

  • 🔻 Low spectral entropy: Narratives collapse into overly similar descriptions — useful patterns are lost in sameness.
  • 🔺 High spectral entropy: Narratives are scattered and inconsistent — making it hard to find reliable patterns.
  • Optimal spectral entropy: Narratives are varied but structured — distinct, yet grouped in meaningful ways.

This score is grounded in the spectral analysis of a Narrative graph constructed from narrative similarities — capturing the underlying structure of the behavioral space.

🔁 Evolution Loop: How Prompts Improve

Each iteration of the evolution includes both exploration and exploitation, combining the strengths of MAP-Elites and Q-Learning:

1. MAP-Elites Exploration
We generate a wide range of prompt candidates by varying elements like tone, abstraction level, and phrasing. These prompts are evaluated and archived based on how well they cover the diversity of possible behavioral structures — measured by entropy and behavioral clustering.

2. Q-Learning Exploitation
We then learn from reward trajectory. Using reinforcement learning, we estimate which prompt transformations (e.g., changing from descriptive to abstract language) are most likely to improve the narrative quality. This helps us selectively mutate or recombine prompts to create stronger candidates faster.

In each cycle, we:

  • Start from the best-performing prompt from prior iterations
  • Combine it with a contrasting or exploratory “inspirational” prompt.
  • Apply mutations guided by Q-values (reinforcement signals).
  • Generate new narratives, re-embed them, and recalculate spectral entropy.

This loop repeats across generations, gradually refining prompts toward the optimal zone — where uniqueness and cohesiveness are in balance.

The Outcome: A High-Performing, Purpose-Ready Prompt

After multiple rounds of evolution, we converge on a prompt that consistently delivers:

  • Rich individual-level insight — Capturing what makes each customer distinct.
  • Segment alignment — Enabling clear groupings based on behavior and intent.
  • Prompt generalizability — Usable across personalization, segmentation, and store planning use cases.

By evolving prompts systematically — rather than relying on trial and error — we build a scalable, repeatable way to generate high-quality customer narratives tailored for strategic decision-making.

Measuring the Difference — Initial vs. Optimized Prompt

To understand the impact of prompt evolution, we compare narratives generated from the initial prompt and the optimized prompt. Both prompts were applied to the same illustrated customer journey, but the results differ significantly in clarity, focus, and behavioral relevance.

📚 Narrative from Initial Prompt

This customer regularly stocks up on essentials like 40-pack purified water (Brand A), purchased multiple times across months. Their baskets often include tortillas, shredded lettuce, cheddar cheese, salsa, and charro beans — suggesting frequent taco or Tex-Mex-style meals. Eggs are purchased in large volumes across multiple trips, alternating between Brand A and Brand B.
Fruits and vegetables appear frequently — bananas, sweet potatoes, tomatoes, and greens — alongside indulgent treats like cherry pie, muffins, donuts with sprinkles, and specialty ice creams. Several items repeat week to week, indicating routine-driven shopping.
There’s also a recurring presence of ready-made or convenience foods like frozen chimichangas, biscuit sandwiches, ramen in multiple flavors, and pre-made ravioli with larger basket includes gourmet cheese, parsley, ground meats, Gatorade in three flavors, and a variety of breakfast cereals. Pantry basics are consistently replenished — mayonnaise (both Brand C and Hellmann’s), multiple bacon types, butter, and pasta sauces. Brand C products dominate across dairy, frozen foods, seasonings, and paper goods, suggesting strong retailer brand preference. Seasonal shifts are evident too, with pie crusts, whipped toppings, eggnog, and stuffing mix appearing near the holidays.

📝 Narrative from Optimized Prompt

This shopper consistently purchases produce staples like premium bananas, sweet potatoes, green onions, and Roma tomatoes, alongside core meal items such as Brand C shredded cheese, tortillas, and salsa — suggesting a pattern of preparing Tex-Mex–inspired family meals. There is strong brand loyalty to Brand C across categories, including dairy (lactose-free milk, whipped topping), pantry (olive oil, egg nog, spices), and personal care (mouth rinse, multivitamins), reflecting both trust and convenience. Occasional indulgences like Brand C’s artisan ravioli, Blue Bell ice cream, and bakery treats such as muffins and turnovers highlight a preference for comfort foods with elevated touches.

What Changed — and Why It Matters

Through multiple generations of refinement, the optimized prompt produces narratives that are more semantically consistent, abstracted away from specific product names or brand names, and aligned to behavioral signals — exactly what’s needed for applications like segmentation, journey modeling, and personalization at scale. We use LLM Evaluator to conduct a comparison between the narratives from initial prompt vs. optimized prompt below:

Narratives Comparison using LLM Evaluators

Technical Deep Dive: Optimizing Prompts via Program Evolution

Inspired by AlphaEvolve, our prompt optimization strategy is a custom framework that evolves better prompts by learning how well each one structures customer narratives. Instead of hand-tuning prompts, we treat each prompt as a “program” and evolve it using a hybrid of MAP-Elites and Q-Learning, guided by a measurable signal of quality: semantic entropy.

Initialize prompt archive and Q-learning memory
For each generation:1. Sample customer inputs
2. For each prompt in the population:
a. Generate narratives using the prompt and customer inputs
b. For each narrative:
- Break text into chunks
- Embed each chunk into a vector
- Use RBF kernel to compute weighted average
- Aggregate into a single narrative embedding

c. Compute similarity graph of narrative embeddings
d. Calculate spectral entropy of the graph
e. Update MAP-Elites archive with prompt and entropy score
f. Update Q-values for prompt selection

3. Select top-performing prompt and a few diverse inspirations
4. Mutate and recombine prompts to create next generation

Return prompt with optimal entropy balance

Measure Narrative Quality: Spectral Entropy

Each customer narrative is segmented into short, overlapping chunks — such as sentences or semantic phrases — which are embedded into a vector space to capture local meaning. These chunks are then aggregated into a single latent vector per narrative using a kernel density method tuned for global semantic density. We compute pairwise similarity across all narrative vectors using a radial basis function (RBF) kernel, construct the resulting similarity graph’s Laplacian, and analyze its spectral distribution. The spectral entropy of this distribution quantifies the diversity and structure of the narrative space, serving as the reward signal for each prompt.

Prompt Evolution with MAP-Elites + Q-Learning

We treat each prompt as a small program that turns raw inputs (like shopping history) into natural language narratives. To evolve better prompts, we use a two-part process:

MAP-Elites: Explore Diverse Prompt Styles
We project and group prompts into cells based on their embeddings, where each cell represents a different style or structure. In each cell, we keep the top-performing prompt (the “elite”) based on narrative quality — measured by semantic diversity. This encourages broad exploration across different prompt types.

MAP-Elites promotes exploration across varied styles and abstraction levels.

Q-Learning: Improve Within Each Style
Within each cell, we use Q-Learning to track which prompts perform best over time. Each prompt acts like an “action” taken in a given “cell” or state. The system learns which prompts to favor and how to evolve them through mutation or recombination, focusing on patterns that consistently perform well.

Q-Learning drives exploitation of the best-performing patterns.

Together, they help evolve prompts that generate narratives with cohesive specificity — rich and personal enough for insight, but structured enough for comparison and planning at scale.

Below is an example of the initial prompt and its evolved version after N iterations:

Initial Prompt

Create a paragraph analyzing the shopper’s unique purchasing patterns using the provided chronological shopping data. Identify key recurring items, distinct brand preferences, and lifestyle indicators such as dietary habits or household needs. Highlight any noticeable frequency trends or shifts in preferences, ensuring observations are insightful, precise, and free of redundancy. The analysis should focus on actionable insights that reflect the shopper’s evolving journey and prioritize meaningful, distinctive elements of their habits.

Final Prompt After Evolution

Craft a succinct, three-sentence analysis that highlights the shopper’s unique purchasing patterns using the provided chronological shopping data. Focus on recurring products, brand affinities, and key lifestyle choices such as dietary preferences, household habits, or gourmet interests. Ensure the observations are precise, impactful, and tailored to reflect meaningful trends in the shopper’s evolving journey, avoiding filler jargon or extraneous details.