Case study
Designing personalized Reasons to Watch for HBO Max
A synopsis tells you what a title is. It never tells you why it is for you.
Reasons to Watch replaces the synopsis with a 135-character pitch written for one viewer — and the interesting part was never the model. It was deciding what the system was allowed to claim.
The surface
Succession
Drama · Series
0 / 135 characters
Both strings are real, straight out of the POC. The personalized one is written for a viewer who liked The Righteous Gemstones — and it is the shorter of the two.
01 — Problem
One reason per segment scales. It cannot explain one person.
What we were working from
- +4%clarity on why a title was recommended
- 17affinity groups covering the whole audience
Reasons to Watch appears when a viewer focuses on a title. Instead of repeating the synopsis, it argues for the title. Internal research had already shown the pattern helped: it lifted clarity around why something was being recommended by 4%.
The scalable version of that idea assigned every viewer to one of 17 affinity groups and wrote one piece of copy per group. It was practical for cost and for a catalog this size. It also flattened taste. The same person can want a prestige drama, a competitive cooking show and a comfort sitcom in one week, for three completely unrelated reasons.
Segmented copy · The Gilded Age
From Julian Fellowes, this sprawling period drama chronicles the great conflict between old and new in New York's glittering Gilded Age.
A gripping portrayal of wealth, power, and the societal change in old New York through the lens of layered, elite families.
Ambitious women navigate intricate relationships and societal pressures in a fiercely stratified 19th-century New York.
Witness the sweeping transformation of New York through the eyes of powerful families during a dramatic social upheaval.
The brief I set myself: write for one viewer rather than their segment, without breaking the 135-character limit, Max's voice, or the spoiler rules.
02 — Insight
A name-drop looks personal. A connection has to be earned.
Two signals were available, and they are not interchangeable. Watch history is broad and behavioral — it shows what someone puts on. Explicit likes are narrow and endorsed — they show what someone would defend. The first model I built used history alone, and it name-dropped titles a viewer had merely watched, sometimes titles they had abandoned halfway.
That version tested worse than the generic copy. Not because the writing was weaker, but because a wrong name-drop is louder than no name-drop. It reads as a system claiming to know you and getting it wrong.
So: history infers the pattern, likes earn the reference. A previous title may be named only when the analyst scores the connection above 85 — a specific, defensible link, not "both are dramas."
Personalization was never the name-drop. It was the explanation of fit.
03 — Decisions
Three decisions turned a prompt into a system.
Write a pitch, not a summary.
The first prompts produced accurate plot summaries. They described the title correctly and gave no one a reason to choose it. I reframed the task three times before it landed — each reframing was a real rewrite of the brief, not a wording tweak.
- 135-character summary
Read like the synopsis it was meant to replace.
- Recommendation
More directional, still too broad to be persuasive.
- Reason they could like it
Closer. Still hedged.
- 135-character pitch
Chosen. "Pitch" is the only one of the four that implies you have to win someone over.
Use history for patterns. Use likes for permission.
The threshold is a dial, not a law. Set it lower and the copy references familiar titles constantly and cheapens them; set it higher and it almost never earns one. 85 was where the references stopped feeling automatic in my test set.
One agent was doing too many jobs.
A single prompt had to research the title, infer taste, write the copy, enforce the product rules and judge its own work. The output stayed generic and quietly missed specifications, and when it failed there was no way to tell which job had failed.
04 — Behavior
Every sentence traces back to the decision that made it.
Splitting the work was not about accuracy. It was about diagnosis. Once each stage returned a structured handoff, a weak sentence had an address: a bad connection is an analysis problem, flat language is a writing problem, a broken character count is an editing problem. Each one gets tuned without disturbing the others.
The four agents, in full
Design engineering prototype · trace one output
Step through a real run for Succession, written for a viewer who liked The Righteous Gemstones.
- 01
Pattern Analyst
Reads the history as family-power drama with a vicious comic streak, then looks for the one title it can defend naming.
affinities Emmy-winning prestige drama · provocative, boundary-pushing comedy overlap family power struggle, vicious comedy, inheritance as weapon new corporate rather than religious setting linked The Righteous Gemstones — connection 91
- 02
Blurb Writer
91 clears the gate, so The Righteous Gemstones may be named. Three drafts, deliberately different angles.
A mood A family that mistakes cruelty for intimacy, and jokes for love. B connection Love the chaos of The Righteous Gemstones? Its corporate cousin is a vicious fight for Daddy's love… C image A birthday party staged as a hostile takeover.
- 03
Editor
Applies the product rules. Draft B keeps its connection and gains the italic title reference; it was already inside 135.
changes italicised The Righteous Gemstones; no other change needed
- 04
Critic
Scores each on creativity and informativeness. B wins on informativeness — it is the only one that says what the show is as well as how it feels.
A 4 / 2 B 4 / 5 C 5 / 2
Final reason
0 / 135 characters
This walkthrough is a recorded run, replayed deterministically. A live one is at the end of the page.
The editor is the design system
Most of what makes the copy sound like Max lives in one agent. These are product rules, not style preferences — several exist because legal or brand would reject the alternative, and every one of them started as something the model kept doing wrong.
- Length135 characters, after editing
- NamingNever "this show", never the title itself
- AwardsNo Emmy, Golden Globe or HBO — use "critically acclaimed"
- ReferencesFull title, italicised, one maximum
- PunctuationNo exclamation points. No em dashes.
- PresumptionNo "If you enjoy…", no "You love…"
- RhythmNo three-adjective stacks. No one-to-three word sentences.
- Banned words — with one exceptionNever "grit" or "chess". Never "delicious" unless the title is genuinely culinary.
That last exception is the one I point at when someone asks whether the rules are real. The pitch for Selena + Chef ends "…all her sassy, competitive fire in a delicious 20-minute duel." The word survives editing because the show is about cooking. A rule with a working exception is a spec; a rule without one is a filter.
05 — What changed
"Personalized" was too easy a bar to clear.
Everything above could be true and the feature could still not be worth shipping. Personalized copy will always beat generic copy when the reader knows which is which. The only honest question was whether it beat the default when nobody knew which was which.
I worked with design technologist Travis Swan to build an internal evaluation tool. Participants entered their own watched and liked titles, then compared default copy against generated copy blind. After each choice the tool asked why — plot and theme balance, a familiar connection, attention, or a sense that the writer understood the title.
Design engineering prototype · blind comparison
Which would better help you decide whether to watch The Gilded Age? Sources stay hidden until you pick.
The real tool asked why before revealing either source. That order matters: reasons given after the reveal are rationalisations.
The tool took weeks to build, and I did not want to wait weeks for signal. While it was in development I ran the same comparison asynchronously in a spreadsheet — uglier, slower to analyze, and it answered the question a month earlier.
Personalized output · three real examples
Talented chefs battle for the chance to take down Chef Bobby Flay.
As if Monica Geller from Friends got her own cooking show. All her sassy, competitive fire in a delicious 20-minute duel.
A bitingly funny drama series exploring themes of power and family through the eyes of an aging media mogul and his four grown children.
Love the chaos of The Righteous Gemstones? Its corporate cousin is a vicious fight for Daddy's love where the jokes cut like glass.
The creators of Game of Thrones bring another epic series of power, politics, and family feuds. It's Succession with dragons.
What the prompts lost along the way
The agent guide keeps every rejected phrasing, because each one failed for a reason worth remembering.
- "You are a blurb writer"
Accurate and impersonal. Produced catalog copy.
- "You are a friend of the user"
Personal, and it started using slang. Max does not use slang.
- "You are a friend who is really good at making recommendations"
Chosen. Warmth without the register slipping.
- Critic scores "Personalized" and "Emotional Appeal"
Killed. The writer produced three near-identical drafts, so the critic had nothing to choose between. Fixed upstream instead, by demanding three distinct angles.
- Longer, more detailed agent prompts
Killed. Past a certain length the copy reads as constrained — it sounds like something obeying rules.
06 — Results
It proved what useful looks like. Not that it moves anyone.
Qualitative feedback landed on three strengths, consistently: the copy matched a mood the viewer was already seeking, it was genuinely enjoyable to read, and the connection to a past favorite made the recommendation easier to believe. The work went on to inform the team's exploration of personalized copy in single-content promotions.
What participants said
- Mood"Intriguing." "Very me."
- Voice"Hilarious, fun to read, and fairly accurate."
- Trust"I like how it draws the connection."
What stayed unproven
This was an internal POC measuring copy preference and perceived relevance. It did not establish that personalized Reasons to Watch increase playback, shorten time to selection, or improve retention. Preference is not behavior, and I would not present it as such.
If it shipped, three metrics would settle it. I would build them with engineering before launch rather than after, since none of them can be reconstructed from a copy test.
-
Play Rate # of plays titles focused
Did the reason earn a watch?
-
Time to Selection seconds browsing sessions
Did it help them decide faster?
The third is the one I would watch hardest: the share of sessions that end without playback. It is the inverse of the first, and it is the only one that catches a reason that persuades someone to open a title and then disappoints them.
Mini prototype · live generation
Pick at least one title you have watched. The four agents run for real against a random HBO Max title — the same pipeline, live.
Your watch history
Titles you would actually defend (optional — these earn a name-drop)
- ●Pattern Analyst
- ●Blurb Writer
- ●Editor
- ●Critic
Reason for
0 / 135 characters
Runs on Claude through a small Cloudflare Worker — four calls, one per agent, so each handoff appears the moment it exists. Rate limited, because it is a portfolio, not a service.
AI did not make the copy personal. The system did. Better output came from better product decisions: choosing signals I could defend, separating failure modes so each had one owner, writing the constraints down, and refusing to evaluate the result without hiding where it came from. The model supplied the language. The system supplied the judgment.