The Power of the Customer Spec
In 2009, Domino's had a problem that no marketing budget could solve. People said the crust tasted like cardboard, and they said it loudly, in public, at length. Rather than quietly reformulate and run an upbeat “new and improved” campaign, Domino's printed the harshest comments and put them on the walls of its headquarters, changed a 50-year-old recipe, and then built an ad campaign around the criticism itself. The result was a 14.3% increase in U.S. same-store sales in Q1 2010, the largest quarterly jump in the category at the time.
The customer feedback served as a product spec. It told Domino's executives what to change and in what order. The reputation metrics came second, as confirmation the change had landed.
Nearly 20 years later, that same loop has picked up a new reader. The AI assistant answering your customers' questions is summarizing the same kind of text, for every location you operate, every day.
When Everyone Sells the Same Thing
Pick a category. Urgent care, quick lube, apartment leasing, coffee. Now describe what separates the top three providers in your market on product alone.
Most customers can't do it. The oil is the same oil. Urgent care clinics run the same protocols against the same insurance codes. The two-bedroom apartment across the street has the same square footage, amenities, and appliances. In an increasingly commoditized world, what's left is everything wrapped around the end-to-end interaction. Whether the phone got answered. Whether the appointment started on time. Whether the person at the counter knew the answer or had to go find somebody who did.
The research has been pointing here for a decade. Gartner's much-quoted figure is that 89% of companies compete primarily on the customer experience they provide rather than the products they sell. Salesforce, asking the other side of the transaction, puts it at 88% of customers saying experience is as important as products or services, the highest reading since they started tracking it. Qualtrics, in a 2026 study of 20,000 consumers across 14 countries, found something sharper: Good customer service produced higher satisfaction (92%) than good value for the money.
This, of course, is not news to anyone in CX. Improving the experience is the outcome the entire discipline exists to produce, and CX practitioners have been making the business case for it since long before a model was reading the results. What's changed is that the work now has a second audience.
Reading Your Reviews Backwards
Earlier in this series we looked at the evidence that reviews are one of the highest-leverage inputs into AI search. Rating, volume, recency, response behavior—all of it shapes whether your location gets named when a consumer asks an assistant for the best option nearby. Yelp licenses its review data to OpenAI and Perplexity. Google Review text flows straight into AI Overviews, AI Mode, and AskMaps. And it goes deeper than the big two. Healthgrades, Autotrader, Angie's List, and the vertical sites your customers trust in your category are all feeding the same engines.
The star average is the least interesting part of that. What an assistant summarizes is the prose. The wait, the surprise line item on the invoice, the tech who took twenty extra minutes to walk someone through the repair, the leasing office that instantly followed up. That text is a transcript of your operation, published publicly, refreshed daily, and now read by systems deciding whether to put you in an answer.
Most brands run that relationship in one direction. Reviews are a scoreboard, a number to raise, a slide in the monthly business review, something the field gets graded on. Sit through enough of those meetings and the pattern is hard to miss. The conversation is almost always about the number and almost never about the sentences underneath it, which is odd, because the sentences are the part with instructions in them.
Run it the other direction and the same body of feedback becomes the cheapest operational telemetry you will ever own. Thousands of unprompted, timestamped, location-specific accounts of where your journey breaks.
The Questions Are Driver-Level Now
Search used to be a noun. Someone typed “urgent care near me,” got a list of ten, and did the comparison themselves. Assistants took that job over, and the questions people bring them have gotten specific in a way that should make operators uncomfortable. Which clinic has the shortest wait? Which shop is straight about pricing? Which complex fixes things when you call?
Those are driver questions. They map almost one-to-one onto the themes sitting in your review data, because your review data is where the answer comes from. An assistant asked about wait times has no wait-time database to consult. It has your customers' words and the experiences they share.
So the theme dragging your rating and the clause describing you in an AI answer are frequently the same sentence, written by the same customer. An output shaped like “well reviewed for staff friendliness, though several patients mention long waits” is illustrative rather than a live result, but any operator who has read their own one-star reviews will recognize where the second half came from.
From Feedback to a Fix List
This is where most reputation programs stall. The data exists—review text, survey responses, call transcripts, social listening—often in four systems owned by three teams. What's missing is a defensible way to sequence the work. So everything looks urgent, and the loudest complaint wins, which usually means whichever review a regional VP happened to read on a Sunday night.
Journey mapping and driver analysis are how you get out of that:
- Align feedback by journey stage. Awareness, scheduling, arrival, service delivery, checkout and billing, follow-up. Departments will fight over ownership of a theme; stages won't. Stage alignment also catches handoff failures, which is where a surprising share of one-star reviews originate. Nobody was rude. Nothing broke. The information just didn't travel between two teams, and the customer experienced that as incompetence.
- Find the drivers that move the rating. Every category has a handful of themes carrying disproportionate weight, and they're rarely the ones your operators assume. Healthcare tends to be wait time, communication clarity, and billing transparency. Automotive service is usually timeline accuracy and unexpected charges. Multifamily lives and dies on maintenance response. You don't have to guess at yours. Run the correlation between theme mentions and star rating across your own feedback and the drivers will announce themselves. This is also the step where somebody's pet theory dies, which is why it's worth doing before the budget conversation rather than after.
- Weight by impact, then by fixability. A theme showing up in 40% of reviews that barely moves the rating is background noise. A theme in 8% that drags the average a half star is your Monday morning. Use a balanced metric: Mention frequency × rating impact × how fixable it is at the location level. That third term is what keeps a roadmap from turning into a wish list.
- Localize before you generalize. Enterprise averages hide almost everything worth knowing, and the theme killing forty locations can be invisible in the brand-level number. Rank locations driver by driver and the pattern usually resolves into something concrete: A region, a shift, a manager tenure band, or one group of stores that never got trained on a system everybody else got in March.
- Read solicited and unsolicited feedback together. Surveys tell you about the questions you thought to ask. Reviews tell you what people cared enough to volunteer. When the two disagree, that gap is the finding—and the most common version is a survey program timed before the billing statement arrives, which is to say before the experience the customer will write about.
Where the Operational Gaps Hide
Drivers vary by industry, but where feedback exposes operational gaps remains remarkably consistent:
- Healthcare: Wait time dominates negative sentiment, and the root cause is frequently a scheduling template rather than a staffing shortage. When wait-time mentions cluster in a handful of clinics instead of spreading evenly across the system, you're usually looking at a fix that costs nothing and a clause that disappears from your AI summary a quarter later.
- Automotive Service: Complaints cluster on timeline accuracy far more than on repair quality. Customers rarely write that the work was wrong. They write that nobody told them it would take another day. A proactive status message at a defined hour changes the review text without changing a single thing in the bay, and the review text is what an assistant repeats back to the next customer comparing three shops.
- Property Management: Maintenance response time correlates with rating more strongly than any amenity theme. Properties that publish a response-time commitment and hold it will outperform properties that spent the capital budget on a clubhouse renovation.
- Retail and Restaurants: Staff knowledge and consistency across locations drive more sentiment variance than product or price, which is where the Domino's lesson bites hardest. Sometimes the feedback is about the product, and the only useful response is to believe it.
Fix the Operation, Change the Answer
You cannot optimize your way into a good AI answer when the underlying experience is bad.
Schema markup, entity consistency, a clean location page, a well-maintained Google Business Profile—all worth doing, and none of it capable of editing what a customer wrote about waiting ninety minutes past their appointment. These systems summarize the text they're given, and your team writes that text one interaction at a time, at every rooftop, every day.
Which makes operational improvement a visibility strategy. Fix the driver and the review text changes. Once the text changes, so does the summary an assistant builds from it. That loop runs slower than a content sprint, and it's the only one that compounds. Domino's ran it in 2010 with a documentary crew and a national ad buy. You can run it with a driver analysis and a Tuesday operations meeting.
It also changes what's worth measuring. A star average reports the score. The answer text reports the reason, and the reason is the part your operators can do something about. Tracking AI visibility means reading the answers themselves, on the queries your customers use, at the location level, then watching which clauses survive after the operational fix ships. A location whose rating held steady while “long waits” dropped out of its summary got better, and the star average never showed it.
Now Is the Time to Act
Reputation metrics have spent a decade as a marketing scorecard. In an AI-mediated market, they work better as a diagnostic instrument, and the brands using them that way get paid twice: A better operation, and a better answer when a consumer asks an assistant where to go. Every competitor in your category can buy the same schema, the same location page template, the same listings coverage. None of them can buy your customers' sentences.
This is the loop the Reputation platform is built to close. Reviews, surveys, and social feedback land in one place. Experience insights surface the drivers moving your scores and rank them by impact. Location-level scoring shows which rooftops are dragging the brand, and on which specific driver. And the GEO Readiness Report shows what AI engines are saying about you, so you can watch the answer move as the operation does.
The bottom line? Your customers already wrote the fix list. You just have to read it as one.




