Your store already gets mystery shopped. You just don't see the results.
Every week, real customers submit a lead, wait, get a generic auto-reply, call in, get bounced to voicemail, and go buy somewhere else. That's a mystery shop. You just never get the scorecard.
A monthly mystery-shop program is you running the same test on purpose, on a schedule, with a form you can act on. Not to catch people. To find the four or five spots where your process leaks every single month and never gets fixed because nobody owns them.
The version most stores try goes like this: someone submits a fake lead, gets mad at the response, forwards the email to the whole team with "THIS IS UNACCEPTABLE," and nothing changes. That's not a program. That's a mood.
Here's how to build one that produces assigned work.
Shop four doors, not one
Most dealership mystery shopping only covers the internet lead. That's the easiest to fake and the least representative. A customer can enter your store four ways, and each one fails differently.
1. Internet lead (third-party and website form) Submit from a real Gmail account with a real phone number you'll answer. Use a plausible name and a real VIN off your lot. Ask one specific question that requires a human to answer — not "what's your best price."
"Hi, is the silver 2022 Highlander with 41k still available? I've got a 2018 Accord to trade and I'm out of state until Thursday. What would delivery look like?"
That question tests three things at once: vehicle-specific response, trade acknowledgment, and whether anyone handles a logistics wrinkle instead of ignoring it.
2. Phone-up Call the main line at 8:15am, at 12:40pm, and at 7:20pm on different shop months. Time of day is the variable most managers never test. Your 6pm coverage is a different dealership than your 10am coverage.
3. Trade appraisal form Fill out the "value your trade" tool on your website. Half the stores I've seen never route this anywhere. The lead lands in a folder. Or it auto-replies with a range so wide it's useless and nobody follows up.
4. Credit application This is the one nobody shops, and it's the highest-intent lead in the building. Submit a full credit app with a soft-spot income and a mid-600s profile. Then start the clock. A credit app that sits three hours is a customer who applied somewhere else.
Four shops. One a month each, rotating time of day and rotating which rep or which source. Twelve months, forty-eight shops. That's enough volume to see patterns instead of anecdotes.
Score the same nine things every time
Long rubrics don't get filled out. Keep it to items you'd actually coach on. Every item is yes/no — no five-point scales, because nobody agrees on what a 3 means.
- Speed to first human contact. Not auto-reply. A human. Log the minutes.
- Channel match. They gave a phone number. Did anyone call, or just email?
- Named person. Did the response come from a person with a name and a direct line, or "Internet Sales Team"?
- Vehicle-specific. Did they reference the actual unit, or send a generic list?
- Answered the question asked. The out-of-state Thursday thing. Did anyone touch it?
- Trade acknowledged. They mentioned an Accord. Did anyone ask about it?
- Appointment ask with a specific time. "When works for you" doesn't count. "Can you do Thursday at 5:30 or is Saturday morning better" counts.
- Second attempt within 24 hours if no answer to the first.
- Different channel on attempt two or three. Called twice and stopped, or called, texted, emailed?
Nine boxes. Takes four minutes to score after the shop. And be honest about what you're measuring — this isn't a personality test. You're measuring whether the process ran.
Record the actual artifact
Screenshot every email. Save the text thread. Pull the call recording. When you sit down with a rep, "you didn't ask for the appointment" is arguable. Playing forty seconds of the call where they said "just let me know" is not.
This is where an internet lead audit stops being opinion. You're not describing what happened. You're showing it.
Keep a folder per month with the artifacts, the scored form, and the timestamps. It takes ten minutes and it's the difference between a coaching conversation and an argument.
The failure-to-fix conversion
Here's the part that makes it a program instead of an exercise.
Every failed item goes into one of three buckets, and each bucket has a different owner.
Person failure. One rep, one time, knew the standard, skipped it. Owner: their manager. Fix: a one-on-one with the recording, and a re-shop of that rep within 30 days.
Process failure. Nobody could have passed because the process doesn't support it. Trade form routes to an unmonitored inbox. Credit apps only get checked twice a day. Phone tree dumps 7pm callers to a voicemail box nobody owns. Owner: whoever controls that system. Fix: a change with a date.
Tool failure. The CRM template auto-sends a generic response before a human touches it. The lead source drops the customer's comment field. Owner: you or your CRM admin. Fix: usually a vendor ticket, and you have to chase it.
Most stores mislabel process failures as person failures. You yell at the BDC for slow credit app response when the actual problem is that nobody assigned credit apps to a queue. The rep gets defensive, nothing changes, and you shop again next month and get the same result.
Sorting into buckets forces the honest answer.
One page, three fixes, named owners
At the end of every month, the whole program produces one page:
| Shop | Score | Top failure | Bucket | Owner | Due |
|---|---|---|---|---|---|
| Internet lead – 3rd party | 5/9 | No trade acknowledgment | Person | Sales mgr | 10th |
| Phone-up 7:20pm | 3/9 | Rolled to general VM | Process | Ops | 15th |
| Trade form | 2/9 | No human contact in 48h | Process | BDC mgr | 15th |
| Credit app | 6/9 | 4h 20m to first call | Tool | CRM admin | 20th |
Three fixes maximum per month. If you assign nine, you get zero. Pick the three that cost the most deals and let the rest wait for next month.
Read the closed items out loud at the next manager meeting. "Last month we said the 7pm phone tree would route to the BDC cell. Did it?" If the answer is no, it carries forward with the same owner and the same name attached. That's what lead handling accountability actually looks like — not a memo, a list with names that gets read back.
What changes after three months
Month one, your scores will be bad. Expect 3s and 4s out of 9. Don't panic and don't hide it.
Month two, the obvious plumbing gets fixed. Trade forms start routing. Somebody owns the after-hours line.
Month three is where it gets useful, because the plumbing failures are gone and what's left is behavior. That's your real BDC quality assurance signal — the process runs, and now you're looking at whether people execute inside it.
By month six you'll be shopping to confirm rather than to discover, and you'll know within two shops when something has drifted.
Four shops a month. Nine yes/no boxes. Three assigned fixes with dates. That's the whole thing.
If you're already scoring live calls and lead responses against a rubric, your mystery shops become the calibration check — a known input, a known correct handling, and proof that what you're scoring at scale matches what actually happens when a stranger contacts your store.