Case study Work Worth Doing Property mentorship AI coach clone

Dean's WhatsApp was the product. Then it became the ceiling.

A full teardown of the AI coach we built for his mentorship programme: the architecture, this week's numbers pulled live, and a dated log of everything that broke.

0
Conversations
last 7 days
0
Knowledge gaps
same period
0
Distinct jobs
it now does
0
Days from
kickoff to live
$30-50
Monthly
running cost

Pulled from the running system on 6 August 2026. Usage figures are the trailing seven days, not all-time.

Before
  • Every question in the programme arrived in one man's DMs
  • Beginners and experienced operators got the same answers
  • Call insights lived in recordings nobody went back to
  • The course sat in a platform the coaching never touched
  • Growth meant more of Dean's evenings
After
  • 205 conversations answered in a week without him in the loop
  • Answers change based on whether you own nought properties or twelve
  • Every coaching call becomes a written summary in the member's portal
  • All 138 lessons are searchable inside the coaching itself
  • Growth costs about forty dollars a month
01The bottleneck

The thing his members paid for was the thing he couldn't scale

Dean Reilly runs Parea Living, a short-let accommodation business across London and the Midlands, and teaches other people to build the same thing through a mentorship programme called Work Worth Doing. The course behind it, The Automatic Airbnb System, runs to 138 lessons across seven modules covering housekeeping, mindset, foundations, operations, leads, sales and systems.

The course was never the product though. What members were really buying was access to Dean's judgement, and that arrived one WhatsApp message at a time. Someone would ask whether to take a mortgage or keep saving. Someone else would ask whether to move a portfolio into an SPV before the next tax year. Both questions landed in the same inbox, and both waited for the same person.

This is the shape of almost every expert business we get called into. The bottleneck isn't marketing or delivery. It's that the founder is the only instance of the thing being sold, and there's a hard ceiling on how many hours that instance has. Dean had already hit it. He could either cap the programme, hire coaches who wouldn't sound like him, or find a third option.

A coach who answers in four hours instead of four days is worth more than a coach with better answers.

The brief we took was narrow on purpose. Answer members in Dean's voice, using only what Dean actually teaches, and adjust the answer to where that member currently stands. Everything else the system does now grew out of that, and none of it was in the original scope.

02Why a wrapper fails

A chatbot on top of his course would have been actively dangerous

The obvious build is a week's work. Take the course, chop it into chunks, put them in a vector database, and pass whatever comes back to a language model with an instruction to answer like Dean. Most agencies would ship that and call it an AI coach. It fails in three specific ways, and each failure taught us something about what the real system had to do.

It gives the same answer to two people who need opposite advice

Take a real question from the programme: should I use a limited company or buy in my own name? For someone with no properties and no income from lettings, the answer is that the question doesn't matter yet and there's something more useful to do this week. For someone with nine units and a full-time job, it's a serious tax conversation. A plain retrieval system gives both people the same paragraph, because both queries hit the same chunk of course content. That's how an AI coach ends up giving beginners advice that only makes sense for experienced operators.

So retrieval had to know who was asking. Every member has a profile that builds up over their conversations, tracking where they are in the programme and what they're working on, and that profile changes which chunks get ranked highest before the model ever sees them.

It answers the question instead of doing the job

People don't only ask questions. They ask for a message to send a landlord who's gone quiet. They ask to be prepped for a call in thirty minutes. They ask why the last eight approaches got no reply. Those are different jobs with different shapes of correct output, and a system that treats all of them as questions produces something that reads fine and helps nobody.

The running system now recognises eleven distinct jobs, which we didn't design up front. We built four and the rest came from watching what members actually sent.

It sounds confident about things it doesn't know

This is the failure that matters most and gets discussed least. When retrieval comes back weak, a language model doesn't stop. It fills the gap with whatever is plausible. In our own evaluation the system confidently named specific software and quoted specific percentages that appear nowhere in Dean's material. To a member, an invented figure in Dean's voice is indistinguishable from Dean's actual advice, which makes it worse than no answer.

That's why the build includes a gap detector that logs any answer where retrieval scored badly, and why weak retrieval on a factual question now gets intercepted before it reaches the member rather than after.

The hard part of an AI coach isn't making it answer. It's making it decline.

03The topology

Five lanes, one machine

What started as a chat endpoint is now five independent lanes sharing one knowledge base. The chat lane answers members. The ingestion lane keeps the knowledge current without anyone uploading anything. The call lane turns coaching calls into written records. The onboarding lane builds a new member's entire environment from a signature. The operations lane watches the other four and tells Dean when something breaks.

Lane 01 Coaching chat trigger: member DM on WhatsApp, Telegram or web
gateVerified email
and phone
→
haikuIntent triage
11 job types
→
graphQuery expansion
3,500+ topics
→
vectorRetrieve
7,787 chunks
→
rerankStage-aware
source quota
→
sonnetAnswer in
Dean's voice
→
reviewSecond-pass
draft check
→
replySend in
the DM
scroll →
Lane 02 Knowledge ingestion trigger: cron, 02:00 daily
pullYouTube, reels,
course, podcast
→
whisperTranscribe
audio and video
→
embedBoundary-aware
chunking
→
storeVector DB and
knowledge graph
→
annealWeekly review
of its own answers
scroll →
Lane 03 Call capture and routing trigger: coaching call ends
routeMatch by verified
email
→
attendConfirm against
meeting log
→
branchGroup call?
alias-scored
→
summariseActions and
decisions
→
deliverMember portal
plus a DM
scroll →
Lane 04 Onboarding trigger: signed agreement
buildPortal, invoice,
course access
→
groupPrivate WhatsApp
group created
→
rosterAdded to the
access roster
→
welcomeGuided first
week
→
chaseNudge at 48h,
flag at 96h
scroll →
Lane 05 Operations trigger: cron, every minute
sweepRecover anything
a webhook missed
→
driftRoster diffed
against backup
→
probeHourly health
and key check
→
alertFailure named
by step
→
dashOwner dashboard
and weekly digest
scroll →

The three decisions that matter

A knowledge graph in front of the vector search. Members ask about things using their own words, and Dean teaches them using his. A graph of more than three and a half thousand topics, built from his own material, expands a query before it reaches the vector database, so a question about "getting landlords to say yes" also searches the language Dean uses when he teaches objection handling.

Retrieval that knows the member's stage. Reranking happens after the vector search and before generation, weighted by what the member's profile says about where they are. The same query from two members returns different source material. This is the single change that stopped the system giving beginners advanced answers.

A cheap model in front of an expensive one. Every message hits a fast, inexpensive model first to work out which of the eleven jobs it is and what to actually search for. Only then does the expensive model generate. That keeps the running cost at tens of dollars a month rather than hundreds, and it makes the routing decision inspectable, which matters when something goes wrong.

What we've left out here is deliberate: no prompts, no similarity thresholds, no reranking weights. The shape of the system is worth more to you than our settings, which change every few weeks anyway.

04One query, traced

What actually happens in those forty seconds

Here's a single exchange with the pipeline decisions shown next to it. The member and the details are invented, because publishing real conversations from a live coaching programme isn't something we'd do. The routing, the retrieval behaviour and the timings are real.

WhatsAppMentee C
Mentee C had a viewing yesterday and the landlord basically said yes on the call then went quiet. been two days. do i chase or leave it
Coach Two days isn't quiet, it's a Tuesday. Quiet is a week.

Send one message today and make it about them, not about you. Something like: "Hi Mark, been thinking about the second bedroom since we spoke. If the guarantee is the bit you want to be sure about I can put the numbers in writing so you've got them in front of you." Then stop.

You've already got the verbal. Chasing it turns your position into a weaker one. What you want is to give him a reason to reply that costs him nothing.
Pipeline38.4s total
▸gate verified0.2s
▸triage closing_playbook, haiku1.1s
▸profile first deal, 0 units0.1s
▸expand +6 graph terms0.4s
▸retrieve 24 candidates1.8s
▸rerank 24 → 4, stage-weighted0.3s
▸quota sales 2 · calls 20.0s
▸patterns asked twice before0.2s
▸generate sonnet, extended26.7s
▸review draft pass, 1 edit7.4s
▸gap check no gap logged0.1s
▸deliver sent, logged0.1s

Two lines in that trace are worth your attention. rerank 24 → 4, stage-weighted is where a beginner and an experienced operator stop getting the same answer. patterns: asked twice before is memory across conversations, which exists because one member asked us the same question more than thirty times and got a patient fresh answer every time.

05The numbers

Pulled from the running system, not from a slide

Everything below came out of the live system on the morning of 6 August 2026. We're showing the trailing seven days rather than all-time totals, and that decision is worth explaining, because it's the kind of thing that usually gets quietly hidden in a case study.

The all-time counter reads higher. We checked it against the weekly series and found the conversation log only holds records from late July, which is consistent with a storage eviction we've already documented and fixed. So the bigger number exists, and we can't stand behind it. The seven-day figures we can.

Adoption, week over week

trailing 7 days vs the 7 before

conversations37
→ this week205
unique members4
→ this week13

Conversations up 5.5x, active members up 3.2x, one week to the next. The jump follows a roster repair and the switch of the course reference described in log entry 01. 13 of 29 people with access used it this week.

Zero knowledge gaps in the logged window

The system scores its own retrieval on every message and logs a gap whenever it answers from weak material. Across the 205 conversations logged this week, it recorded none. That's a claim worth being suspicious about, since a silent detector and a working one look identical from outside, so here's the evidence it's alive: the same detector has flagged 87 queries over the life of the system, routing twenty to Dean as missing content and closing another twenty on its own once the material existed.

The honest reading is that the knowledge base finally covers what members ask, which it didn't in June.

It stopped being a question-answering bot

This is the chart we find most useful, and the one nobody publishes. In May the system did three things and 99 percent of that was answering questions. Here's the same measurement now. Ten of the eleven job types it recognises got used this week, and they add up to the same 205 conversations in the chart above.

question89
outreach_advice51
closing_playbook18
pre_call_coaching15
bottleneck_diagnostic15
navigation8
scaling_strategy4
content_script3
call_recap1
content_angle1

Nearly a third of all traffic is now outreach and closing help, which is work Dean used to do on calls. Another fifteen messages were classified as pre-call coaching, someone getting ready for a conversation they're about to have. That's the moment coaching is worth the most and the moment a human coach is least likely to be free.

Our own quality scores, including the bad one

We grade the system on five dimensions against a standing set of cases, with a separate model doing the grading against a written rubric. Every time we find a new way for it to fail, that failure becomes a permanent case. Publishing this is unusual and we think it should be normal. Two things make it worth your time: the low score, and that we re-ran the whole thing against the live system this morning, so you're seeing movement rather than a snapshot.

Quality evaluation, before and after

scored 1 to 5 by an independent model · 11 June vs 6 August 2026

voice · jun4.62
→ aug4.63
register · jun4.31
→ aug4.37
focus · jun4.15
→ aug4.84
usefulness · jun3.85
→ aug4.58
grounding · jun3.08
→ aug3.58

Usefulness up 0.73, focus up 0.69, grounding up 0.50. Intent misclassifications halved, two down to one.

One caveat you should have, because it's the first thing we'd attack if someone showed us this chart: the June run covered 13 cases and the August run covered 19. We added six, and most of them are deliberately nasty grounding tests written from the failures in log entry 06. So the case set got harder between the two columns, which makes the grounding move more meaningful and every other column less exactly comparable. We're showing it rather than quietly re-running June's 13.

Grounding is still the weakest dimension at 3.58. On one case the system produced sensible weekly activity targets that weren't the specific targets Dean teaches, which is a milder version of the same mistake. The grounding gate fired and corrected a figure on three cases in this run, so the mechanism works. It doesn't catch everything yet.

06The engineering log

Everything that broke, and what we did about it

This section is the reason the page exists. Anyone can show you an architecture diagram. What tells you whether a system is actually being run is the record of it failing, and the record of somebody noticing. These are real entries, dated, in the format we keep them in internally.

Read them as a checklist for your own build. Five of the six are mistakes you'd make too.

2026-06-19

01 · The course reference held 171 characters

symptomAnswers about course content were vague and kept falling back to general property advice, despite the course being connected.
causeThe workspace page the system read as its course reference contained six module rows and nothing else. Everyone involved, us included, had assumed the module titles were pointers to content. They were the content.
fixExported the real course and built a full-text mirror: 138 lessons, roughly 988,000 characters, verified block for block. Retrieval was re-embedded against it.
what the system was reading 171 characters
what the course actually was 988,000 characters
lessonMeasure the size of every source you connect, on day one, and put the number somewhere you'll see it. A source that resolves without erroring can still be empty.
2026-07-18
2026-07-21
2026-07-27

02 · The access roster silently halved, three times

symptomPaying members messaged and got no reply. No error was raised, because from the system's point of view they weren't members.
causeThe roster lived in the hosting platform's managed key-value store, which evicts entries under memory pressure. It evicted some and kept others, so the roster stayed readable and became wrong. It went from 26 entries to 16, then 26 to 23, then 24 to 10.
fixMoved the source of truth to a persistent volume with the key-value store as a cache. Removals now write a tombstone so a restore can't resurrect somebody who was taken off. An hourly job diffs live against backup and alerts on any drift.
lessonNever let a managed cache be the source of truth for anything that controls access. Partial data loss is worse than total loss, because total loss sets off alarms.
2026-08-04

03 · The call webhook had been dead since launch

symptomCall summaries were arriving in member portals between thirty and ninety minutes after the call rather than within a few minutes. Nobody had complained, which is why it took a while to notice.
causeThe webhook that fires when a call ends was rejecting every request over a signature check and had been since the day it went live. Twenty-one of the twenty-five calls in the period had been picked up instead by a recovery job that sweeps for anything missed, which is why nothing was actually lost and why nothing looked wrong.
fixCorrected the signature verification, and made the recovery sweep raise an alert when it catches something rather than fixing it quietly. A safety net that works in silence hides the hole it's covering.
lessonBuild the fallback, then make the fallback noisy. Every recovery path should report how often it fires, or it will end up carrying your system while you think the main path works.
2026-07-18

04 · One member asked the same question thirty times

symptomNothing, technically. Every answer was correct and the member kept coming back to ask again.
causeSessions had no memory of each other, so the system couldn't tell the difference between a new question and the same question for the thirtieth time. A human coach would have noticed on the third and changed approach.
fixAdded recurring-topic detection across a member's whole history. When a subject keeps returning, the system names it and coaches the pattern instead of answering the question again.
lessonA system with no memory can be correct every single time and still be failing the person using it. Repetition is a signal, so log it and act on it.
2026-08-02
2026-08-04

05 · Group calls routed to nobody at all

symptomSome group coaching calls produced no summary for anyone, with no error logged.
causeRouting matched a member's portal name against the name on the call. Real people use shortened names, married names and nicknames, and some attendees had no email on record at all, so the match returned nothing and the call was skipped as unroutable.
fixAlias scoring against known name variants, attendance confirmed against the meeting log rather than the invite, and group-shaped calls detected explicitly so the single-member path stops trying to claim them.
lessonIdentity is the hardest part of any system that spans channels. Pick one identifier a human can't accidentally change, verify against it, and treat every name as a guess.
2026-06-11

06 · It invented software and percentages

symptomOur own evaluation caught the system naming specific tools and quoting specific figures that appear nowhere in Dean's material, in Dean's voice, with total confidence. Grounding scored 3.08 out of 5.
causeWhen retrieval came back weak the model carried on regardless. Nothing in the pipeline treated low-confidence retrieval as a reason to change behaviour on a factual question.
fixA grounding gate that checks factual claims against what was actually retrieved and either corrects the number or declines to give one. Tested against the real failing cases, including the specific fee question that started it.
statusGrounding moved from 3.08 to 3.58 when we re-ran the same case set on 6 August 2026, and the gate corrected a figure on three of the cases in that run. It's still the lowest of the five dimensions. This one isn't closed.
lessonGrade your own system on a fixed set of cases before you launch and after every change, and grade it on grounding, not only on tone. Tone is what you'll notice. Grounding is what will hurt someone.
07Rules it wrote for itself

The system watched its owner work and wrote down what it learned

Every week the system reviews recent coaching call transcripts and its own answers, and proposes behavioural rules for itself. Dean's judgement is the source and the rules are additions to how it responds. Seven are currently active. Here they are in summary, in our words rather than the system's, because the originals quote his teaching material directly.

01

When someone asks how fees or payment terms work, give the actual structure with a worked example. Deflecting to "it depends" is what causes the follow-up question.

02

After stating a profit figure, stop. Let the number sit there. Adding explanation immediately afterwards weakens it.

03

When someone is confused about numbers, offer a short recorded walkthrough they can pause rather than explaining the same thing again in text.

04

Give the next physical action, not the principle. "Walk into three agencies on Thursday" beats "build relationships with agents".

05

When someone brings a problem that's really about confidence, name that before solving the mechanics. Solving the wrong layer sends them back in a week.

06

Match the length of the answer to the size of the decision. A two-line question doesn't get a framework.

07

Never invent a tool name, a supplier or a percentage. If it isn't in the material, say it isn't and say what would settle it.

Rule seven is the interesting one, because the system proposed it after being caught doing exactly that. Rule two came from watching Dean handle silence on a sales call, which is not something anybody wrote into a brief.

This is the part of the build we'd point at if you asked what separates a system that's running from a system that was delivered. It gets better on a schedule, without us touching it.

08What it costs to run

Fourteen days to build, tens of dollars a month to keep

The arithmetic is the part most people get wrong, in both directions. The build is real work and it isn't cheap. What surprises people is that once it exists, the running cost barely moves as usage grows, because the expensive part is thinking and the cheap part is volume.

Generationthe expensive model, on every answer$20 to $35
Hosting and computealways-on container, five scheduled jobsabout $5
Embeddings and transcriptiondaily ingestion of new contentunder $5
Total monthly$30 to $50

Rebuilding the entire knowledge base from scratch, all 7,787 chunks, costs under ten cents. We mention that because it's the number that tends to change how people think about this. The knowledge base isn't a precious asset to be protected. It's cheap enough to throw away and rebuild whenever the source material changes, which is what the daily job does.

Set against that: 205 conversations in a week, including fifteen people prepped in the half hour before a landlord call. The comparison isn't between this system and a cheaper system. It's between this system and Dean's evenings.

Fourteen days from kickoff to production. Nothing here needed a research team.

What it did need was somebody willing to look at the source connections and check their size, grade the output honestly enough to find the grounding problem, and keep running the thing after the invoice cleared. That's the whole job.

Next

If you're the bottleneck in your own business, that's a solvable problem

We start by working out what's actually constraining you, which is usually not the thing that feels most annoying. If it turns out an AI build isn't the answer, we'll tell you that on the call.

Book a diagnostic call