Booked Solid Inspector · Web chat testing · 13 September 2026
We put over 500 scripted conversations through the web chat on all four early-adopter tiles. Nine faults turned up that a customer would have seen. The biggest one was that the web chat could not finish a booking at all. It can now.
Test a lot of times. Test till you found errors, then fix. Test and fix, test and fix, in the in-house. Dil, 13 September
The answer on one page
The web chat is the one to fix first. There is no phone line in the middle of it, so nothing can be misheard. The customer types, and the spelling is already right. Anything that goes wrong there is ours.
The biggest finding, and it came last
Late in the day we stopped asking Josh questions and tried to actually book an inspection, the way a customer would. It did not work — on any tile.
| What we tried | Turns | Booked |
|---|---|---|
| A caller who answered every question patiently | 22 | No |
| A caller who gave every detail in one message | 5 | No |
The saving step never ran once, in either. And the caller was told this:
Nothing was saved. Reading their own address and details back is exactly what makes it believable. That customer goes away certain they have an inspection booked.
Every reply Josh writes has a length limit. The closing message of a booking reads back the name, address, size, year, foundation, occupancy, utilities, who is attending, the package, the price, the day, the time and the inspector — and only then does it save. The read-back alone used up the whole limit. The reply was cut off, and the save that was coming next never happened. One was cut off at the exact words "Perfect. Let me get you booked."
A second cause sat underneath it. Josh only recorded the caller's email if he happened to ask for it in one of eight approved wordings. "What is the best email for the confirmation?" was not one of them. So the caller answered, nothing was recorded, and Josh asked again — over and over instead of booking.
Same conversations, after the fixes:
| Tile | Result |
|---|---|
| GC — normal pace | Booked and saved |
| GC — everything in one message | Booked, and correctly recognised as the same person |
| GC — an agent booking for a client | Booked and saved |
| Staffordshire | Booked and saved |
| Property Masters | Booked and saved |
| Quality | Still not finishing — see below |
Four real records now exist to look for: Dana Whitfield and Sarah Brennan on GC, Priya Nair on Staffordshire, Glen Okafor on Property Masters.
The uncomfortable part
473 conversations had already run before we found this, and they reported five faults — all of them about how Josh words things. None of them noticed that the one job the product exists to do was not happening.
The reason is simple and worth stating plainly. Every check we had read what Josh said. Not one read what the system did. We had a check for claiming a booking that did not exist — but silence is not a claim. A booking that quietly never saved passed as a clean conversation.
The booking test was also only seven turns long, which is not enough to reach the end of a real booking even when everything works. So it never got close.
A booking test now fails unless a real record comes back, and it prints the record's reference so nobody has to take it on trust. The booking conversations are eighteen turns on all four tiles. A reply cut off mid-sentence is now reported instead of passing silently.
Did it improve?
Every row was measured the same way on the same live tiles. The left column is what we found. The right column is the same test run again after the fix went out.
| What we measured | Before | After |
|---|---|---|
| Price answers that used the real price list | 1 in 5 | 5 in 5 |
| Price answers that used it without a “hi” first | 4 in 5 | 5 in 5 |
| Wrong price given for a 2,200 sq ft home | 4 in 5 | 0 in 5 |
| Hardest six price questions answered from the list | 0 of 6 | 6 of 6 |
| Replies showing formatting symbols on screen | 19 in 68 | 0 in 12 |
| Replies where Josh talked about himself to the customer | 9 in 15 | 0 in 15 |
| Saying a job was arranged when nothing was saved | seen live | 0 in 15 |
| Accusing a customer of attacking the system | seen live | 0 in 15 |
| Booking conversations that finished and saved | 0 of 6 | 5 of 6 |
| Replies cut off mid-sentence in a booking | 4 of 4 | 0 |
Those two faults were rare — each turned up once in 171 chats. Fifteen clean runs afterwards means the fault did not come back. It does not prove the new guard caught it, because it never had to. Both are proven in our own tests against the exact sentence. We will keep watching for them on the real tiles.
What a customer would have read
These are real replies from the live tiles, copied exactly.
Josh was writing in a style meant for a web page, but the chat window shows plain text on purpose. So the stars showed up. It is untidy rather than costly, and it was the most common thing we found. Fixed: the stars are now stripped before the reply leaves.
Our system sometimes sends Josh a private note mid-reply — for example, “you did not answer what they asked.” The note says plainly that the customer cannot see it. Josh was answering it out loud anyway. Fixed: those sentences are now removed before the reply is sent, and the real answer is kept.
The home was 2,200 sq ft, where GC's price is $450. The $425 is real — it is GC's flat pre-drywall fee — so our existing checks let it through as a price Josh is allowed to say. Said about a normal inspection, it undercharges. Fixed: a fourth check now looks at what the sentence is about, not just the number. “Pre drywall is a flat $425” still goes out untouched.
Nothing was arranged and nothing was saved. That customer drives away believing a callback is booked. We already caught “let me get you booked” — a promise can still be kept in the same breath. “You're all set” cannot; it is already untrue when it is said. Fixed: the phone has had this guard since August, and the web now has it too.
The customer asked an ordinary question. Josh answered it correctly and then accused them of attacking us, and told them how he is built. This is the same root as fault 2 — he read our private note as something the customer sent. Fixed: the accusation is removed. The price and the closing question still go out.
How the testing works now
We used to print the conversations and read them. That works for five chats and fails at five hundred. Now every reply is checked automatically against the rules we already hold — never quote a price no tool produced, never text a homeowner, never claim a booking that did not save, never deny being an AI, never invent a discount. Only the ones that look wrong get printed.
Seven times the testing flagged a reply that was actually correct. Josh added $450 and $85 and $95 to get $630, which is right, and our check did not understand adding up. He repeated a customer's own figure back to them to turn it down, and our check read it as a price he had invented. Each one cost time to chase. All are fixed now — but it is the reason we re-check a finding before calling it a fault.
The 57 situations go well past someone asking a price. An agent booking for a client. A seller who wants a pre-listing. Someone who says the size, then changes it, then changes it again. A caller who asks for a Saturday when the firm works weekdays. A commercial building. A house too big to quote. Someone asking for a discount. Someone quoting a rival's price. A rude caller. A caller who says up front that this is only a test. Someone asking to be texted. Someone typing in capitals, or typing just a question mark, or writing seventy words in one go.
One decision, for Ken and Dil
When anyone publishes an update, the web app restarts. Anyone mid-conversation at that moment sees “Connection problem — please try again.” We measured seven of these during testing. It happens several times on a normal working day, because more than one person publishes updates.
We already knew this about the phone and wrote it down. Nobody had written it down for the web chat, and it happens far more often there.
The obvious answer is to let the old version finish its conversations before the new one takes over. That is a real change to how the app runs, not a setting. The other answer — have the chat window quietly try again — risks booking the same job twice, which is worse than showing an error. Both need a decision rather than a quick edit.
What this does not cover