Daily report · Sunday 13 September 2026

The web chat could not finish a booking. Three faults were in the way. All three are fixed, and the phone got better too.

Three Joshes were worked on today. Josh in the web chat got 473 test conversations, and then a harder question: does he actually book? He did not. That is fixed. Josh on the phone got 26 test calls on Chad’s line, six deploys, and a long list of repairs. Josh on the test line had more than twenty faults found in-house, without spending a single test credit. Nothing was sent to a client, and no customer heard any of it.

26sets of changes merged today
26test calls on Chad’s line, in four runs
473web chats graded by computer
6deploys to the live phone
0clients touched

Read this first We do not text homeowners. The only text this system sends is the booking alert to the inspector's own mobile. A home buyer, seller or agent hears from us by email only. That did not change today and it is not going to.

Nothing reached a client. Every test call went to Chad's line, which is a test line now. Dil ruled on 9 September that all four early-adopter lines carry no real inspections. Test bookings landing in the real system are expected.

The phone is current Dil pressed the deploy button six times today, the last two at 10:17 a.m. and 10:36 a.m. Texas time. Everything in this report that touches the phone is on Chad's and Ted's lines now, including the three fixes that came out of reading run 8's log. Nothing is waiting on a press.

Three Joshes, one day


It helps to know which Josh each part of this report is about.

 Josh in the web chatJosh on the phoneJosh on the test line
Who uses itAnyone who types on one of the four early-adopter pagesAnyone who rings Chad's or Ted's lineNobody. A test number and a made-up firm
How he worksHe decides what to ask next. The software checks every replyHe decides what to ask next. The software says the numbersHe follows a checklist, like Ken asked for
How we tested473 scripted chats, graded by computerFour runs, 26 calls, and the server's own logHundreds of whole calls driven through the code
TodayFive faults fixed, then the booking itselfKen's email rule proved. Numbers said rightMore than twenty faults found and fixed

Part one — the web chat


Dil's instruction for the day: “test in house. test a lot of times. test till you found errors/issues then fix.” So that is what happened.

We used to print the conversations and read them. That works for five chats and fails at five hundred. Now every reply is checked by computer against the rules we already hold. Never quote a price no tool produced. Never claim a booking that did not save. Never text a homeowner. Only the replies that look wrong get printed. Here are the same tests, before and after the fixes.

What we measuredBeforeAfter
Price answers that used the real price list1 in 55 in 5
Wrong price given for a 2,200 sq ft home4 in 50 in 5
Replies showing formatting symbols on screen19 in 680 in 12
Replies where Josh talked about himself to the customer9 in 150 in 15
Saying a job was arranged when nothing was savedseen live0 in 15
Accusing a customer of attacking the systemseen live0 in 15

The full story of those five faults, in Josh's own words, is in its own report: What 473 test chats found, and what we fixed. It is written for Ken and takes five minutes.

Then Dil asked the question that mattered “Every successful test from these four must be saved or recorded in the client's portal. We need the web chat Josh to behave the same as the phone call Josh.”

So we ran a booking on Chad's page with a caller who answered every question. Twenty-two turns. Not booked. Then a caller who handed over every detail in one message. Not booked. The save was never called once. And the caller was told:

“You're all set, Dana. Here's what I have: Dana Whitfield at 1827 Willow Bend Drive, Pearland, 77584 — 2,200 square feet, built 1998, slab foundation, vacant with utilities on, and you'll be attending.”

Nothing was saved. That is the worst thing this system can do, and the full read-back is what makes it believable. Three faults were in the way. All three are fixed below.

What we fixed in the web chat


Each one came from a real test conversation we can point at. All are live on the web now.

1

The email answer did not count unless Josh asked in an exact way

Ken's rule is no email, no booking. So the email is required. But the software only counted the caller's answer if Josh's question matched one of eight set phrasings. “What is the best email for the confirmation?” did not match. Neither did “Best email?” So the caller answered, nothing recorded it, and Josh asked for the email again and again instead of booking.

Fixed. An email address is now recognised by its shape, wherever it turns up. Ten cases were checked both ways, including five that must not count, like an agent's address the caller mentions in passing. This is the same family as the street address fault fixed on 10 September. A few other required fields have the same gap. Each needs its own check before it gets the same fix.

Live on the webPR 779
2

The reply limit was smaller than a booking read-back

At the end of a booking Josh reads back the name, address, size, year, foundation, who is attending, the package, the price, the day, the time and the inspector. Then he has to save. His replies had a size limit of 400 tokens, and the read-back alone filled it. He was cut off before he could save. We measured it on all four pages: four booking conversations, four cut-off replies, and the save never called. One was cut off at the exact words “Perfect. Let me get you booked.”

The limit is now 900. And a reply that does get cut off carries on once, with one instruction: if you were about to save, save now, before any more words. The limit is a safety net. It is not what keeps Josh brief; his instructions do that.

Live on the webPR 783
3

“You're all set” is now removed before it reaches the screen

Earlier today we added a check that asks Josh to put it right when he claims a booking that did not save. He said it again on the retry, and there was no check left. So asking is not enough. Now the claiming sentence is removed before the reply goes out. The read-back stays, because the read-back is how a caller catches a wrong detail.

It cannot damage a real booking. It only runs until the CRM has confirmed the save. After that, every confirmation Josh makes goes out untouched. This is the phone's rule, brought over to the web.

Live on the webPR 783
4

“Josh had a problem” — one retry, and a reason we can read

Running the same booking three times, it died once at an ordinary turn (“Yes, that works.”). Nothing was wrong with the conversation. The shape of it was an overloaded call, not a bad request. And from outside, nobody could tell: the error message gave no code.

Now a passing fault like that is retried once, with a short wait. The retry only happens if nothing has been saved in that turn. Booking twice is worse than failing once, so a turn that already saved is never replayed. The error now carries a status code, so the next one can be diagnosed from outside.

Live on the webPR 783
5

The tester graded the manners, not the work

This is the honest part. All 473 chats were graded on what Josh said. Not one check asked what the system did. A booking conversation that quietly never saved came back clean, run after run. The booking test was also only seven turns long, not enough to reach a save even on a healthy day.

Fixed. A new check fires when a booking case ends with no contact saved, and says whether the save was even attempted. Another reads whether a reply was cut off. The booking cases are now 18 turns of a patient caller, on all four pages, plus one that hands over everything at once.

Live in the testerPR 779
6

Publishing an update hangs up on whoever is chatting

Seven errors in one test run, all in a row. It was our own update landing mid-run. Every update restarts the web app, and anyone typing at that moment sees “Connection problem — please try again.” It happens several times on a normal working day. We had written this down for the phone and never for the web, where it happens far more often.

Not fixed, on purpose. The obvious fix is a gentle restart, but that only works in a mode the app does not run in. The other fix, a quiet retry in the chat window, risks booking the same job twice. Both need a decision from Dil, not a quick edit. The tester now retries on its own, so an update is no longer reported as seven faults.

Dil's callPR 775

What is proven, and what is next Each fix above was proved against the exact conversation that failed. The next check is the one that matters: a full booking, saved, on all four pages, with the new tester watching. That is the first thing to run.

Part two — the phone


Four test runs on Chad's line, 26 calls in all, and the server's own log read after each one. Dil pressed the deploy button six times. The rule that held all day: read the log, not the transcript.

Every fault fixed today on the phone was invisible in the test transcript. The transcript is what the test company heard Josh say. The log is what our code actually did. Half of today's phone faults turned out to be the instrument, not the behaviour. Both got fixed.

1

Ken's rule is on the live phone: no email, no booking

Ken, 7 September: “we can't take the inspection without your email — have a blessed day. Boom, hang up… that's not even a qualified lead.” Dil confirmed it on 10 September. The web and the test line already had it. The phone did not: it booked a caller who refused her email on five runs in a row.

Now it is live. A caller is asked at most twice. Then Josh says Ken's sentence, saves nothing, and the line drops. A caller who was asked and simply never answered still books, with no email promised. That is Ken's own carve-out: refusal, not absence.

Proved on runs 7 and 8. Nobody who refused an email was booked. One caller who was booked on a refusal on run 7 was correctly turned away on run 8.

Live on the phonePR 764
2

Josh reworded Ken's sentence, so now the software says it

Run 8 showed the next problem. Josh saw the refusal, and then ignored the instruction to say the sentence word for word. On three calls out of three he improvised instead:

“Here's what I can do… your inspector will reach out to confirm your email.”
“No problem at all — we'll confirm your email… let me get you booked.”

Nothing was booked, because the save gate held every time. But the caller heard three different things, none of them Ken's. So the sentence is no longer trusted to Josh. Once the refusal is counted, the software speaks Ken's sentence itself, once, and drops anything else Josh writes for the rest of the call. Eight lines that said “the team will…” now say “your inspector will…”, which is what Ken asked for on 6 September.

Live on the phonePR 777
3

Chad's “he never gets my email” — the cause was found in-house

Two hundred generated calls with a caller who answers whatever Josh asks. Two faults fell out. First, seven yes-or-no questions (is there a pool, any outbuildings, any add-ons, and so on) treated a bare “No.” as a filler word, not an answer. So the caller answered, nothing recorded it, and Josh asked again. Second, the add-ons question could be asked again and again, and it blocked every question behind it. The email was one of them.

Both fixed. Failures on those 200 calls went from 650 to 60, and the 60 that remain are one deliberately unfair case. Live on the phone since the morning's deploys.

Live on the phonePR 774
4

Four number faults found by fuzzing, and one was a wrong price

The software that says numbers out loud was pushed through 120 ZIP codes in six written forms, every number sentence Josh has ever produced, and every reply chopped up 28 different ways. Four faults:

Josh wroteThe caller heardNow
Pearland, seventy seven zero four.a four-digit ZIP7-0-7-0-4, five digits
Pearland, 77,584.“seventy-seven thousand five hundred eighty-four”7-7-5-8-4
8:30 works.“eight thirtyworks”, one word“eight thirty works”
$4.25, arriving in two pieces“four dollars .25” — a wrong price“four dollars and twenty-five cents”

The last one is the serious one. Nothing was lost; the figure was simply wrong, and only when the network split the reply in the wrong place. So it was right nine times in ten and every existing test passed. None of the four would have been found by a phone call. All four are live since the morning's deploys.

Live on the phonePR 772
5

The first test run after the deploy found the ZIP still wrong

Six test calls were written to prove the morning's fixes, on the client lines, because Ken's standing rule sends behaviour tests there. On the closing read-back Josh said, in his own audio:

“Mike Chen, twelve Oak Street, Pearland, seventy seven thousand five hundred and eighty four.”

The morning fix caught the ZIP written in digits after a comma. Nobody checked the ZIP written in words after a city. They are two separate doors in the same file, and both have to be opened every time. Fixed the same day, with 26 checks.

And most of the “one pass out of six” was our test file, not Josh. It held five Texas cases and one Delaware case, and the test company dials one number per file. On Chad's line the Delaware caller was correctly refused as out of area. Josh was right every time. The file is now split into one per firm, and a new check refuses any test file that names two.

Ted's suite is now built from Ted's own facts, not Chad's with the town swapped: his towns, his Saturday-morning rule, his $650 package, and Kennett Square in Pennsylvania. His area crosses two state lines, and a Josh who has learned “we serve Delaware” would turn that caller away. Nothing tested that before today.

Live on the phonePR 785, 786
6

Two houses on one call keep both prices, and a blocked price calls the tool

On run 7 a caller asked about a 3,400 sq ft house and then a 2,200 sq ft one. Josh quoted $500 and then $450. Our guard read the second quote as a correction, threw the $500 away, and then blocked Josh saying it eight times over two minutes. That was the dead air and the “still with you” on that call. Quotes are now kept per size. Run 8 proved it: “$450 for the 2,200 square foot home, and $500 for the 3,400.”

Also from run 7: when the guard blocks a price Josh read off the table by eye, it used to replace it in the audio and tell Josh nothing. He carried on as if the caller had heard it. Now he is told the price was not spoken and to call the price tool. Live since the 8:02 a.m. deploy.

Live on the phonePR 773
7

The closing promise is built from what we actually hold

The phone used to promise an email and an agreement at the end of every booking, whether or not we had an address or the agreement was switched on. Now the server builds the closing line from what is true for that booking, and Josh says it word for word. The web already worked this way; both brains now use the same sentence.

Proved on run 8. One caller heard “you'll get an email with a secure payment link” and nothing about an agreement, because the agreement switch is off on the server.

Live on the phonePR 773
8

The spelled email — the guard fired on a live call for the first time

This has been the stubborn one. A guard that catches a spelled-out email shipped on 11 September, and Dil's searches of the server log kept showing it had never once fired on a live call. Three reasons were found, one per run, and each was fixed: the caller's turn arrives in pieces and the guard was shown one piece at a time; the phone's transcriber puts a full stop after every piece; and it mishears the letter but not the code word, so “november” now decides the letter. And the address the caller spelled is now the one we save, whatever Josh typed.

Then Dil pasted run 8's log, and the guard had fired. On the call where the caller spelled D-A-N-A-K-I-R-K-B-R-I-D-E, the log shows the guard caught the whole address and replaced three wrong read-backs (danikirkbright, danikirkbride) with the letters she gave. The same log showed three more faults, all fixed the same hour: a correct read-back behind a preamble was being “corrected”; two second refusals (“No, thanks, I don't give my email”, “I don't wanna give my email”) were shapes the rule did not read, so Ken's sentence came at the third refusal instead of the second; and “saving fifty-eight dollars” was blocked as an invented figure when it was the difference between two prices the tool had quoted.

What is left is the transcriber, not our code. On one call eleven code words (“N for November…”) reached us as NIABEHLY. On another, letters said one at a time arrived as ordinary words, so there was no spelling to catch. Only logging the caller's exact words for one run will show what the transcriber writes. That is Dil's call.

Live on the phoneDil — one decisionPR 767, 773, 784, 787
9

Two recap guards: the wrong weekday and the changed house number

Run 8 had a recap that said “Wednesday, October 1” for a Thursday. Run 7 had a recap that said 1727 Willow Bend for a house number the caller had confirmed as 1827. Both are now caught before they are spoken. A weekday that contradicts its date is corrected. A confirmed house number with one digit changed is never spoken; Josh asks for it again instead.

Live on the phonePR 784
10

Five ways of saying “I need to ask my husband” were missed

Ken's own method says the spouse deferral is about 90% of lost bookings, and that it is a trust problem, not a price problem. Both brains had a list of phrases that count as that deferral, and every one needed the word with. So “I have to talk to my husband”, “I'll run it past my wife” and “I need to ask my wife” all fell straight through. Four phrasings became eight, with nine sentences pinned that must not match, like “can I ask you a question?”

Live on the phonePR 781
11

The instruments, fixed so the next fault can be seen

Four log lines that exist to say which field was missing had been printing the word “%s” instead of the field name since they were written. Fixed, with a check that fails the build on any repeat. The tool that reads a test run now catches today's four faults by name, and says on screen which of its own answers to trust: a transcript is Josh's audio, transcribed, so a “none” from it is not proof on its own. And Dil's deploy presses were checked rather than reassured: every one is recorded, and the button refuses to ship old code.

LivePR 767, 780

Part three — the test line


A second Josh runs on a test number for a made-up company. This is where the checklist Ken asked for is being built. Today it was tested without a single phone call.

Instead of ringing it, we drove the code itself. Every caller sentence we have ever recorded, 1,645 of them from 72 test calls, was run through every function that reads what a caller said. Then hundreds of whole calls were driven through the checklist with a perfect Josh assumed, so anything that went wrong was the code's. Then Josh's own 6,227 recorded sentences were run back through the guards that police his mouth. No test credits were spent on any of it.

1

Six faults sitting in the recorded calls all along

A curly apostrophe is a different character from a straight one, and real transcripts are full of them. So “I don't want to give my email” did not read as a refusal, and “No, it's Peachtree” came back as the surname Speachtree. A full stop was read as a dot in an email. “I won't be there, but the site manager will let the inspector in” read as a refusal. “Yes, that's fine.” was recorded as the name Fine. And an email said as two words, “anna king at gmail dot com”, lost its first half, five separate times. All six fixed.

2

One of them booked the wrong hour

When Josh offers two openings and the caller picks one, the code only understood the literal word “first”. “9am”, “morning”, “the 9 AM one” and an empty answer all booked the 1 p.m. when she had picked the 9 a.m., and nothing on the call said so. It now reads her words, and falls back to the opening Josh offered first. Two more from the same method: a question whose answer never made sense was asked forever, up to 59 times, and an empty diary sent 41 calls out of 200 round in a loop until the ten-minute cap. Both now end the call honestly, with the inspector told.

3

Eleven booking claims and seven promises walked past the guards

Every one of these is a sentence Josh has really said on a test call. “I've got you down for 88 Cypress Lane” is how a person says “booked”, and the guard had never heard of it. “The office will follow up with the details” is the exact promise Ken banned on 6 September, because most inspectors have no office. Both are caught now. Nothing that was blocked before is unblocked.

4

The fuzzer is now part of every test run

Four hundred whole calls, with chatter, corrections, mid-call changes, a refused email and a diary that sometimes has nothing, in about four seconds. Every call must end properly, save at most once, never say “you're all set” without a save, and never promise an email we do not hold. It also proves Dil's checklist ask on every call: 36 questions on the list, none missed. Delete one from the walk and 187 of the 400 fail.

5

Ken's spouse rebuttal is on the test line

Dil: “lets do this now.” Until today the test line answered every hesitation at the price with “No problem at all. Thanks for calling, and have a great day”, all but word for word the sentence Ken names as the failure. Now the code spots the deferral, picks her own word for her partner, and speaks Ken's sentence from the shared script. One question, then stop. She can still book from there, which is the whole point. Nothing here is written by us.

Live on the test linePR 781

What is still wrong


Named, so nobody thinks it is done.

The spelled email — now the transcriber's turn The guard has been seen working on a live call. But on two of the four spelling calls the phone's transcriber handed us text no code can read: eleven code words squashed to NIABEHLY, and letters said one at a time turned into ordinary words. The next step is to log the caller's exact words for one run and read what the transcriber actually writes. Dil decides whether to switch that on.

A full web booking, end to end Three faults in its way are fixed. A booking that saves on all four pages has not yet been watched happening with the new tester. That is the next run.

Not built yet The last name is the only field with no read-back, so Benschley and Rays for Reyes go uncaught. And Josh sometimes asks two things in one breath, which the caller then answers as one word. Both are known, neither is built.

What needs Ken


Five decisions. None of them is code.

Your call

  1. Words for the “I can't catch that” line. When Josh cannot make out a detail after three tries, he now says: “I'm not catching that clearly enough to book it safely. I'll pass what I have to your inspector and they'll reach out to you.” It promises no office. Change it if you want different words.
  2. Words for the failed-save line. The web still says “someone from our team will follow up” when a save fails. It is waiting on your own words, as agreed on 6 September.
  3. A caller who answers nothing. Today he is asked for the city and ZIP about twenty times over a long call, because the rule parks the question and brings it back. There is no cap. Whether there should be one is a judgement about what Josh should do with a caller who will not answer.
  4. The ISN report. ISN forms vs Josh ends with two calls that are yours: whether the deeper questions are standard for every firm but Chad, and whether to build the six access-code boxes.
  5. The test company's own pass marks. It reported a call that booked on a refused email as “all passed”. Replacing its four generic checks with ours needs a yes from you or Dil, because it changes the shared workspace.

What needs Dil


In the order to take them.

Next steps

  1. Run the two new suites — five cases on Chad's line and six on Ted's, built from Ted's own facts — and the six earlier cases (A6-1 to A6-4, S17-3, S17-4). Everything they test is on the phone now. Send the export and the same search of the log: pm2 logs austin-voice --lines 20000 --nostream | grep -E "\[email\]|\[book-gate\]|\[promise\]|\[quote\]|\[price-guard\]|\[spelling\]"
  2. Re-import the two rewritten test rubrics for A6-3 and S17-4. They used to require a booking, which is the old behaviour. They live in the test company's workspace, so only you can update them.
  3. Decide on logging the caller's exact words for one run, now that the transcriber's text is the blocker on spelled emails. All four lines are test lines, so the rule against logging a real caller does not bite. It is still your call, and it goes off again afterwards.
  4. Decide what to do about updates hanging up on live web chats (part one, item 6).
  5. Run the new voice comparison when there is time. Six cases are built for each side, on the test number and on Chad's line, and the run sheet fixes what counts as a win before the numbers arrive. They grade what a voice actually controls: digits that reach the ear, hard words, talking over Josh, dead air. The old way of scoring mostly graded his brain, which is the same on both sides.
  6. After the runs, one search clears the test bookings from the CRM: every caller in both suites carries the surname Fixcheck.

Since the last report


The last daily report was 8 September. The days in between, in one line each.

DayWhat changed
10 SepJordan was retired. Code work now goes straight to Claude Code sessions, not through an agent.
10 Sep“Add an inspection” is back in the Inspector Portal, under Inspections.
10 SepKen's thirteen faults from the first website test turned out to be five causes. All five were fixed in code.
10 SepThe firm's own ticks on the setup form now decide which extra questions Josh asks.
10 SepDil ruled that “no email, no booking” applies to the web chat too. Built the same day.
10 SepParker, James and the meeting analyser moved to Sonnet 5.
11 SepElevenLabs is on the Scale plan. The phone can voice 15 calls at once, up from 3.
11 SepThe mobile being asked twice was traced to its root cause and fixed.
11 SepThe new voice (v3) stays on the test number. It does not go to Chad and Ted on the evidence so far.
11 SepReview replies moved to Sonnet 5. No live job runs Opus 5 any more.
12 SepThe direct price question is answered 12 times out of 12, measured after the fix went live.
12 SepJosh no longer answers our private system notes out loud to the caller: 0 of 6 on a targeted test.

How we know all this


Every claim above has something behind it. Here is what.

The evidence

  1. 26 sets of changes merged, numbers 763 to 788, each with its own automatic checks.
  2. Four test runs on Chad's line, 26 calls: run 6 (4 calls), run 7 (10), run 8 (6) and the six-case suite after the second deploy. The exports for runs 6, 7 and 8 are filed in the repo, and each run has a written read-out.
  3. Six deploy presses, all recorded, from 4:28 a.m. to 10:36 a.m. Texas time. One was checked file by file against the code it was meant to ship, and the button refuses to ship old code.
  4. The full test suite: 297 steps, 289 green. The eight red are the known list, unchanged by today's work, and every one has a name.
  5. 473 web chats graded by computer, then the booking runs on all four pages.
  6. Hundreds of whole calls driven through the test line's code, with no phone call and no test credit spent.
  7. Every quote in this report is copied word for word from a transcript, a server log line or the written record of the change.