Retest — Question Set Builder v1.1.0
INTERNAL TEST ARTIFACT. Never ships. Run against a real client session; contains client-identifiable content and stays in the repo.
Input: Stevan and Kathryn, Blueprint Mapping Session, 7/27/2026. Mode A, full timestamped transcript, 0:01:56–1:43:06. 456 speaker turns, 15,231 words. Purpose: verify the merged rules produce reproducible counts, and that no diagnosis or characterization survives.
The counts, against both prior versions
| v1.0.0 (Cole) | Kathryn's version | v1.1.0 (merged) | |
|---|---|---|---|
| Questions in the set | 27 shown in body | 6 | 6 in body, 27 in collapsed appendix |
| Dead answers | 9 | 5 | 4 counted + 1 borderline |
| Answered for them | 7 | 4 | 5 |
| First 15 talk ratio | 44/56 | 44/56 | 44/56 |
| Whole call | not reported | 58/42 | 58/42 |
| Benchmark / verdict | "Target 30/70 · Gap 14 points over" | none | none |
| Diagnosis lines | present | none | none |
The ratio is the cross-check. Three independent passes, same numbers. Computed here from speaker labels: 1,016 words her / 1,283 his across 0:01:56–0:16:56.
Why the dead-answer count moved to 4
This is the reproducibility test, and it worked. Three of the four fall out mechanically — the prospect's own words trigger the counts list with no judgment involved:
| Time | What came back | Rule triggered |
|---|---|---|
| 0:10:59 | "So repeat the question one more time, just so understand the perspective." | request to repeat or rephrase |
| 1:18:28 | "Yeah, I'm I'm drawing a blank on this one." | "I'm drawing a blank" |
| 1:22:16 | "Sorry, repeat that one more time." | request to repeat or rephrase |
The fourth needs the rule, not taste. At 0:05:43 she asked about a conversation that worked; the reply at 0:06:00 went to preparing video ads, and she restated the frame at 0:07:22 before an answer came. Two criteria fire: an answer to a different question than the one asked, and an answer the user has to restate the frame to get.
And one came out. 1:16:11 — "What would you say next? What would Steven say here?" The reply begins to answer and then turns back: "I mean, typically, what would I say would come after this?" That is on the doesn't-count list — a question the prospect turns around and asks back for a legitimate clarification. Excluded from the count, listed as borderline, exactly as the rule directs. Kathryn's version counted it; v1.1.0 doesn't, and the difference is traceable to a written line rather than to a judgment call.
What v1.0.0 was counting that v1.1.0 isn't: short-but-complete answers, caught by the old "under about ten words with no new information" proxy. "No, not typically" answers the question. Brevity is not death.
What the no-diagnosis rules removed
Every "Why it died" line is gone, replaced by Move broken, which names M1–M4 and takes the question as its subject.
Cut from the v1.0.0 output, verbatim:
- "They picked the one they had an answer ready for." — a claim about him.
- "Nobody answers this on a call." — a claim about people in general, unsupported.
Kept, because the subject is the question:
- M2 — the question carried its own answer.
- M1 — asks for a figure not in front of them.
At 0:10:59 and 1:22:16 the reply is a request to repeat. Per the rule these are noted as landing problems, not design problems, and the replacement still stands. No cause is assigned — the line, the connection, and the listener are all invisible to a transcript.
Talk ratio, as it now reports
| Her share, first 15 | 44% |
| His share, first 15 | 56% |
| Whole call | 58% / 42% |
| Method | Measured from speaker labels, ±2 pts. 15,231 words counted. |
| Window | 0:01:56–0:16:56. Recording starts at 0:01:56, so the window is wall-clock — offset noted, not adjusted. |
No target. No gap. No verdict. From roughly 0:38 the session turns from interviewing into building the deck, which moves the whole-call figure for reasons unrelated to the questions.
On this call the number needed no verdict anyway — she was running a paid working session she was hired to talk in. v1.0.0 printed "Gap: 14 points over" against an invented standard and told her she was failing.
Open, and honest
- The projection path was not retested. The denominator changed from share-of-call to share-of-answer-time, and that needs its own recall-path run before it ships. Not verified.
- The appendix rendering was not produced here — this run tested counts and rules, not layout. The 800-word body cap needs one real output to confirm against.
- Both are Cole's, before any link goes anywhere.