Investment Plans workspace
Open raw ↗

V3 — RE-VERIFICATION AFTER CR-004

Agent: V3, independent verifier (GOV-C3.2, GOV-E2.3). Did not build any artefact, and did not participate in the corrections under CR-004. Date: 2026-08-20 · Baseline: v1.0.0 rebuilt under CR-004 · Prior report: 02_Work/scratch/V3_deliverable_verification.md (unchanged, retained as evidence). Method: the coordinator's list of twelve fixes was treated as a claim to be tested, not as information. Every status below was reached by opening the artefact and, where the fix is arithmetic, recomputing it. Nothing was edited. The checker suite was deliberately not run: check C23 tampers with a pack file and rewrites checker_run_log.txt, which would modify artefacts under verification. The suite's source was read instead, and its log was read as delivered.


HEADLINE

Ten of the twelve anomalies are genuinely fixed. Two are partially fixed. Three new defects were introduced or left behind by the fixes, one of them in the same tile the original defect was in. The corrections are real work, not relabelling: the arithmetic errors are gone, the assurance layer is materially harder to fool than it was, and Appendix B2 publishes all twelve findings unprompted. Requirement verdicts move from 18 PASS / 8 FAIL to 21 PASS / 5 FAIL. The project's own claim of "26 of 26 verified" is still not supportable.


PART 1 — ANOMALY-BY-ANOMALY

#StatusWhat I opened, and what I found
ANOM-01 employed-manager break-evenFIXEDmodel_params.breakeven_hours_employed_manager() exists and returns 411.37. Study §1 ACT-001 box now reads "If you employ a manager on about $95,000 a year instead, it is about 411"; the dashboard ACT-001 row reads the same. The string "past 126" survives only inside the DEF-004 history record, correctly labelled as the error. Recomputed independently: (847.33 + 95,000/12) / 21.3042 = 411.4. ✔ One overstatement in the fix record — see NEW-03.
ANOM-02 support-coordination runwayFIXEDrunway_downside() now takes an explicit workers parameter and the SC branch passes workers=0. Study Table 5.1 reads $8,757. Recomputed: 3,673.20 + 6 × 847.33 = 8,757.20. ✔
ANOM-03 ASM-008 wrong consequencePARTIALLY FIXEDASM-008 now reads: "At 0.40 hours … NEGATIVE $2.07 … At 0.10 hours it is $15.46." Both verified (21.3042 − 58.4292 × 0.40 = −2.07; 21.3042 − 58.4292 × 0.10 = 15.46). The $10.85 is gone. But the second stranded register figure documented in the same anomaly is untouched: ASM-013 still says support coordination "earns roughly $71,000 a year gross instead of $112,000", where the model computes gross $69,257 (at 35%) and $108,832 (at 55%), and net $59,089 / $98,664. Neither published figure is reproducible from any of the four candidate readings.
ANOM-04 three open-item countsPARTIALLY FIXEDproject_data.open_item_count(), open_rows(), open_breakdown() and _STATUS_COL all exist. I verified every index against the CSV headers: RQ 5, ACT 6, CR 13, DEF 8, RSK 9, ISS 6, BKL 7 — all correct. README, the dashboard tile and HANDOVER.txt now all publish 26, which equals the canonical count and the published lineage formula. C11 enforces the README and the tile against the registers. But the fix introduced NEW-01 and left NEW-02 (below), both in the tile it repaired.
ANOM-05 appendices overstateMOSTLY FIXEDAppendices A and B are now generated and each line names a checker ID or a file. Gate 2 and gate 3 cite 02_Work/scratch/V3_deliverable_verification.md, which now exists — the "verified before the verifier ran" defect is closed. Gate 6's claim is now backed by a C09 that really does open the workbook. New Appendix B2 publishes all fourteen defects with severity and correction, including the 3x error, unprompted. Residuals: gate 5 still asserts "LibreOffice reports 398 formulas" where the delivered workbook contains 394 formula cells by direct count; gate 13 asserts C29 "measures 1,051 text runs" where the run log says 1,114 and the file actually contains 2,205. And see NEW-04 — the evidence register still attributes to V3 the evidence for four requirements V3 failed.
ANOM-06 workbook cost bandFIXEDCosts!B19:D20 now carry 200/=Drivers!$B$7/500 and 150/=Drivers!$B$8/350. Cached values B28 = 601.5, C28 = 847.333, D28 = 1438.167 — identical to model_params.monthly_fixed() and to the study's "$602 to $1,438 (base $847)". ✔
ANOM-07 no cached valuesFIXEDRecalculation is a real step in build_all.py. The delivered workbook now carries 501 cached numeric values: UnitEconomics!B11 = 21.30424, BreakEven!B12 = 7,318.61, SupportCoordination!B12 = 98,664.15. A reader with no calculation engine now sees answers. ✔ But this fix introduced NEW-02.
ANOM-08 checks assert more than they testMOSTLY FIXEDC09 now opens the workbook and compares twelve figures across study, dashboard, README and registers, including the employed-manager figure. C11 reconciles the published open-item count against the registers in both the README and the tile. C20 now scans the delivered study. C27, C28 and C29 are new and are real measurements, not assertions. Residuals: (i) C27's reference pattern accepts REQ-, DEC-, RSK-, DEF-, BKL-, ACT- and the literal strings "Financial Model" / "model_params" — so its 1.000 is not the test REQ-MOP-01 states; measured strictly on SRC/ASM the ratio is 0.860 (see NEW-05). (ii) C29 measures only the 1,114 runs that carry an explicit size and silently ignores the 1,091 that inherit from a style — it happens to be safe here, but it would not catch a regression introduced by changing Normal. (iii) C11 checks the tile's headline number and not the breakdown printed beside it, which is why NEW-01 survived.
ANOM-09 Windows path in the studyFIXEDZero occurrences of C: in the rebuilt study. Gate 10 now reads "Paths in this document are RELATIVE by design (GOV-F4.2)" and states that the absolute delivery location belongs in the covering message. C20 scans the study. ✔ (A different absolute-path problem now exists elsewhere — NEW-02.)
ANOM-10 unquantified qualifiersFIXEDREQ-SYS-03, REQ-SYS-08, REQ-CON-04 and REQ-MOP-01 rewritten with measurable tests. C01's banned list widened from 7 literals to 26 covering the qualifier class. Independently re-scanned all 26 statements: every one has exactly one "shall" and none contains a banned qualifier. ✔ Note REQ-MOP-01's rewrite narrowed the requirement as well as sharpening it — see NEW-05.
ANOM-11 weights-before-scores assertedFIXEDDEC-003 and DEC-004 now record "WEIGHTS FIXED AND RECORDED AT 2026-08-20T09:40 AEST" against "Scored 2026-08-20T11:05 AEST". The ordering is now stated with a clock time and is testable against the AC. ✔
ANOM-12 ISS row and roundingMOSTLY FIXEDISS.csv ISS-001 now reads Owner Zaid, Status Open — the column shift is gone and the row is counted correctly. ACT-009 added and it is substantive: "Ask one operating Victorian NDIS provider how many non-billable administration hours they spend for every billable hour", owner Zaid, due 2026-09-05, blocking "the base case … the service model recommendation, and the go/no-go itself (ASM-008)". The break-even rounding is now consistently "about 40" in both the study and the dashboard. Residual: study §1 finding 4 still states the 300-hour wage bill as $15,684 where four other places say $15,683 (true value 15,682.73).

PART 2 — WHAT THE FIXES BROKE OR LEFT BEHIND

NEW-01 — MEDIUM — The repaired open-items tile now contradicts itself inside one sentence

Dashboard.html, first status tile: "26 OPEN ITEMS across every register — the first number to read (ACT 9, RSK 10, ISS 1, BKL 6, DEC parked with Zaid 1)". The breakdown sums to 27. Cause: open_item_count() totals seven registers (RQ, ACT, CR, DEF, RSK, ISS, BKL) per the published lineage formula, while open_breakdown() appends an eighth line, DEC parked with Zaid, that the count excludes. C11 tests the headline number and never adds up the label beside it. This is the same defect class as ANOM-04 — a reader who does the arithmetic gets a different answer from the one published — reintroduced by the fix for ANOM-04, in the same tile, under the same instruction to treat it as the first number to read.

NEW-02 — MEDIUM — The documented rebuild command is now machine-dependent and will fail on any other computer

01_System/build_all.py lines 13-14 and 19-20 hard-code two absolute paths: /root/.claude/skills/synced/xlsx/scripts/recalc.py and /root/.claude/skills/synced/docx/scripts/office/soffice.py. There is no existence check and no fallback: main() calls sys.exit(r.returncode) on the first non-zero return, so on any machine without that tree the build hard-fails at step 4 of 9, before the study, the deck, the dashboard or the transfer pack are regenerated. README.md §6 instructs the next operator to run exactly this command. Introduced by the fixes for ANOM-07 (recalculation step) and CI-014 (PDF render). It also sits squarely against the GOV-F4.2 portability posture that Appendix A gate 10 and checker C20 assert — C20 scans five text files and the study, and no build script.

NEW-03 — LOW — DEF-004's correction record claims more than the artefacts show

DEF-004 records that the corrected employed-manager figure is read by "Study, dashboard and decision pack". I extracted all text from NDIS_Decision_Pack_v1.0.pptx (10 slides): it contains no occurrence of 411, and none of 126 either. The deck states no employed-manager figure at all, so it cannot be reading from the new function. The deck is not wrong — its other figures ($15,875, $21.30, $6.70, −$2.07) are all current — but the fix record overstates its own reach, which is the ANOM-05 pattern in miniature.

NEW-04 — MEDIUM-HIGH — The evidence register attributes to V3 the evidence for four requirements V3 failed

03_Registers/EVD.csv records 15 rows produced by V3 (13 solely, 2 jointly), among them EVD-001 (REQ-SYS-01), EVD-010 (REQ-SYS-10), EVD-022 (REQ-CON-04) and EVD-025 (REQ-MOE-01). All four are recorded as FAIL in the V3 report that Appendix A gates 2 and 3 now cite as their evidence. Two of the four have since been genuinely fixed (REQ-CON-04, REQ-MOE-01); two have not (REQ-SYS-01, REQ-SYS-10). Either way the register asserts that the independent verifier produced evidence proving requirements the independent verifier's own report says are unproven. This is DEF-008's root cause — an outcome attributed to a verifier who did not produce it — surviving in a register that the appendix fix did not reach, and "26 of 26 verified" rests on it. Two of the same rows also cite evidence that cannot be opened: EVD-001 names "NDIS_Financial_Model_v1.0.xlsx sheet 'Capital'" (no such sheet exists — the workbook has README, Inputs, Drivers, Costs, UnitEconomics, BreakEven, Scenarios, SupportCoordination, Sensitivity) and EVD-015 names "01_System/build_dashboard.py" (no such file — the builder is build_web.py). C16 and C17 test ID linkage only, never path resolvability.

NEW-05 — MEDIUM-HIGH — REQ-MOP-01's checker tests something weaker than REQ-MOP-01

The rewritten requirement reads: "The proportion of paragraphs and table cells in the delivered study that contain a currency amount and also cite a SRC or ASM identifier shall be at least 90 percent." C27 reports 1.000 over 86 units. I re-ran C27's own extraction with its own exemption list and then applied the requirement's own test — SRC or ASM only:

TestAttributedRatio
C27 as coded (SRC, ASM, REQ, DEC, RSK, DEF, BKL, ACT, "Financial Model", "model_params")86 / 861.000
REQ-MOP-01 as written (SRC or ASM only)74 / 860.860 — below the 0.90 threshold

The twelve units carrying the difference are attributed to the project's own artefacts, not to a source — for example "Contribution per billable hour [Financial Model, UnitEconomics] $21.30 $6.70". Citing the model for a derived figure is defensible practice; the problem is that the requirement says SRC or ASM and the checker that certifies it does not. The acceptance criterion has also been written as "PASS if checker C27 reports an attribution ratio >= 0.90", which makes the criterion true by definition of the checker rather than by the property the requirement names. Zero of the 86 units carry no reference at all, which is a real and substantial improvement over the first build — the finding is a definition mismatch, not a research failure.


PART 3 — INDEPENDENT ARITHMETIC ON THE CHANGED FIGURES

Recomputed from register values, not read from the project's outputs.

Employed-manager break-even (ANOM-01). ASM-011 manager cost $95,000/yr → $7,916.67/month. (847.3333 + 7,916.6667) / 21.30424 = 8,764.00 / 21.30424 = **411.37** hours/month. Project publishes 411.37 (model_params), "about 411" (study §1), "about 411" (dashboard ACT-001). ✔ The superseded figure 126.53 is the paid-coordinator case (847.3333 / 6.69693) and now appears only in the defect history. ✔

Support coordination six months, no revenue (ANOM-02). one-off, workers = 0, registration excluded: 636 + 108 + 139.20 + 290 + 2,500 = **3,673.20** 3,673.20 + 6 × 847.3333 = **8,757.20** → study reads $8,757. ✔ (model_params returns 8,757.18 because it rounds the monthly figure to 847.33 before multiplying — a 2-cent presentation difference, not an error.) Prior published value $10,044.78; overstatement removed: $1,287.58.

Fixed-cost band (ANOM-06). low (45 + 78 + 200 + 150) + (342 + 1,200)/12 = 473 + 128.50 = **601.50** base (45 + 78 + 300 + 200) + (342 + 2,350)/12 = 623 + 224.33 = **847.33** high (125 + 143 + 500 + 350) + (342 + 3,500)/12 = 1,118 + 320.17 = **1,438.17** Delivered workbook cached values: B28 = 601.5, C28 = 847.333…, D28 = 1,438.167. Exact match, in the artefact itself and not only in the Python. ✔

Unchanged figures re-confirmed against the recalculated workbook (previously unverifiable, since the file carried no values): gross margin UnitEconomics!B11 = 21.30424; paid-admin contribution B14 = 6.69693; break-even BreakEven!B6 = 39.7730, B7 = 126.5257; hours to draw $5,000 B8 = 274.468, B9 = 873.137; working capital B12 = 7,318.61, 30-day B13 = 15,682.73; SC net per year 98,664.15. All match my own arithmetic to the cent. ✔

ASM-008's new figures. 21.3042 − 58.4292 × 0.40 = **−2.07** ✔ · 21.3042 − 58.4292 × 0.10 = **15.46** ✔

ASM-013's unchanged figures (still wrong). 55%: 38 × 0.55 × 52/12 = 90.57 hr/mo × $100.14 = $9,069.35/mo → gross $108,832/yr, net $98,664/yr. 35%: 57.63 hr/mo → $5,771.40/mo → gross $69,257/yr, net $59,089/yr. Register says $71,000 gross and $112,000. Neither reconciles on any reading.

Open items. RQ 0 + ACT 9 + CR 0 + DEF 0 + RSK 10 + ISS 1 + BKL 6 = **26** — matches README, dashboard tile and HANDOVER.txt. The tile's own parenthetical sums to 27 (NEW-01).

Mobile legibility. 2,205 runs in the delivered study: 835 at 10.0pt, 275 at 10.5pt, 2 at 11.5, 1 at 14, 1 at 26, and 1,091 inheriting — all from Normal (10.5pt), Heading 1 (17pt) or Heading 2 (13pt). The 8.0pt Body Text 3 and 9.0pt Caption styles exist in the template but are used by no run. Minimum effective size 10.0pt; 28 tables, maximum width 5 columns; PDF renders to 47 pages. ✔


PART 4 — COLD-START TEST, ITEM (c) RE-RUN

Item (c) was the one cold-start failure: what is open, who owns it, and what the next action is.

Now PASS, with one residual defect. From the folder alone: README.md says 26 open items; the dashboard's first tile says 26; 00_Handover/HANDOVER.txt says "Total open: 26"; the canonical function returns 26; the published lineage formula produces 26. One number, four artefacts, agreeing — and C11 now fails the build if the README or the tile drifts from the registers. Ownership is answerable: 9 actions on Zaid, 4 backlog items on the AI, 10 open risks, 1 open issue, 1 decision parked with Zaid. The next action is unchanged and clear.

Two things still stop this being clean. The tile's parenthetical breakdown sums to 27 (NEW-01). And two of the 26 counted items appear nowhere on the dashboard: BKL-002 (first-party read of the SCHADS pay guide) and BKL-003 (Victorian payroll tax threshold and disability-sector WorkCover rate) are open and owned by Zaid, but the "On Zaid" list renders only the 9 ACT rows and the "On the AI" list filters BKL to owner != Zaid. The count says 26; the page displays 24 of them. Both are material to the confidence of the model's cost side. Items (a), (b), (d) and (e) re-confirmed as PASS, with the (e) caveat now much reduced: running the checker would today catch the open-item drift, the workbook divergence, the font floor and the employed-manager figure.


PART 5 — RE-ISSUED VERDICTS ON THE EIGHT FAILED REQUIREMENTS

REQ IDWasNowBasis
REQ-SYS-01FAILFAILStatement and AC unchanged: capital "for each candidate service model". Study Table 5.1 still gives capital for two of the three candidates scored in §7; supported accommodation has none. DEC-005 now records the exclusion as a decision — a real improvement over a silent omission — but the requirement was not amended to match, so the AC is still not met. Its evidence row EVD-001 also points at a workbook sheet that does not exist (NEW-04).
REQ-SYS-10FAILFAILTable 8.1 gives break-even for core supports in two administration modes; SupportCoordination!B13 gives 8.5 hr/month for the second model; supported accommodation still has none. Same root cause, same unamended AC.
REQ-SYS-15FAILPASSThe tile that previously mapped to no register state now reads project_data.open_item_count(); "On the AI — 4 open" now resolves the BKL owner column correctly (was 0 of 6). No value on the page is hand-entered, which is what the AC tests. NEW-01 and the two undisplayed BKL rows are logged as defects but do not breach this AC.
REQ-SYS-16FAILFAILMETRIC_LINEAGE.csv is unchanged at 10 rows. The dashboard still publishes $6.70 and −$2.07 as standalone tiles with no lineage entry, and the "Gross margin per billable hour" formula in words ("price limit minus the casual wage multiplied by the on-cost multiplier") does not produce either — both require the Level 3 administration deduction. C11 was extended to reconcile counts but still never compares the tile set to the lineage set, so the AC's "metrics-without-a-lineage-entry equals zero" remains untested and unmet.
REQ-CON-01FAILFAIL (materially improved)Every one of the 86 dollar-bearing units in the study now carries a reference — zero unattributed, against several before. But the AC covers "every external fact, figure, price, rate, date and legal position", and the checker that certifies it now looks only at units containing a $. Non-currency external facts are still untested and instances remain: §9 recommendation 6 states "a national utilisation rate of 53.7%" and "4,638 places" with no inline reference, and §7.1 states "269,000+ providers" the same way. All three are sourced elsewhere (SRC-050, SRC-045), so this is an attribution gap in the prose, not an unsourced claim.
REQ-CON-04FAILPASSThe statement was rewritten into a measurable test and the artefact genuinely meets it. Independently measured over all 2,205 runs: minimum effective size 10.0pt, no run below it, 28 tables with a maximum of 5 columns. The document grew 32 → 47 pages, which is the honest cost of the fix. C29's own measurement is narrower than its description (1,114 of 2,205 runs), but the requirement holds on the full population.
REQ-MOE-01FAILPASSACT-009 now exists, owned by Zaid, due 2026-09-05, explicitly blocking "the base case in the financial model, the service model recommendation, and the go/no-go itself (ASM-008)". The study's own highest-value pre-commitment question now carries an owned, dated action instead of sitting inside a roadmap that only activates after the decision. This was the sharpest gap in the first pass and it is properly closed.
REQ-MOP-01FAILFAILPasses its acceptance criterion (C27 reports 1.000 ≥ 0.90) and fails its own statement (SRC/ASM-only attribution = 0.860 < 0.90). The two do not test the same property. See NEW-05.

Tally: 21 PASS, 5 FAIL (was 18/8). No previously-passing requirement regressed: I re-checked the two rewritten statements (REQ-SYS-03's three alternatives, timestamps 09:40 before 11:05, sensitivity and rationale all present; REQ-SYS-08's five items each naming a register ID and a checkable party), and re-ran the register integrity scans — 53 sources all accessed 2026-08-20 with no empty "why below High", 15 assumptions all carrying confidence and failure consequence, 26 requirements each with exactly one "shall" and no banned qualifier.

The project still publishes "26 of 26 requirements verified" on the dashboard and in Appendix A gate 1. On independent verification against the recorded acceptance criteria it is 21 of 26. That claim, and the EVD rows underneath it (NEW-04), are the last piece of the assurance layer that has not been brought into line with the evidence.


PART 6 — VALIDATION AGAINST PURPOSE

"Zaid can decide, on evidence rather than impression, whether to commit personal capital to starting an NDIS provider business in Victoria."

A reasonable person can now make this decision from these artefacts. That is a change from my first report, and it rests on three specific things rather than on the volume of corrections.

First, the three defects that actually distorted the decision are gone and I have recomputed each of them. The employed-manager branch of ACT-001 — the question the study itself calls the single biggest swing factor — now says 411 hours rather than 126, so Zaid is being shown the true consequence of the option he is being asked to choose. The support-coordination capital figure is $8,757 rather than $10,045, so the option the trade study ranks first is no longer penalised by a phantom $1,288 of costs for workers it does not employ. And ASM-008 now records that at 0.40 administration hours every billable hour loses money, instead of recording a survivable $10.85 — which matters because that single assumption spans viable to unviable and a reader consulting the register would previously have been reassured by it.

Second, the highest-value thing Zaid can do before spending anything now has his name and a date on it. ACT-009 turns "find a provider who will tell you their administration ratio" from a line in a post-decision roadmap into an owned action due 2026-09-05. Combined with ACT-005 to ACT-008, the study's recommendation — spend two weeks and no capital closing the four cheapest uncertainties — is now fully instrumented.

Third, the assurance layer is no longer self-certifying in the places that mattered. C09 opens the workbook; C27 tests attribution at the unit a reader reads; C28 would fail the build if the delivered file carried no values or diverged from the model; C29 measures the font floor; C11 fails the build if the open-item count drifts. Appendix B2 publishes all twelve of my findings, with severities, unprompted. A project that publishes the verifier's criticism next to the deliverable is a project whose remaining claims can be weighed.

What still qualifies the decision. Supported accommodation is scored as a candidate and costed as none — the exclusion is now a recorded decision (DEC-005) rather than a silence, but a reader comparing three options is comparing two sets of numbers and one argument. Two of the twenty-six open items are invisible on the dashboard, and both concern the confidence of the wage and tax inputs. ASM-013's support-coordination earnings figures still contradict the model by roughly $3,000 a year in both directions, in the register a careful reader would consult when weighing the recommended option. And the claim of 26 of 26 verified is five requirements ahead of the evidence. None of these changes the direction of the recommendation; all of them are the kind of thing that erodes trust in it when found later rather than now.

Decision-readiness: the artefacts now support the decision the study actually recommends — proceed to the four free confirmations and ACT-009, commit no capital yet, lodge registration in month one whichever model is chosen, and do not start with SIL or SDA. They do not yet support a capital commitment on the supported-accommodation option, because that option was never costed.


FINAL VERDICT

PASS WITH DEFECTS — fit to deliver, not yet fit to claim a clean baseline. Ten of twelve anomalies fully fixed, two partially; all three decision-distorting errors corrected and independently recomputed; the cold-start failure closed; requirement verdicts 21 PASS / 5 FAIL, up from 18 / 8. Five defects remain open on my count: NEW-01 (repaired tile contradicts itself, 26 v 27), NEW-02 (the documented rebuild command hard-fails on any other machine), NEW-04 (evidence register attributes to V3 evidence for requirements V3 failed, two of them citing files and sheets that do not exist), NEW-05 (REQ-MOP-01's checker tests a weaker property than the requirement states; strict ratio 0.860), and the ASM-013 residual under ANOM-03. None is decision-distorting. The one claim that should be withdrawn or qualified before delivery is "26 of 26 requirements verified" — independently it is 21 of 26, and the four EVD rows that underwrite the difference name the verifier who failed them.

Verified by V3, independent verifier. No artefact under verification was modified, and the checker suite was not run in order to avoid altering the pack and the run log. This file and the first report are the only outputs written.