V3 — INDEPENDENT DELIVERABLE VERIFICATION
Agent: V3, independent verifier (GOV-C3.2, GOV-E2.3). Did not build any artefact in this project. Date of verification: 2026-08-20 · Baseline verified: v1.0.0 · Project class: STANDARD Scope: the whole delivered set at /home/claude/ndis. Method: cold-start test first; then requirement-by-requirement verification against the recorded acceptance criteria; then adversarial cross-artefact and arithmetic checks; then validation against purpose. Nothing in this project was edited. The only file written is this one.
PART 1 — COLD-START TEST (GOV-F9.5)
Run before reading anything else, from the folder alone, time-boxed to ~15 minutes. Everything below was answerable from README.md and Dashboard.html without opening a deliverable.
| # | Question | Result | What answered it, and what was wrong |
|---|---|---|---|
| (a) | Purpose | PASS | README.md §1, first line: "Zaid can decide, on evidence rather than impression, whether to commit personal capital to starting an NDIS provider business in Victoria." Unambiguous, and it states what the project did not do. |
| (b) | What is baselined, at what version | PASS | README.md header table: baseline v1.0.0, class STANDARD, jurisdiction Victoria, governance v3.7, facts current at 2026-08-20 and stale after 2027-08-20. Confirmed against 03_Registers/REQ.csv (every row baselined v1.0.0) and the Dashboard.html subtitle. |
| (c) | What is open, who owns it, what the NEXT ACTION is | FAIL — documentation defect | The next action is clear and owned (README.md §5 and the dashboard banner: answer ACT-001 and ACT-002, owner Zaid, due 2026-08-27). But the count of what is open is stated three different ways and no two agree: README says 19, the dashboard's first tile says 28 and instructs the reader "the first number to read", and the published lineage formula in 03_Registers/METRIC_LINEAGE.csv ("count of rows whose status is not Closed, in RQ, ACT, CR, DEF, RSK, ISS and BKL") yields 25. A cold-start reader cannot state what is open. Root cause is diagnosed at ANOM-04. |
| (d) | What you must NEVER do | PASS | README.md §3 lists INV-1 to INV-7 as inviolable and non-waivable, with the prompt-injection rationale. Repeated on the dashboard under "What must never happen". This is the clearest part of the pack. |
| (e) | How to verify your own work | PASS, with a material caveat | README.md §6: run python 01_System/checkers/run_checks.py; change only 01_System/project_data.py and rebuild with build_all.py. The instruction is findable and executable. Caveat: following it would not have surfaced a single one of the twelve anomalies below. Several checks do not test what their own descriptions claim (ANOM-08). "26 of 26 checks passed" is not evidence that the artefacts agree. |
Cold-start verdict: four of five clean, one defective. Item (c) is a documentation defect, not a user error. A stranger can pick this project up and act; they cannot state its open position correctly.
PART 2 — REQUIREMENT VERIFICATION
Every requirement in 03_Registers/REQ.csv, verified against its own recorded acceptance criterion. The project's own "Verified" status was disregarded; each was re-checked. 18 PASS, 8 FAIL.
| REQ ID | Result | Evidence examined | Note |
|---|---|---|---|
| REQ-SYS-01 | FAIL | Study §5.1 Table 5.1 (docx); Costs sheet of NDIS_Financial_Model_v1.0.xlsx; model_params.ONE_OFF_COSTS; study §7 trade study | Capital is stated for two of the three candidate service models. REQ-SYS-04 names core supports, support coordination and supported accommodation as candidates; SIL/SDA carries no capital figure anywhere (disclosed as BKL-005, but the AC grants no exemption). Separately the support-coordination capital row does not reconcile with its own components — see ANOM-02. Cost lines themselves do all carry SRC/ASM references. |
| REQ-SYS-02 | PASS | Study §2 and §6 (verification 4–6 months, certification 9–12, SRC-021); RAMP_BASE and ASM-015 (first client month 4); roadmap §10 week-by-week | Month counts and a cited lead-time source exist for the two modelled models; SIL/SDA carries the certification 9–12 month figure with SRC-021. |
| REQ-SYS-03 | PASS | Study §6.1–6.2: alternatives A1/A2/A3, weights block, scores, sensitivity, rationale | Three alternatives, weights, sensitivity and rationale all present. Qualification: the AC requires weights "timestamped before the scores block"; every artefact in the project carries the single date 2026-08-20, so the ordering is asserted, not evidenced (ANOM-11). Also see ANOM-10 on the word "credible" in the requirement statement. |
| REQ-SYS-04 | PASS | Study §7 and §7.1; DEC-003 in DEC.csv | All three models scored on identical weighted criteria; sensitivity stated (3.70 v 3.30, flipping to 3.40 v 3.45 at a 40% regulatory-stability weight); escalated rather than decided under GOV-B6.4. This is the strongest single piece of work in the pack. |
| REQ-SYS-05 | PASS | UnitEconomics!B11 = SUM(B5:B10), each term referencing Inputs!B5/B8/B10/B11/B12/B13; Inputs column D = SRC-002, SRC-009, SRC-011, SRC-012, SRC-013 | No hard-coded margin anywhere. Independently recomputed: $21.30 (Part 4). |
| REQ-SYS-06 | PASS | Study §4.1 obligations table (authority, cost, frequency, lead time, SRC per row); SRC-025 to SRC-042 | No empty cells found in the delivered table. |
| REQ-SYS-07 | PASS | Scenarios sheet rows 37–46 (DOWNSIDE, twelve months, zero hours); model_params.runway_downside(); $15,875 in study §1 finding 3, README and dashboard | Month-by-month position and a single runway figure both present, and the figure recomputes exactly. Caveat at ANOM-07 (the workbook holds no cached values). |
| REQ-SYS-08 | PASS | Study §9 box "WHAT WOULD CHANGE THIS ANSWER" — five named, checkable items with confidence levels | Exceeds the three required. Genuinely falsifiable, e.g. item 1 names the number that would invert the recommendation. |
| REQ-SYS-09 | PASS | Study §10, explicitly labelled conditional; sequences ASIC, ABN/TFN/GST/PAYG, myID+RAM, insurance, policy manual, registration lodgement, screening, first aid | Every dependency named in §4.1 appears in the sequence. |
| REQ-SYS-10 | FAIL | BreakEven!B6, B7; SupportCoordination!B13; study Table 8.1 | Break-even is computed for two of the three candidate models. Nothing for supported accommodation. Same root cause as REQ-SYS-01: the requirement set is internally inconsistent about how many candidates exist. |
| REQ-SYS-11 | PASS | BreakEven!B12 = Drivers!B14*(Inputs!B8*(1+B10+B11+B12+B13))*(Drivers!B6/30) | Derived, not asserted — from wage rate, hours driver and the stated 14-day lag (ASM-007). Recomputed to $7,318.61 (Part 4). |
| REQ-SYS-12 | PASS | Study §1 finding 1, §3.1, §7 (per-model July-2027 consequence); §5.1 and roadmap §10 for cost to be ready | Each modelled candidate carries a stated consequence and a readiness cost; SIL/SDA carries "already in force". |
| REQ-SYS-13 | PASS | 10 PNGs in 02_Work/diagrams/; 10 images embedded in word/media/ of the study (verified by unzip) | Note: checker C12 only counts images ≥10; it does not verify that each process narrative maps to a diagram. Verified by hand for D1, D3, D4, D8. |
| REQ-SYS-14 | PASS | 09_Help_Hub/index.html HLP-10; checker C10 source read at run_checks.py:245-263 — nine literal regex probes over the actual file, not a self-report | Patterns are loose (e.g. looks wrong) but they are real searches over the delivered page. |
| REQ-SYS-15 | FAIL | 01_System/build_web.py:130-138 and :204-211; Dashboard.html status tiles | The dashboard is generated by a script (that half of the AC holds), but the AC also requires that every value maps to a register row. The "28 OPEN ITEMS" tile maps to no register state — it counts three closed change requests as open — and "On the AI — 0 open" contradicts six open BKL rows. See ANOM-04. |
| REQ-SYS-16 | FAIL | METRIC_LINEAGE.csv (10 rows); dashboard tiles (9 values); checker C11 source at run_checks.py:265-272 | AC: "metrics-without-a-lineage-entry equals zero." The dashboard publishes $6.70 and -$2.07 as standalone tiles; neither has a lineage row (only "Gross margin per billable hour" exists). C11 never tests the mapping — it only checks that lineage rows are non-empty and that a heading string exists on the page. The criterion is unproven and, on inspection, not met. |
| REQ-SYS-17 | PASS | Checker C22/C23 source read at run_checks.py:386-433 | Genuine: C22 verifies SHA-256 per file against MANIFEST.txt and detects MISSING, ALTERED and EXTRA. C23 really appends bytes to a live pack file, re-hashes, and restores in a finally block. This is the one place where the project's assurance claim is fully supported by code I read. |
| REQ-SYS-18 | PASS | Independent scan of SRC.csv: 53 rows, 20 High, 33 below High, zero with an empty "Why below High" | Re-computed by me, not taken from the checker. |
| REQ-CON-01 | FAIL | Checker C05 source at run_checks.py:100-162; study §9 recommendation 6 | The AC requires the checker to report zero material external claims without a SRC-### reference. C05 does not do that. It extracts numeric tokens from the study and asks whether each number appears somewhere in any register or computed result — it never tests whether a claim carries a reference. It also pre-loads range(0,101) into the allowed set and discards every integer below 100. A wrong figure passes as long as the same wrong figure is in a register (which is exactly how ANOM-03 survived). Instances exist in text: study §9 recommendation 6 states "a national utilisation rate of 53.7%" and "4,638 places" with no inline reference. |
| REQ-CON-02 | PASS | Study §4.1 (WorkSafe Victoria, Victorian Worker Screening, Victorian WWCC, Victorian Portable Long Service Authority, Victorian payroll tax); SRC-012, SRC-013, SRC-031, SRC-032, SRC-033 | No obligation stated for another state. |
| REQ-CON-03 | PASS | Independent scan of all 53 SRC rows: every Accessed value = 2026-08-20; every Source URL begins http | Re-computed by me. |
| REQ-CON-04 | FAIL | python-docx inspection of NDIS_Business_Enabling_Study_v1.0.docx: 27 tables, none over six columns — but 973 runs set at 8.5 pt inside tables | AC: "no body text is below 10 point." Nearly all of the study's substance lives in those tables. Appendix A gate 13 asserts "no body text below 10pt" as PASS; the file contradicts it. Nothing mechanical tests this requirement. |
| REQ-CON-05 | PASS | Independent scan of ASM.csv: 15 rows, all with a confidence level and a non-empty "What breaks if it is wrong" | Criterion met. Separately, two of those failure-consequence fields state figures that contradict the model (ANOM-03, ANOM-06) — that is a content defect, not an AC breach. |
| REQ-CON-06 | PASS | Checker C08 source at run_checks.py:190-221 (real regex sweep for credentials, 9-digit identifiers, card numbers, BSBs, private keys across the controlled set), plus my own spot scan | No credential, identifier or third-party personal datum found. |
| REQ-MOE-01 | FAIL | ACT.csv (8 rows), BKL.csv (6 rows), study §9 recommendation 2, study §10 roadmap week 1-2 | AC: "every open decision-blocking question carries an ACT-### with an owner and a due date." The study's own recommendation 2 calls closing ASM-008 (administration hours) "the single highest-value thing you can do before committing capital", and the Sensitivity sheet says at 0.40 hr every billable hour loses money. There is no ACT for it — not in ACT-001..008, not in BKL-001..006. It appears only inside the §10 roadmap, which states in its own header that it "has no force unless and until you decide to proceed". The one question that most changes the go/no-go answer is filed behind the go decision. |
| REQ-MOP-01 | FAIL | Checker C05 source at run_checks.py:100-162; METRIC_LINEAGE.csv row "Sourced-claim ratio" | The published lineage says this ratio is "material external claims carrying a SRC-### reference, divided by all material external claims found by checker C05". C05 computes no such thing — it computes a number-recognition rate. The metric as defined has never been measured, so the ≥0.90 threshold is unproven. |
PART 3 — ANOMALIES
Twelve findings. Severity: HIGH = changes what a reader would conclude or decide, or falsifies a published assurance claim. MEDIUM = wrong or unsupported, contained. LOW = cosmetic.
ANOM-01 — HIGH — The study understates employed-manager break-even by roughly 3x, in the box that poses the decision to Zaid
Where: 05_Outputs/NDIS_Business_Enabling_Study_v1.0.docx §1, box "FOUR THINGS I NEED FROM YOU", ACT-001 line. The study says: "Owner-operator break-even is about 40 billable hours a month. An employed manager pushes it past 126." Dashboard.html, ACT-001 row, says of the same question: "an employed manager pushes it past 400." The project's own ASM-011 says a passive-investor model "adds roughly $95,000 per year of fixed cost and moves break-even by about 370 billable hours per month" — i.e. 39.8 + 371.7 ≈ 411 hours. The dashboard is right; the study is wrong. 126 is the paid-administration-coordinator break-even, a different scenario. ACT-001 is described in the same document as "the single biggest swing factor in the model". The consequence of one of its two branches is misstated by a factor of three, in the artefact Zaid is told to read first.
ANOM-02 — HIGH — Arithmetic error: the support-coordination six-month capital figure is overstated by $1,288 and does not reconcile with its own row
Where: study §5.1 Table 5.1, row "Support coordination, no registration, no workers"; 01_System/model_params.py:169-174. The row reads: one-off $3,673, fixed cost $847 a month, six months with zero revenue $10,045. $3,673.20 + 6 × $847.33 = $8,757.20, not $10,045. Root cause: runway_downside() calls one_off_total(include_registration=include_registration) without passing workers=0, so the default WORKERS_AT_START = 3 applies and the figure silently includes worker screening and first-aid costs for four people ($139.20 + $290) × 4 = $1,716.80 — in the scenario whose own label is "no workers". The overstatement is 14.7%, and it falls on the option the trade study ranks first (3.70 v 3.30), making that option look more capital-hungry than the project's own model says it is.
ANOM-03 — HIGH — The Assumptions Register contradicts the study's headline finding, and reverses its sign
Where: 03_Registers/ASM.csv ASM-008; reproduced in study §12 "Assumptions, and what breaks if they are wrong". ASM-008 states: "At 0.40 hours per billable hour, gross margin per hour falls from $21.30 to about $10.85 and break-even roughly doubles." Every other artefact says the same input produces negative $2.07: study §1 finding 2, study Table 5.2 final row ("Every hour you sell would lose money"), the Sensitivity sheet row 6, the Drivers sheet note ("At 0.40 the core-supports margin turns NEGATIVE"), README.md headline findings, and the dashboard tile. Recomputed independently: −$2.07 is correct (Part 4). $10.85 is not derivable from any combination of the model's inputs. It is stranded from an earlier draft. The register that exists to record what breaks if an assumption is wrong states the wrong consequence for the project's single most dangerous assumption — and states it as a survivable margin rather than a loss.
ANOM-04 — HIGH — Three different open-item counts, two column-index bugs, and a dashboard that hides six open items
Where: README.md §current state; Dashboard.html first status tile; 01_System/build_handover.py:31-37; 01_System/build_web.py:130-138 and :211. Register truth: RQ 0 open, ACT 8, CR 0, DEF 0, RSK 10, ISS 1, BKL 6 → 25 by the formula published in METRIC_LINEAGE.csv. Published values: README 19, dashboard 28. Causes, both off-by-one column indexes:
BKLstatus is column 7; both builders read column 6 (Owner).build_handover.py:32testsb[6]=="Open"against a name, so all six open backlog items vanish from the README count.build_web.py:211does the same and prints "On the AI — 0 open" on the dashboard while six items are open and owned by the Master Brain.
ISSstatus is column 6;build_handover.py:37testsi[5]=="Open"against the Owner field, dropping the open issue from the README.
CRstatus is column 13;build_web.py:133testsx[12] != "Closed"against theVersionfield, so three closed change requests are counted as open — that is the entire 25→28 gap.
README 19 = 8 ACT + 10 RSK + 1 parked DEC. Dashboard 28 = 8 + 3 phantom CRs + 10 + 1 ISS + 6 BKL. Neither is correct, they are wrong in opposite directions, and the dashboard labels its number "the first number to read".
ANOM-05 — HIGH — Appendix A and Appendix B overstate; three gates and one audit line are not supported by the evidence they name
Where: study Appendix A (Definition of Done) and Appendix B (GOV Part I compliance audit).
- Gate 5, "No collateral damage", PASS: "Financial model recalculated: 398 formulas, zero errors. Every value reconciled cell by cell against model_params.py." The delivered workbook contains no cached values at all —
openpyxlwithdata_only=TruereturnsNonefor every formula cell, which is what a file written by openpyxl and never opened by a calculation engine looks like. No recalculation left any trace in the artefact. And the reconciliation claim is false: see ANOM-07.
- Gate 6, "Synchronised", PASS: "Checker C09 confirms the study, the model and the registers report identical figures. Zero contradictions between controlled artefacts." C09 (
run_checks.py:222-243) tests four strings for presence in the study text and the dashboard HTML. Its own inline comment concedes it never opens the workbook. ANOM-01 to ANOM-04, ANOM-06 and ANOM-07 are six contradictions between controlled artefacts.
- Gate 3 and Appendix B lines 42-49, PASS: "V3 the deliverables" and "Cold-start run by agent V3, which did not build the project." The delivered v1.0 study records the outcome of an independent verification and a cold-start test as already passed, in an artefact baselined before that verification ran.
02_Work/scratch/contains V1 and V2 records and no V3 record. GOV-I1.2 forbids marking a line PASS without naming the evidence; the evidence named here did not exist when the line was written.
- Appendix B lines 33-37, PASS: "both trade studies carry >=2 credible alternatives" — REQ-SYS-03 requires three. The audit tests a weaker bar than the requirement and reports PASS against it.
- Table B.1 caption: "All 49 lines pass. Any line that had failed would be named here." On this evidence at least three should have been named.
ANOM-06 — MEDIUM-HIGH — The delivered spreadsheet computes a different fixed-cost band from the study
Where: Costs!B19:D20 of NDIS_Financial_Model_v1.0.xlsx v model_params.RECURRING_MONTHLY v study Table 5.1. Study Table 5.1 states fixed cost per month "$602 to $1,438 (base $847)", which is exactly what model_params.py computes (bookkeeping 200/300/500, overhead 150/200/350). The workbook wires all three columns of both those lines to a single Drivers cell (=Drivers!$B$7, =Drivers!$B$8), so the delivered file computes $751.50 to $1,088.17 for the same band. The low case is overstated by $150 a month and the high case understated by $350 a month. README §6 tells Zaid the workbook is where he changes assumptions; the low/high columns there do not respond to the drivers they claim to model, and the range he would read differs from the range the study reports.
ANOM-07 — MEDIUM — "Every value reconciled cell by cell" cannot be true, and the workbook ships with no computed values
Where: 05_Outputs/NDIS_Financial_Model_v1.0.xlsx; study Appendix A gate 5. Two passes with openpyxl (formulas, then data_only=True) return formulas in the first and None in the second for every computed cell. Anyone reading the workbook without a calculation engine — any script, any preview pane, any mobile viewer — sees blank cells where the answer should be. Combined with ANOM-06, the cell-by-cell reconciliation claim is falsified by the first two cells I checked.
ANOM-08 — MEDIUM-HIGH — Four checks assert more than they test, and their descriptions are quoted as evidence elsewhere
Where: 01_System/checkers/run_checks.py.
- C09 ("no two controlled artefacts contradict"): four string-presence tests; never opens the model or the registers.
- C05 ("every material figure exists in the source register"): number-recognition, not claim-to-source attribution; pre-loads every integer 0–100 and drops all sub-100 integers. Wrong-but-registered figures pass, which is how ANOM-03 survived 26 green checks.
- C11 ("every metric on the dashboard has a lineage entry"): never compares dashboard metrics to lineage rows.
- C20 ("no drive letter or absolute machine path in any generated artefact"): scans five text files for
/home/,/mnt/user-dataandC:\Users\only. It does not scan the delivered study — which contains a Windows drive-letter path (ANOM-09).
"26 of 26 checks passed" is quoted in the study, the README, the dashboard and Appendix B as proof of consistency. The green result is real; what it proves is narrower than what is claimed for it.
ANOM-09 — MEDIUM — A Windows absolute path in the delivered study, under a PASS for cloud portability
Where: study Appendix A gate 10: "All deliverables in C:\AI Projects\NDIS Project\05_Outputs." The delivered folder has no such path; the README's own folder map is relative. Appendix B lines 21-25 marks "cloud portability … all files inside the project folder" PASS, and C20 reports "no machine-specific path in any generated artefact", while the study itself directs a reader to a machine-specific location that will not exist on their disk. Breaches GOV-F4.2 on the evidence of the artefact that claims compliance with it.
ANOM-10 — MEDIUM — Requirement statements carrying unquantified qualifiers, and a banned-word list too narrow to catch them
Where: 03_Registers/REQ.csv; checker C01 at run_checks.py:56-72. All 26 statements carry exactly one "shall" — that half of GOV-B3.6 holds, independently confirmed. The qualifier half does not: REQ-SYS-03 "at least three credible alternatives" (credible by whose test?), REQ-SYS-08 "the specific evidence", REQ-CON-04 "shall be legible on a mobile screen", REQ-MOP-01 "material external claims". C01's banned list is seven literal strings (and/or, etc., as appropriate, as required, user-friendly, robust, fast enough) and matches none of them. REQ-SYS-03 is the clearest breach: the acceptance criterion silently redefines "credible" as "counted", which is not the same test.
ANOM-11 — LOW-MEDIUM — Weights-before-scores is asserted, not evidenced
Where: study §6 preamble: "The weights below were set on 20 August 2026 before any alternative was scored." REQ-SYS-03's acceptance criterion requires "a weights block timestamped before the scores block". Every artefact and register row in the project carries the single date 2026-08-20, so no ordering can be demonstrated from the evidence. The claim may well be true; it is not verifiable from the delivered set, and Appendix B records it as PASS.
ANOM-12 — LOW — Register data defect and rounding inconsistencies
03_Registers/ISS.csv, ISS-001: theMitigationandOwnercolumns hold the same sentence, andStatusholds the valueZaid. The row is shifted, which is the proximate cause of half of ANOM-04. Appendix B's claim that every open item "carries an owner and a due date" is not met by this row — it has no due date field at all.
- Study §1 finding 4 states the 300-hour wage bill as $15,684 where Table 5.1 and the model give $15,683 (15,682.73). Same quantity, two roundings, one page apart.
- Break-even is "about 40" in the study §1 box and "about 39" on the dashboard (both round 39.77).
PART 4 — INDEPENDENT ARITHMETIC
Recomputed from SRC.csv and ASM.csv values only, without reference to the project's computed outputs.
On-cost multiplier (SRC-011 super 12.00%, SRC-012 WorkCover 1.80%, SRC-013 portable LSL 1.65%, ASM-003 payroll tax 0%): 1 + 0.1200 + 0.0180 + 0.0165 + 0 = 1.1545 → 1.1545 (15.45% on-cost). Matches the model. ✔
Fully loaded wages (SRC-009): L2 45.28 × 1.1545 = 52.2758 → $52.28/hr ✔ · L3 50.61 × 1.1545 = 58.4292 → $58.43/hr ✔
Gross margin per billable hour (SRC-002 price $73.58): 73.58 − 52.2758 = 21.3042 → $21.30 ✔ · as % of price 21.3042 / 73.58 = 28.95% → study's "29.0%" ✔ Component check: 45.28 + (45.28×0.12 = 5.4336) + (45.28×0.018 = 0.8150) + (45.28×0.0165 = 0.7471) = 52.2757 ✔
Contribution with paid administration (L3 loaded, ASM-008 at 0.25 hr): 58.4292 × 0.25 = 14.6073; 21.3042 − 14.6073 = 6.6969 → $6.70 ✔ At 0.40 hr: 58.4292 × 0.40 = 23.3717; 21.3042 − 23.3717 = −2.0674 → −$2.07 ✔ → ASM-008's "about $10.85" is not reproducible from any input combination. ANOM-03 confirmed.
Monthly fixed cost: base (45 + 78 + 300 + 200) + (342 + 2350)/12 = 623 + 224.33 = 847.33 → $847 ✔ low (45 + 78 + 200 + 150) + (342 + 1200)/12 = 473 + 128.50 = 601.50 → $602 (study) ✔ high (125 + 143 + 500 + 350) + (342 + 3500)/12 = 1118 + 320.17 = 1438.17 → $1,438 (study) ✔ Same bands from the delivered workbook: low 623 + 128.50 = 751.50; high 768 + 320.17 = 1088.17 → $751.50 to $1,088.17 ≠ $602 to $1,438. ANOM-06 confirmed.
Break-even billable hours per month, both ways: owner does the administration 847.33 / 21.3042 = 39.77 → 39.8 hr/month administration is paid 847.33 / 6.6969 = 126.53 → 126.5 hr/month ✔ both match employed manager (ASM-011, +$95,000/yr = $7,916.67/mo) (847.33 + 7,916.67) / 21.3042 = 411.4 → 411 hr/month → the study's "past 126" for this case is wrong; the dashboard's "past 400" is right. ANOM-01 confirmed. To draw $5,000/month: (847.33 + 5,000) / 21.3042 = 274.5 → 274 hr ✔ · paid admin 5,847.33 / 6.6969 = 873.1 → 873 hr ✔
Working capital (300 hr/month, ASM-007 14-day lag): 300 × 52.2758 = 15,682.73 wage bill; × 14/30 = 7,318.61 → $7,319 ✔ · at 30 days → $15,683 ✔
Six-month runway, core supports, no revenue, no drawings: one-off 636 + 108 + 0 + 0 + (139.20 × 4 = 556.80) + (290 × 4 = 1,160) + 0 + 2,500 + 1,080 + 4,750 = 10,790.80 10,790.80 + 6 × 847.33 = 15,874.78 → $15,875 ✔ · with $6,000/mo drawings 10,790.80 + 6 × 6,847.33 = 51,874.78 → $51,875 ✔
Six-month runway, support coordination: one-off as labelled (no workers, no registration) 636 + 108 + 139.20 + 290 + 2,500 = 3,673.20 3,673.20 + 6 × 847.33 = 8,757.20 → $8,757, against the study's $10,045. The published figure implies a one-off of 10,044.78 − 5,084 = 4,960.80, which is the 3-worker one-off with registration stripped out — the default-argument bug at model_params.py:169. ANOM-02 confirmed, overstatement $1,287.58 (14.7%).
Support coordination annual (38 hr/wk × 55% utilisation, SRC-005 $100.14): 38 × 0.55 = 20.9 hr/wk; × 52/12 = 90.57 hr/mo; × 100.14 = $9,069.35/mo; − 847.33 = $8,222.02/mo; × 12 = $98,664/yr ✔ (study §1 finding 5) At 35%: 57.63 hr/mo → $5,771.40/mo revenue → $69,257/yr gross → $59,089/yr net. → ASM-013's "roughly $71,000 a year gross instead of $112,000" matches neither the gross ($69,257 / $108,832) nor the net ($59,089 / $98,664) figures the model produces. Second stranded register figure — see ANOM-03's pattern.
Open items: RQ 0 + ACT 8 + CR 0 + DEF 0 + RSK 10 + ISS 1 + BKL 6 = 25, against README 19 and dashboard 28. ANOM-04 confirmed.
Checks that reproduced exactly (11): on-cost multiplier, both loaded wages, gross margin and its percentage, $6.70, −$2.07, $847, $15,875, $51,875, $7,319, $15,683, 274 and 873 hours. The model's core arithmetic is sound. Checks that failed (4): ASM-008's $10.85, ASM-013's $71k/$112k, the SC runway $10,045, the workbook's fixed-cost band.
PART 5 — VALIDATION AGAINST PURPOSE (GOV-C3.8)
Stated purpose: "Zaid can decide, on evidence rather than impression, whether to commit personal capital to starting an NDIS provider business in Victoria."
Judgement: the deliverable substantially serves its purpose, and is not yet safe to decide from unaided.
What it genuinely delivers. A reasonable person reading §1, §5 and §9 would come away knowing: that roughly $15,875 of business cash is at risk over six months before any revenue, plus $7,319 to $15,683 of working capital that the earlier plan had no line for at all; that gross margin is $21.30 an hour and that the entire business case turns on who does the administration, not on price; that the unregistered route into core supports expires in July 2027 and the application therefore starts in month one; and that the service-model choice is a genuine coin-toss escalated rather than concealed. The trade studies record their weights, their sensitivity and the exact weight at which the answer flips. §9's "what would change this answer" is a real falsification list, not a disclaimer. The honesty about what could not be verified — §13, ISS-001, eleven Low-confidence assumptions named as such — is well above the standard of most commercial feasibility work. On the specific adversarial test of whether Low-confidence assumptions are dressed as facts: they are not. ASM-008 is flagged as "the model's most dangerous assumption" in four separate places, and the Sensitivity sheet states plainly that "there is a plausible, unfalsified version of this business in which every hour of core support you sell LOSES money."
What blocks the decision.
- The pack's own assurance is unreliable, and that is worse than a wrong number. A reader is told 26/26
requirements verified, 26/26 checks passed, 49/49 compliance lines passed, "zero contradictions between controlled artefacts". I found six contradictions in under an hour, three of them decision-relevant. A reader who trusts the assurance will not re-derive anything; a reader who discovers ANOM-01 will not trust the rest.
- Two of the four inputs the study says it needs are still missing, by the study's own admission (ACT-001
to ACT-004). Until they are answered the artefacts report bands, not a number — the study says so itself, and that is a correct and honest position, not a defect.
- The single highest-value pre-commitment task has no owner and no due date. Closing ASM-008 is what
determines whether core supports makes $6.70 an hour or loses $2.07. It sits only in a roadmap that, by its own header, does not activate until after the decision it is supposed to inform (REQ-MOE-01 FAIL).
- One of the three candidate models was never costed. Supported accommodation is scored in the trade study
and eliminated on qualitative grounds; it has no capital figure and no break-even (BKL-005 discloses this). The elimination is well-argued, but "the evidence says no" and "we did not model it" are different statements.
Could a reasonable person decide from these artefacts? They could reach a well-supported conditional position — proceed to the four free phone calls, do not spend capital yet, do not start with SIL or SDA, lodge registration in month one whichever model is chosen. That is exactly what the study recommends, and it is the right answer on this evidence. They could not yet commit capital on it: the employed-manager branch of ACT-001 is misstated by 3x, the support-coordination capital figure is 15% overstated, and the assumption register misstates the downside of the assumption that decides the whole case.
What is missing to close the gap — none of it requires new research: correct ANOM-01, ANOM-02, ANOM-03 and ANOM-06; reconcile the open-item counts and fix the three column-index bugs; raise an ACT for ASM-008 with an owner and a due date; correct or withdraw the four overstated Definition-of-Done and compliance lines; and either cost the supported-accommodation model or restate the requirement set as two candidates rather than three.
FINAL VERDICT
CONDITIONAL PASS — DO NOT TREAT AS BASELINE-CLEAN. 18 of 26 requirements PASS; 8 FAIL — REQ-SYS-01, REQ-SYS-10, REQ-SYS-15, REQ-SYS-16, REQ-CON-01, REQ-CON-04, REQ-MOE-01 and REQ-MOP-01, of which REQ-SYS-01/10 share one root cause (the uncosted third candidate model) and REQ-CON-01/REQ-MOP-01 share another (checker C05 does not implement the criterion it is cited for). The cold-start test passes on four of five items. The research, the sourcing discipline, the trade studies and the core financial arithmetic are sound and in several places exemplary. The assurance layer is not: the Definition of Done, the Part I compliance audit and four checker descriptions claim more than the evidence supports, and the artefacts contradict each other on six figures — four of which reach the reader's eye on page one of the study, on the dashboard's first tile, or in the box that asks Zaid the decisive question. v1.0.0 should not be presented as verified until ANOM-01 through ANOM-06 are corrected and the affected Definition-of-Done and compliance lines are restated.
Verified by V3, independent verifier. No artefact under verification was modified. This file is the only output written.