Investment Plans workspace
Open raw ↗

V5 — NEGATIVE TEST RECORD, BUSINESS LAYER CHECKS

OWNER: 02_Work/scratch/V5_negative_tests_business_layer.md REVIEW: whenever C33, C34, C35 or the business-plan generator changes. Date: 2026-09-07 Rule: GOV-F3.12 and GOV-E2.4 — a check is only a check once it has been PROVEN able to fail.

A green check that has never been seen red is an assumption wearing a checkmark. Each test below broke one specific thing, recorded the failure verbatim, and restored the file.


Test 1 — the generator's own twelve checks (C-BP1..C-BP12)

Tampered: removed the metric from one value-proposition benefit; deleted cba.roi_pct; replaced the verdict with the unrecognised string "LOOKS GOOD TO ME".

Result: three distinct checks caught three distinct tampers, and generation refused.

C-BP4    FAIL benefits without a measurable metric: The administration break-point is known
                before hiring, not discovered afterwards
C-BP5    FAIL cost-benefit missing: roi_pct
C-BP11   FAIL value validation needs a recognised verdict and a justification
  9 of 12 checks passed
--check exit code 1 ; generation exit code 1 (the generator refuses to render an invalid plan)

Restored: 12 of 12, exit 0.


Test 2 — C33, the twin is validated by the suite, not only by its own generator

Tampered: deleted cba.roi_pct from business_plan.json.

C33   FAIL  business plan checks failing: C-BP5; BUSINESS_PLAN_A3.html is staler than
            business_plan.json; regenerate it

Two findings from one tamper, which is the point: C33 tests the twin AND the currency of its projection. Restored: PASS.


Test 3 — C33 refuses a self-certified PROCEED

Tampered: set the verdict to PROCEED while footer.verified_by still reads PENDING — independent verification pass V5 has not yet run against this plan.

C33   FAIL  verdict 'PROCEED' declared while independent verification is still PENDING
            (GOV-C3.2 forbids self-certification)

This is the same structural protection as C31 on the requirements register, applied to the business plan: the party that built the model cannot be the party that declares it sound. Restored: PASS.


Test 4 — C34, the one-sheet claim is measured

Background. The first build of BUSINESS_PLAN_A3.html rendered to 2,169 px against the 1,123 px an A3 landscape sheet holds — very nearly two pages — while the file name, the generator docstring and the Standard all said one page. Nothing caught it, because nothing had rendered the page. The layout was rebuilt (masonry column flow instead of a row grid, which recovered roughly 600 px of whitespace under the shorter card in each row, plus tightened typography) and C34 was written to measure rather than assert.

Tampered: injected a card of 400 repeated words into the generated page.

C34   FAIL  content overflows the A3 sheet by 184 px (1307 wanted against the 1123 a 297 mm
            sheet holds)

C34 also fails when the content fits by fewer than 15 px, because a margin inside measurement noise is not a margin. Current state: 1,103 px of content, 20 px to spare. Restored: PASS.


Test 5 — C35, placement and debris

Not tampered — this one failed on its own, unforced, on its first run:

C35   FAIL  7 placement finding(s): DEBRIS 01_System/__pycache__;
            DEBRIS 01_System/portable/templates/business_plan/__pycache__; ...

The checker was itself creating the debris it reported: importing the generator and tidy.py to run their validators wrote bytecode caches. Fixed at the source with sys.dont_write_bytecode = True in run_checks.py rather than by widening the exclusion list — excluding the finding would have hidden the next one too. The existing caches were swept to 06_Archive/_debris/2026-09-07/ with a SHA-256 manifest. Nothing was deleted. Restored: PASS.


Standing after these tests

ALL CHECKS PASSED — 35 of 35

Three new checks, five recorded failures, five recorded restorations. The one-sheet defect in Test 4 was real and shipped in the first build; it is logged as a defect, not written out of history.


SECOND ROUND — CHECKS ADDED AFTER VERIFICATION PASSES V6 AND V7

Date: 2026-09-07. Same rule: a check is only a check once it has been proven able to fail.

Test 6 — C36, a model citing a register row that does not exist

Tampered: changed SRC-063 to SRC-999 in model_agedcare.py.

C36   FAIL  model_agedcare.py cites SRC-999, which is in no register

Restored: PASS.

Test 7 — C37, the workbook disagreeing with the model

Tampered: changed the care-management share on the workbook's AC_Inputs sheet from 10 per cent to 15 per cent and recalculated.

C37   FAIL  care management revenue: workbook B9 = 375.00, model = 250.00;
            aged care billable hours: workbook B11 = 20.61, model = 21.82;
            AGED CARE CONTRIBUTION: workbook B17 = 1079.99, model = 1002.06

One tamper, five disagreements caught. Restored: PASS.

Test 8 — C34, BOTH failure modes

C34 has failed honestly twice in its life, and the second time is the important one.

8a. Pagination. Added 900 words of filler as a tenth card.

C34   FAIL  BUSINESS_PLAN_A3.html prints to 2 A3 pages, and it is called a single-sheet document.

8b. Truncation — the failure the FIRST version of this check could not see. Independent verification V7 established that the earlier fix made the page fit by clipping it: overflow:hidden produced a one-page PDF that had silently lost a table, a whole section, the value-validation verdict block and the footer, while the check reported "20 px to spare". V7 then proved the check was incapable of failing by quadrupling the card content and still getting a PASS. So C34 now extracts the printed text and requires every section heading, the verdict block and the footer to be present on the paper.

Re-clipping the sheet to 150mm to reproduce the original defect:

C34   FAIL  the page printed on one sheet but 2 element(s) never reached the paper, which means
            it was clipped rather than fitted: Lifecycle cost-benefit; Value validation

Restored both times: exactly 1 page, 420mm x 297mm, and all 9 section headings plus the verdict block and the footer are present in the printed text — fitted, not clipped.

The lesson worth keeping. The first C34 measured a box. The requirement was about pagination. A check that measures a proxy for the thing will pass while the thing is broken, and it will do so confidently. Where a claim is about how something RENDERS, the check has to drive the renderer.

Test 9 — C41, a builder overruling a verifier

Independent verifier V7 defeated the first version of C41 in a sandbox: it wrote a verification report whose text read "EVERY REQUIREMENT FAILS. 0 of 32 PASS", transcribed all thirty-two verdicts as PASS into the project, and got C31 — 32 of 32 independently verified and C41 — 32 of 32 against the CURRENT baseline, with 1 of 41 checks failing. The check verified that the report file EXISTED. It never read it.

C41 now parses the machine-readable verdict block every verification report must end with, and compares it line by line against what the project transcribed. Its first run after the change:

C41   FAIL  the verifier recorded REQ-AC-01=PASS and the project transcribed nothing for it;
            ... (all 32) ...

It passed only once every transcribed verdict matched the verifier's own block — including the two the verifier recorded as FAIL, which the project therefore cannot mark Verified.

Test 10 — C42, a source edited without a rebuild

Not tampered — this one failed on its own, unforced, on its first run, on exactly the condition V7 had found by hand:

C42   FAIL  STALE DELIVERABLES — a source was edited and not rebuilt:
            Dashboard.html is 1663 seconds older than 01_System/build_business_plan_json.py;
            03_Registers/REQ.csv is 1676 seconds older; ... 14 artefacts ...

Every other check had passed, because each one compared an artefact against another artefact of the same vintage. Nothing compared the clock. Restored by a full rebuild: PASS.


Standing after the second round

ALL CHECKS PASSED — 42 of 42, with 30 of 32 requirements independently verified against the current baseline and 2 held at Open because verification pass V7 recorded them as FAIL and the project cannot mark its own work Verified.

Ten checks have now been proven able to fail. Two of them — C34 and C41 — were proven to be INCAPABLE of failing before that, by a verifier, and both had been passing for a while.