# V5 — NEGATIVE TEST RECORD, BUSINESS LAYER CHECKS

OWNER: 02_Work/scratch/V5_negative_tests_business_layer.md
REVIEW: whenever C33, C34, C35 or the business-plan generator changes.
Date: 2026-09-07
Rule: GOV-F3.12 and GOV-E2.4 — a check is only a check once it has been PROVEN able to fail.

A green check that has never been seen red is an assumption wearing a checkmark. Each test
below broke one specific thing, recorded the failure verbatim, and restored the file.

---

## Test 1 — the generator's own twelve checks (C-BP1..C-BP12)

**Tampered:** removed the metric from one value-proposition benefit; deleted `cba.roi_pct`;
replaced the verdict with the unrecognised string "LOOKS GOOD TO ME".

**Result:** three distinct checks caught three distinct tampers, and generation refused.

```
C-BP4    FAIL benefits without a measurable metric: The administration break-point is known
                before hiring, not discovered afterwards
C-BP5    FAIL cost-benefit missing: roi_pct
C-BP11   FAIL value validation needs a recognised verdict and a justification
  9 of 12 checks passed
--check exit code 1 ; generation exit code 1 (the generator refuses to render an invalid plan)
```

Restored: 12 of 12, exit 0.

---

## Test 2 — C33, the twin is validated by the suite, not only by its own generator

**Tampered:** deleted `cba.roi_pct` from `business_plan.json`.

```
C33   FAIL  business plan checks failing: C-BP5; BUSINESS_PLAN_A3.html is staler than
            business_plan.json; regenerate it
```

Two findings from one tamper, which is the point: C33 tests the twin AND the currency of its
projection. Restored: PASS.

---

## Test 3 — C33 refuses a self-certified PROCEED

**Tampered:** set the verdict to `PROCEED` while `footer.verified_by` still reads
`PENDING — independent verification pass V5 has not yet run against this plan`.

```
C33   FAIL  verdict 'PROCEED' declared while independent verification is still PENDING
            (GOV-C3.2 forbids self-certification)
```

This is the same structural protection as C31 on the requirements register, applied to the
business plan: the party that built the model cannot be the party that declares it sound.
Restored: PASS.

---

## Test 4 — C34, the one-sheet claim is measured

**Background.** The first build of `BUSINESS_PLAN_A3.html` rendered to **2,169 px** against the
**1,123 px** an A3 landscape sheet holds — very nearly two pages — while the file name, the
generator docstring and the Standard all said one page. Nothing caught it, because nothing had
rendered the page. The layout was rebuilt (masonry column flow instead of a row grid, which
recovered roughly 600 px of whitespace under the shorter card in each row, plus tightened
typography) and C34 was written to measure rather than assert.

**Tampered:** injected a card of 400 repeated words into the generated page.

```
C34   FAIL  content overflows the A3 sheet by 184 px (1307 wanted against the 1123 a 297 mm
            sheet holds)
```

C34 also fails when the content fits by fewer than 15 px, because a margin inside measurement
noise is not a margin. Current state: **1,103 px of content, 20 px to spare.** Restored: PASS.

---

## Test 5 — C35, placement and debris

**Not tampered — this one failed on its own, unforced, on its first run:**

```
C35   FAIL  7 placement finding(s): DEBRIS 01_System/__pycache__;
            DEBRIS 01_System/portable/templates/business_plan/__pycache__; ...
```

The checker was itself creating the debris it reported: importing the generator and `tidy.py`
to run their validators wrote bytecode caches. Fixed at the source with
`sys.dont_write_bytecode = True` in `run_checks.py` rather than by widening the exclusion list —
excluding the finding would have hidden the next one too. The existing caches were swept to
`06_Archive/_debris/2026-09-07/` with a SHA-256 manifest. Nothing was deleted. Restored: PASS.

---

## Standing after these tests

`ALL CHECKS PASSED — 35 of 35`

Three new checks, five recorded failures, five recorded restorations. The one-sheet defect in
Test 4 was real and shipped in the first build; it is logged as a defect, not written out of
history.

---

# SECOND ROUND — CHECKS ADDED AFTER VERIFICATION PASSES V6 AND V7

Date: 2026-09-07. Same rule: a check is only a check once it has been proven able to fail.

## Test 6 — C36, a model citing a register row that does not exist

**Tampered:** changed `SRC-063` to `SRC-999` in `model_agedcare.py`.

```
C36   FAIL  model_agedcare.py cites SRC-999, which is in no register
```

Restored: PASS.

## Test 7 — C37, the workbook disagreeing with the model

**Tampered:** changed the care-management share on the workbook's `AC_Inputs` sheet from 10 per
cent to 15 per cent and recalculated.

```
C37   FAIL  care management revenue: workbook B9 = 375.00, model = 250.00;
            aged care billable hours: workbook B11 = 20.61, model = 21.82;
            AGED CARE CONTRIBUTION: workbook B17 = 1079.99, model = 1002.06
```

One tamper, five disagreements caught. Restored: PASS.

## Test 8 — C34, BOTH failure modes

C34 has failed honestly twice in its life, and the second time is the important one.

**8a. Pagination.** Added 900 words of filler as a tenth card.

```
C34   FAIL  BUSINESS_PLAN_A3.html prints to 2 A3 pages, and it is called a single-sheet document.
```

**8b. Truncation — the failure the FIRST version of this check could not see.** Independent
verification V7 established that the earlier fix made the page fit by clipping it: `overflow:hidden`
produced a one-page PDF that had silently lost a table, a whole section, the value-validation
verdict block and the footer, while the check reported "20 px to spare". V7 then proved the check
was incapable of failing by quadrupling the card content and still getting a PASS. So C34 now
extracts the printed text and requires every section heading, the verdict block and the footer to
be present on the paper.

Re-clipping the sheet to 150mm to reproduce the original defect:

```
C34   FAIL  the page printed on one sheet but 2 element(s) never reached the paper, which means
            it was clipped rather than fitted: Lifecycle cost-benefit; Value validation
```

Restored both times: `exactly 1 page, 420mm x 297mm, and all 9 section headings plus the verdict
block and the footer are present in the printed text — fitted, not clipped`.

**The lesson worth keeping.** The first C34 measured a box. The requirement was about pagination.
A check that measures a proxy for the thing will pass while the thing is broken, and it will do so
confidently. Where a claim is about how something RENDERS, the check has to drive the renderer.

## Test 9 — C41, a builder overruling a verifier

Independent verifier V7 defeated the first version of C41 in a sandbox: it wrote a verification
report whose text read *"EVERY REQUIREMENT FAILS. 0 of 32 PASS"*, transcribed all thirty-two
verdicts as PASS into the project, and got `C31 — 32 of 32 independently verified` and
`C41 — 32 of 32 against the CURRENT baseline`, with 1 of 41 checks failing. The check verified that
the report file EXISTED. It never read it.

C41 now parses the machine-readable verdict block every verification report must end with, and
compares it line by line against what the project transcribed. Its first run after the change:

```
C41   FAIL  the verifier recorded REQ-AC-01=PASS and the project transcribed nothing for it;
            ... (all 32) ...
```

It passed only once every transcribed verdict matched the verifier's own block — including the two
the verifier recorded as FAIL, which the project therefore cannot mark Verified.

## Test 10 — C42, a source edited without a rebuild

**Not tampered — this one failed on its own, unforced, on its first run,** on exactly the condition
V7 had found by hand:

```
C42   FAIL  STALE DELIVERABLES — a source was edited and not rebuilt:
            Dashboard.html is 1663 seconds older than 01_System/build_business_plan_json.py;
            03_Registers/REQ.csv is 1676 seconds older; ... 14 artefacts ...
```

Every other check had passed, because each one compared an artefact against another artefact of the
same vintage. Nothing compared the clock. Restored by a full rebuild: PASS.

---

## Standing after the second round

`ALL CHECKS PASSED — 42 of 42`, with **30 of 32 requirements independently verified against the
current baseline and 2 held at Open** because verification pass V7 recorded them as FAIL and the
project cannot mark its own work Verified.

Ten checks have now been proven able to fail. Two of them — C34 and C41 — were proven to be
INCAPABLE of failing before that, by a verifier, and both had been passing for a while.
