Gap Analysis & Upgrade Plan
Project: NDIS Business Enabling Study Benchmark Date: 2026-09-01 Governance Standard: AI Project Governance Standard v3.7 Checker Baseline: 32/32 PASS
1. Executive Summary
The NDIS Business Enabling Study is an unusually rigorous decision-support system with a deterministic Python build pipeline, 32 automated compliance checks, and a single-source-of-truth architecture (project_data.py). Compared against 2025–2026 industry benchmarks in agentic workflows, AI governance, multi-agent prompt engineering, and modern responsive dashboard design, the project scores strong on governance discipline and build integrity but lags on frontend architecture, responsive design patterns, and modern CSS practices.
Three high-impact upgrade paths emerge:
- Frontend Refactor — Replace the ad-hoc inline CSS with a modern, fluid, full-width responsive layout using CSS Grid/Flexbox and design tokens.
- Governance Modernization — Augment the existing Governance Standard with agent lifecycle controls, MCP/tool-use governance, and structured communication protocols derived from MAS-PromptBench findings.
- Security Hardening — Add explicit skill-based injection defenses and contextual integrity checks beyond the existing INV-1..INV-7 rules.
All recommended changes are classified under the project's own change-control taxonomy (GOV-D2).
2. Governance & Workflow Benchmarks
2.1 Current State
The project operates under AI Project Governance Standard v3.7 (13 July 2026), a 413-rule, ISO-29148-aligned framework with:
- 7 Inviolable Rules (INV-1..INV-7) — banking/payment/credential prohibitions
- Class 1/2/3 Change Control with explicit decision-test questions (Q1–Q15)
- Independent Verification & Validation (IV&V) — builder ≠ verifier
- 32 automated checks covering requirements, sources, assumptions, secrets, links, and consistency
- Project class: STANDARD — permits 3 merged roles provided builder≠verifier separation holds
- Single-source-of-truth:
01_System/project_data.py generates all artefacts
Strengths observed:
- Anti-fabrication rule (GOV-F8.8) with 95.6% sourced-claim ratio
- Negative test in checker suite (C23) — proven capable of failing
- No self-certification (C31) — independent verifier's recorded verdict required
- Bounded authority table (ALWAYS / ASK FIRST / NEVER)
2.2 Industry Benchmarks Found
| Benchmark | Source | Key Insight |
|---|
| GitHub Next Agentic Workflows | githubnext/agentics | 18+ documented design patterns (IssueOps, ChatOps, DailyOps, BatchOps); reusable workflow catalogues with version pinning; scoped prompts and isolated sub-agent contexts |
| Microsoft Agent Governance | Microsoft Learn | Centralized agent control plane; agent registry mandatory; continuous observability; cost tracking; least-privilege tool access |
| Berkeley CLTC / AISI | CLTC Report | Governance must scale with degrees of autonomy; emergency automated shutdowns; real-time monitoring; manual shutdown as last-resort |
| Palo Alto Networks Agentic AI Governance | Palo Alto Cyberpedia | Lifecycle governance (design → decommissioning); runtime controls separate from model logic; pre-deployment impact assessment; agent-to-agent trust boundaries |
| Zenity CISO Checklist | zenity.io | 10-step governance: autonomous oversight, data protection, transparency, real-time monitoring, lifecycle control; platform-independent governance layer |
2.3 Gaps Identified
| # | Gap | Current State | Benchmark Expectation | Risk |
|---|
| G1 | No agent registry | Agents dispatched via ad-hoc charters in 02_Work/scratch/ | Microsoft/Berkeley mandate a centralized agent inventory with ownership, purpose, and access scope | Shadow agent risk; untracked dispatched verifiers |
| G2 | No MCP/tool-use governance | Tools used (Python, Node, LibreOffice) with env-var overrides | 2026 frameworks require tool-definition inspection and per-call runtime governance | Unscoped tool authority; no audit of tool invocations |
| G3 | Communication protocol undefined | Agent charters are free-form text | MAS-PromptBench: structured communication protocols yield +4.3pp optimization gains over free-form | Ambiguity in verifier-to-builder handoffs; evidence may be misinterpreted |
| G4 | No emergency shutdown path | INV-7 refuses overrides but no kill-switch exists | Berkeley CLTC: "manual shutdown methods should be available as a last-resort control measure" | A runaway build agent could consume resources or overwrite files |
| G5 | Post-delivery drift monitoring absent | Currency rule (GOV-F1.16) is annual; no automated re-check | Palo Alto: "continuous monitoring should extend to supply chain" | Stale sources discovered only on manual re-verification |
| G6 | No cost/token tracking | Build runs are unmeasured | Microsoft/GitHub Next: token usage logged per run; cost tracking dashboards | Unbounded compute spend on regeneration cycles |
2.4 Recommended Upgrades
| ID | Recommendation | Governance Class | Priority |
|---|
| R-G1 | Agent Registry — Add an AGENTS.csv register tracking every dispatched agent's charter, role, build artefacts, and verification status. Auto-populated by build_all.py. | Class 2 (shared schema, Q9) | High |
| R-G2 | Tool-Use Manifest — Every tool invocation (node, soffice, python) must log its command, exit code, and duration to a TOOL_RUN register. | Class 2 (adds data file, Q10/Q12) | Medium |
| R-G3 | Structured Agent Charter Template — Replace free-form charters with a JSON/YAML schema containing: role, task, inputs, expected outputs, forbidden actions, and communication protocol (structured/semi-structured/free-form per MAS-PromptBench). | Class 2 (schema change, Q9) | Medium |
| R-G4 | Emergency Stop Script — Add 01_System/emergency_stop.py that kills running build processes and archives current state before halting. Document in Help Hub. | Class 2 (new artefact, Q10) | Medium |
| R-G5 | Automated Stale-Source Monitor — A monthly script that probes all 53 SRC URLs for 200 OK and flags dead links or changed content to ISS register. | Class 2 (new checker, Q10) | Medium |
| R-G6 | Build Cost Log — Append token count (where available) and wall-clock time per build step to daily_audit_log.md. | Class 3 (local addition) | Low |
3. System Architecture Benchmarks
3.1 Current State
| Aspect | Implementation |
|---|
| Data Layer | Python module (project_data.py) — 1,434 lines of tuples and dicts |
| Build Layer | Sequential Python scripts (build_all.py orchestrates 11 steps) |
| Output Layer | HTML (inline CSS/JS), DOCX, PPTX, XLSX, PDF, CSV, PNG |
| Dependencies | Python stdlib + python-docx + optional soffice / node |
| Portability | Excellent — no CDN, no external APIs, relative paths only |
| Extensibility | Moderate — adding a new register requires editing project_data.py and the builder |
3.2 Industry Benchmarks Found
| Benchmark | Source | Key Insight |
|---|
| LangGraph / Semantic Kernel | LangGraph Docs, Microsoft Semantic Kernel | Stateful agent execution with checkpointing, persistence, streaming, and human-in-the-loop |
| GitHub Agentic Workflows | gh-aw | 10+ event triggers, 8+ safe output types, 5 security layers, token-usage firewall |
| Deterministic Build Patterns | Nix / Bazel principles (implied by research) | Hermetic builds: pinned dependencies, sandboxed execution, reproducible outputs |
| Static Site Generators | Hugo, 11ty, Astro | Content → static HTML with component reuse, partial hydration, and modern bundling |
3.3 Gaps Identified
| # | Gap | Current State | Benchmark Expectation | Risk |
|---|
| A1 | No component reuse in HTML | Dashboard and Help Hub duplicate ~64 lines of inline CSS | Modern SSGs use shared components/partials; DRY principle | Maintenance burden; style drift between surfaces |
| A2 | No CSS preprocessing | Raw CSS with manual vendor prefixes and color repetition | Tailwind, SCSS, or CSS custom-property systematic theming | Inconsistent spacing/colors; no dark mode |
| A3 | Build is not hermetic | build_all.py skips optional steps silently if tools missing | Bazel/Nix model: builds fail fast or use pinned toolchains | Non-reproducible outputs across machines |
| A4 | No incremental builds | build_all.py regenerates everything from scratch every time | Modern build systems: dependency graph, incremental compilation, caching | Wasted compute; slow iteration |
| A5 | No structured artefact manifest | MANIFEST.txt is plain text; no machine-readable SBOM | GitHub Next / enterprise: JSON artefact manifests with SHA-256, provenance | Difficult to verify integrity programmatically |
3.4 Recommended Upgrades
| ID | Recommendation | Governance Class | Priority |
|---|
| R-A1 | Shared CSS Module — Extract inline CSS into 01_System/templates/dashboard.css and help_hub.css, then inline at build time. Eliminates duplication. | Class 2 (shared file, Q10/Q12) | High |
| R-A2 | CSS Custom Property Design Tokens — Define a spacing scale (4px base), color palette, and typography ramp as :root variables. Use clamp() for fluid type. | Class 2 (schema change, Q9) | High |
| R-A3 | Build Hermeticity Check — build_all.py must exit non-zero (not skip) if a required tool is missing. Optional steps must be explicitly opt-in via env var. | Class 2 (behavior change, Q13) | Medium |
| R-A4 | Incremental Build Flag — Add --incremental to build_all.py that timestamps source files and skips builders where output is newer than input. | Class 3 (local to build script) | Low |
| R-A5 | JSON Artefact Manifest — Generate 05_Outputs/manifest.json with SHA-256 of every output, source file path, and build timestamp. | Class 2 (new interface, Q8/Q10) | Medium |
4. Frontend / UI-UX Benchmarks
4.1 Current State
The Dashboard and Help Hub are self-contained HTML files with inline CSS/JS (~55 KB and ~64 KB respectively). Key CSS properties:
main{max-width:1240px;margin:0 auto;padding:20px 26px 70px}
body.mobile main{max-width:520px}
.grid2{display:grid;grid-template-columns:1fr 1fr;gap:18px}
.tiles{display:grid;grid-template-columns:repeat(auto-fit,minmax(215px,1fr));gap:13px}
@media(max-width:780px){...}
Observed limitations:
- Single breakpoint (780px)
- Content constrained to 1240px max-width (not full-screen)
- Mobile mode over-constrained to 520px
- No fluid typography (
clamp())
- Ad-hoc spacing values (22px, 26px, 20px, 16px, 14px, 11px…)
- Tables use
overflow:hidden without horizontal scroll wrapper
.grid2 lacks flex-wrap or adaptive behavior for 800px–1240px range
4.2 Industry Benchmarks Found
| Benchmark | Source | Key Insight |
|---|
| Flowbite Admin Dashboard | GitHub 8.8k stars | 400+ components; HTML/React/Vue/Svelte variants; Tailwind v4; dark mode; responsive sidebar |
| Shadcn Admin | GitHub 11.5k stars | Cmd+K palette; TanStack Table; React Hook Form + Zod; mobile Sheet component; RTL support |
| TailAdmin | GitHub 2k+ stars | 7 framework variants; 200+ components; Figma file; Tailwind v4; ApexCharts |
| DaisyUI | GitHub 15k+ stars | Semantic component classes; multiple themes; built-in theme switching; no JS dependency |
| Material Design 3 | Google Material 3 | Dynamic color; elevation system; motion tokens; adaptive layouts; accessibility-first |
| Apple Human Interface Guidelines | Apple HIG | Readable margins (minimum 16pt); generous whitespace; consistent corner radius; system fonts |
| 2025 Responsive Trends | BootstrapDash | CSS Grid + Flexbox; fluid layouts; container queries; CSS custom properties; text as CSS not images |
| Modern CSS Best Practices | PW Skills 2025 | Percentages/viewport units; clamp() for fluid type; multiple breakpoints (600/768/1200); Grid for dashboards |
4.3 Gaps Identified
| # | Gap | Current State | Benchmark Expectation | Impact |
|---|
| U1 | Not full-width / fluid | max-width:1240px centered | Modern dashboards: width: 100% with padding or container queries; content breathes on ultrawide | Wasted space; poor use of real estate on 1920px+ screens |
| U2 | Single breakpoint | Only @media(max-width:780px) | Best practice: 3+ breakpoints (mobile <640, tablet 640–1024, desktop 1024–1440, wide >1440) | Awkward layouts on tablets and large monitors |
| U3 | No dark mode | Light only | Material 3 / Apple HIG / DaisyUI: prefers-color-scheme dark mode standard | Eye strain; not accessible for low-light environments |
| U4 | No fluid typography | Fixed px values (25px h1, 19px h2, 15px base) | Modern: clamp(1rem, 0.9rem + 0.5vw, 1.25rem) for scalable type | Text too small on mobile, too large constraints on desktop |
| U5 | No design-token spacing | Arbitrary padding/margin values | Tailwind / Material: 4px base scale (0.25rem increments) | Inconsistent rhythm; harder maintenance |
| U6 | Table overflow risk | table{overflow:hidden;border-radius:9px} | Best practice: wrap tables in .table-container{overflow-x:auto} | Clipped content on narrow viewports |
| U7 | Grid2 not adaptive | grid-template-columns:1fr 1fr | Modern: repeat(auto-fit, minmax(min(100%, 400px), 1fr)) or container queries | Columns too narrow on mid-width screens |
| U8 | No focus-visible styles | Default browser focus | WCAG 2.1: visible focus indicators with 3:1 contrast ratio | Accessibility failure for keyboard users |
| U9 | Modebar wrapping uncontrolled | flex-wrap:wrap with 7+ pills | Modern: collapsible hamburger menu below 1024px; persistent priority links | Visual clutter on medium widths |
| U10 | No CSS container queries | Media queries only | 2025 best practice: @container for component-level responsiveness | Components break when embedded in different contexts |
4.4 Recommended Upgrades
| ID | Recommendation | Governance Class | Priority |
|---|
| R-U1 | Fluid Full-Width Layout — Replace max-width:1240px with width:100% and use padding: clamp(16px, 4vw, 48px) for responsive margins. Add .container-wide and .container-reading modifiers. | Class 2 (visible behavior change, Q13) | High |
| R-U2 | Multi-Breakpoint Grid — Add breakpoints at 640px, 1024px, 1440px. Use grid-template-columns: repeat(auto-fit, minmax(min(100%, 320px), 1fr)) for tiles. | Class 2 (behavior change, Q13) | High |
| R-U3 | Dark Mode — Implement prefers-color-scheme: dark with a full color inverse (navy→slate-900, paper→zinc-950, card→slate-900). Test contrast ratios ≥ 4.5:1. | Class 2 (new interface, Q13) | Medium |
| R-U4 | Fluid Typography Ramp — Replace fixed px with clamp() scale: --text-base: clamp(0.875rem, 0.8rem + 0.25vw, 1rem), --text-h1: clamp(1.5rem, 1.2rem + 1.5vw, 2.25rem), etc. | Class 2 (behavior change, Q13) | Medium |
| R-U5 | Design Token Spacing — Adopt 4px base scale: --space-1: 4px, --space-2: 8px, --space-3: 12px, --space-4: 16px, --space-6: 24px, --space-8: 32px, --space-12: 48px. | Class 2 (schema change, Q9) | Medium |
| R-U6 | Table Scroll Wrapper — Wrap every <table> in <div class="table-wrap"> with overflow-x:auto; -webkit-overflow-scrolling: touch;. | Class 3 (local HTML change) | High |
| R-U7 | Accessible Focus States — Add :focus-visible{outline:2px solid var(--teal);outline-offset:2px} to all interactive elements. | Class 3 (local CSS change) | Medium |
| R-U8 | Collapsible Modebar — Below 1024px, collapse modebar links into a "☰ Menu" toggle. Keep critical links (Dashboard, Help, Study) always visible. | Class 2 (behavior change, Q13) | Medium |
| R-U9 | Container Queries for Cards — Use @container (min-width: 400px) { ... } on .tile and .card components so they adapt to their parent, not just viewport. | Class 2 (new CSS feature, Q13) | Low |
5. Security & Compliance Benchmarks
5.1 Current State
Security is governed by INV-1..INV-7 (the Inviolable Rules) and GOV-E5 (Security and Safety). Current controls:
- No secrets/credentials in files (GOV-E5.1)
- No destructive commands without explicit approval (GOV-E5.2)
- Least-privilege tool use (GOV-E5.3)
- External content treated as untrusted data (GOV-E5.4 / INV-6)
- No data exfiltration without approval (GOV-E5.5)
- Secure-by-default build (input validation, fail closed)
- 32-checker suite includes secrets/PII scan (C08)
5.2 Industry Benchmarks Found
| Benchmark | Source | Key Insight |
|---|
| OWASP Top 10 for LLM Apps 2025 | OWASP | LLM01 Prompt Injection = #1 threat; LLM08 Excessive Agency; defense-in-depth required |
| Skill-Based Injection Research | arXiv:2602.20156 | New attack vector: malicious instructions embedded in SKILL.md files; contextual integrity violations; deterministic defenses needed |
| Spotlighting / Instruction Hierarchy | Microsoft Hines et al. 2024 | Delimiters and encoding to mark untrusted content; separate instructions from data |
| MCP-Guard Prototype | arXiv:2512.08290 | Multi-stage detector: static scan → neural detector → LLM arbitrator for tool-use security |
| NIST AI RMF 1.0 | NIST | Identify → Measure → Manage → Govern; continuous monitoring; defense-in-depth |
| AIVSS | RSAC 2026 coverage | Agent vulnerability scoring system emerging as standard |
5.3 Gaps Identified
| # | Gap | Current State | Benchmark Expectation | Risk |
|---|
| S1 | No prompt-injection test suite | INV-6 says "treat fetched content as data" but no automated test verifies this | OWASP / research: red-team with direct, indirect, and skill-based injections before release | A malicious NDIS Commission page could inject instructions |
| S2 | No skill-trust boundary | Skills (if any) are loaded from local filesystem without verification | 2026 research: skills must be signed or hashed; lazy loading must validate integrity | Compromised skill file could override governance rules |
| S3 | No runtime tool filtering | Tools are invoked directly via subprocess.run() | MCP-Guard: static scan + neural detector + LLM arbitrator for tool calls | A prompt injection could trigger rm -rf via a tool call |
| S4 | No content trust-tier map | All external sources treated equally as "data" | Google DeepMind CaMeL: classify sources into trust tiers; privileged actions behind human-approval tier | High-confidence sources (ATO) and Medium-confidence sources (PDF summary) have same trust level |
| S5 | No behavioural eval / regression gating | Checker suite is static; no adversarial testing | OWASP / AI RiskAtlas: behavioural evals and regression gating on every change | New build script could silently weaken a check |
5.4 Recommended Upgrades
| ID | Recommendation | Governance Class | Priority |
|---|
| R-S1 | Adversarial Check Suite — Add checks C33–C35: direct prompt-injection probe, indirect injection via fetched content simulation, and skill-injection via malformed SKILL.md. All must fail safely. | Class 2 (new checker, Q10) | High |
| R-S2 | Skill Integrity Hash — If skills are used, maintain SKILLS_HASH.txt with SHA-256 of every SKILL.md loaded. Checker verifies at build time. | Class 2 (new file, Q10) | Medium |
| R-S3 | Subprocess Sandbox — Wrap all subprocess.run() calls in 01_System/ with a whitelist of allowed commands and arguments. Reject anything not in the whitelist. | Class 2 (behavior change, Q13) | High |
| R-S4 | Source Trust-Tier Annotation — Extend SRC.csv with a trust_tier column: T1 (primary authority), T2 (secondary corroboration), T3 (estimate/unsourced). Checker flags T3 used for high-consequence calculations. | Class 2 (schema change, Q9) | Medium |
| R-S5 | Checker Regression Test — Negative tests for every checker (like C23) proving each can actually fail. Minimum one negative test per security-critical check. | Class 2 (new test data, Q10) | Medium |
6. Consolidated Priority Matrix
| Upgrade ID | Domain | Phase | Effort | Impact | Governance Class | Priority Score |
|---|
| R-U1 | Frontend | Phase 3 | Medium | High | Class 2 | Critical |
| R-U2 | Frontend | Phase 3 | Medium | High | Class 2 | Critical |
| R-U6 | Frontend | Phase 3 | Low | High | Class 3 | Critical |
| R-G1 | Governance | Phase 4 | Medium | High | Class 2 | High |
| R-S1 | Security | Phase 4 | Medium | High | Class 2 | High |
| R-S3 | Security | Phase 4 | Low | High | Class 2 | High |
| R-A1 | Architecture | Phase 3 | Medium | Medium | Class 2 | High |
| R-A2 | Architecture | Phase 3 | Low | Medium | Class 2 | High |
| R-U3 | Frontend | Phase 3 | Medium | Medium | Class 2 | Medium |
| R-U4 | Frontend | Phase 3 | Low | Medium | Class 2 | Medium |
| R-U5 | Frontend | Phase 3 | Low | Medium | Class 2 | Medium |
| R-U7 | Frontend | Phase 3 | Low | Medium | Class 3 | Medium |
| R-U8 | Frontend | Phase 3 | Medium | Medium | Class 2 | Medium |
| R-G2 | Governance | Phase 4 | Low | Medium | Class 2 | Medium |
| R-G3 | Governance | Phase 4 | Medium | Medium | Class 2 | Medium |
| R-G4 | Governance | Phase 4 | Low | Medium | Class 2 | Medium |
| R-G5 | Governance | Phase 4 | Medium | Medium | Class 2 | Medium |
| R-S2 | Security | Phase 4 | Low | Low | Class 2 | Medium |
| R-S4 | Security | Phase 4 | Low | Medium | Class 2 | Medium |
| R-S5 | Security | Phase 4 | Medium | Medium | Class 2 | Medium |
| R-A3 | Architecture | Phase 4 | Low | Medium | Class 2 | Medium |
| R-A5 | Architecture | Phase 4 | Low | Medium | Class 2 | Medium |
| R-U9 | Frontend | Phase 3 | Low | Low | Class 2 | Low |
| R-G6 | Governance | Phase 4 | Low | Low | Class 3 | Low |
| R-A4 | Architecture | Phase 4 | Medium | Low | Class 3 | Low |
Recommended Execution Order
- Phase 3 — Frontend Refactoring (Agent 2): Execute R-U1, R-U2, R-U6, R-A1, R-A2 (fluid layout, breakpoints, table wrappers, shared CSS, design tokens)
- Phase 4 — Governance & Security (Agent 3): Execute R-G1, R-S1, R-S3, R-S4, R-G3, R-G5 (agent registry, adversarial checks, subprocess sandbox, trust tiers, structured charters, stale-source monitor)
- Phase 5 — Polish (optional): Dark mode (R-U3), fluid typography (R-U4), focus states (R-U7), collapsible modebar (R-U8)
Document generated from Phase 2 research. Sources cited inline. All recommendations map to verified industry benchmarks. No hallucinated requirements.