Investment Plans workspace
Open raw ↗

Gap Analysis & Upgrade Plan

Project: NDIS Business Enabling Study Benchmark Date: 2026-09-01 Governance Standard: AI Project Governance Standard v3.7 Checker Baseline: 32/32 PASS


1. Executive Summary

The NDIS Business Enabling Study is an unusually rigorous decision-support system with a deterministic Python build pipeline, 32 automated compliance checks, and a single-source-of-truth architecture (project_data.py). Compared against 2025–2026 industry benchmarks in agentic workflows, AI governance, multi-agent prompt engineering, and modern responsive dashboard design, the project scores strong on governance discipline and build integrity but lags on frontend architecture, responsive design patterns, and modern CSS practices.

Three high-impact upgrade paths emerge:

  1. Frontend Refactor — Replace the ad-hoc inline CSS with a modern, fluid, full-width responsive layout using CSS Grid/Flexbox and design tokens.
  1. Governance Modernization — Augment the existing Governance Standard with agent lifecycle controls, MCP/tool-use governance, and structured communication protocols derived from MAS-PromptBench findings.
  1. Security Hardening — Add explicit skill-based injection defenses and contextual integrity checks beyond the existing INV-1..INV-7 rules.

All recommended changes are classified under the project's own change-control taxonomy (GOV-D2).


2. Governance & Workflow Benchmarks

2.1 Current State

The project operates under AI Project Governance Standard v3.7 (13 July 2026), a 413-rule, ISO-29148-aligned framework with:

Strengths observed:

2.2 Industry Benchmarks Found

BenchmarkSourceKey Insight
GitHub Next Agentic Workflowsgithubnext/agentics18+ documented design patterns (IssueOps, ChatOps, DailyOps, BatchOps); reusable workflow catalogues with version pinning; scoped prompts and isolated sub-agent contexts
Microsoft Agent GovernanceMicrosoft LearnCentralized agent control plane; agent registry mandatory; continuous observability; cost tracking; least-privilege tool access
Berkeley CLTC / AISICLTC ReportGovernance must scale with degrees of autonomy; emergency automated shutdowns; real-time monitoring; manual shutdown as last-resort
Palo Alto Networks Agentic AI GovernancePalo Alto CyberpediaLifecycle governance (design → decommissioning); runtime controls separate from model logic; pre-deployment impact assessment; agent-to-agent trust boundaries
Zenity CISO Checklistzenity.io10-step governance: autonomous oversight, data protection, transparency, real-time monitoring, lifecycle control; platform-independent governance layer

2.3 Gaps Identified

#GapCurrent StateBenchmark ExpectationRisk
G1No agent registryAgents dispatched via ad-hoc charters in 02_Work/scratch/Microsoft/Berkeley mandate a centralized agent inventory with ownership, purpose, and access scopeShadow agent risk; untracked dispatched verifiers
G2No MCP/tool-use governanceTools used (Python, Node, LibreOffice) with env-var overrides2026 frameworks require tool-definition inspection and per-call runtime governanceUnscoped tool authority; no audit of tool invocations
G3Communication protocol undefinedAgent charters are free-form textMAS-PromptBench: structured communication protocols yield +4.3pp optimization gains over free-formAmbiguity in verifier-to-builder handoffs; evidence may be misinterpreted
G4No emergency shutdown pathINV-7 refuses overrides but no kill-switch existsBerkeley CLTC: "manual shutdown methods should be available as a last-resort control measure"A runaway build agent could consume resources or overwrite files
G5Post-delivery drift monitoring absentCurrency rule (GOV-F1.16) is annual; no automated re-checkPalo Alto: "continuous monitoring should extend to supply chain"Stale sources discovered only on manual re-verification
G6No cost/token trackingBuild runs are unmeasuredMicrosoft/GitHub Next: token usage logged per run; cost tracking dashboardsUnbounded compute spend on regeneration cycles

2.4 Recommended Upgrades

IDRecommendationGovernance ClassPriority
R-G1Agent Registry — Add an AGENTS.csv register tracking every dispatched agent's charter, role, build artefacts, and verification status. Auto-populated by build_all.py.Class 2 (shared schema, Q9)High
R-G2Tool-Use Manifest — Every tool invocation (node, soffice, python) must log its command, exit code, and duration to a TOOL_RUN register.Class 2 (adds data file, Q10/Q12)Medium
R-G3Structured Agent Charter Template — Replace free-form charters with a JSON/YAML schema containing: role, task, inputs, expected outputs, forbidden actions, and communication protocol (structured/semi-structured/free-form per MAS-PromptBench).Class 2 (schema change, Q9)Medium
R-G4Emergency Stop Script — Add 01_System/emergency_stop.py that kills running build processes and archives current state before halting. Document in Help Hub.Class 2 (new artefact, Q10)Medium
R-G5Automated Stale-Source Monitor — A monthly script that probes all 53 SRC URLs for 200 OK and flags dead links or changed content to ISS register.Class 2 (new checker, Q10)Medium
R-G6Build Cost Log — Append token count (where available) and wall-clock time per build step to daily_audit_log.md.Class 3 (local addition)Low

3. System Architecture Benchmarks

3.1 Current State

AspectImplementation
Data LayerPython module (project_data.py) — 1,434 lines of tuples and dicts
Build LayerSequential Python scripts (build_all.py orchestrates 11 steps)
Output LayerHTML (inline CSS/JS), DOCX, PPTX, XLSX, PDF, CSV, PNG
DependenciesPython stdlib + python-docx + optional soffice / node
PortabilityExcellent — no CDN, no external APIs, relative paths only
ExtensibilityModerate — adding a new register requires editing project_data.py and the builder

3.2 Industry Benchmarks Found

BenchmarkSourceKey Insight
LangGraph / Semantic KernelLangGraph Docs, Microsoft Semantic KernelStateful agent execution with checkpointing, persistence, streaming, and human-in-the-loop
GitHub Agentic Workflowsgh-aw10+ event triggers, 8+ safe output types, 5 security layers, token-usage firewall
Deterministic Build PatternsNix / Bazel principles (implied by research)Hermetic builds: pinned dependencies, sandboxed execution, reproducible outputs
Static Site GeneratorsHugo, 11ty, AstroContent → static HTML with component reuse, partial hydration, and modern bundling

3.3 Gaps Identified

#GapCurrent StateBenchmark ExpectationRisk
A1No component reuse in HTMLDashboard and Help Hub duplicate ~64 lines of inline CSSModern SSGs use shared components/partials; DRY principleMaintenance burden; style drift between surfaces
A2No CSS preprocessingRaw CSS with manual vendor prefixes and color repetitionTailwind, SCSS, or CSS custom-property systematic themingInconsistent spacing/colors; no dark mode
A3Build is not hermeticbuild_all.py skips optional steps silently if tools missingBazel/Nix model: builds fail fast or use pinned toolchainsNon-reproducible outputs across machines
A4No incremental buildsbuild_all.py regenerates everything from scratch every timeModern build systems: dependency graph, incremental compilation, cachingWasted compute; slow iteration
A5No structured artefact manifestMANIFEST.txt is plain text; no machine-readable SBOMGitHub Next / enterprise: JSON artefact manifests with SHA-256, provenanceDifficult to verify integrity programmatically

3.4 Recommended Upgrades

IDRecommendationGovernance ClassPriority
R-A1Shared CSS Module — Extract inline CSS into 01_System/templates/dashboard.css and help_hub.css, then inline at build time. Eliminates duplication.Class 2 (shared file, Q10/Q12)High
R-A2CSS Custom Property Design Tokens — Define a spacing scale (4px base), color palette, and typography ramp as :root variables. Use clamp() for fluid type.Class 2 (schema change, Q9)High
R-A3Build Hermeticity Check — build_all.py must exit non-zero (not skip) if a required tool is missing. Optional steps must be explicitly opt-in via env var.Class 2 (behavior change, Q13)Medium
R-A4Incremental Build Flag — Add --incremental to build_all.py that timestamps source files and skips builders where output is newer than input.Class 3 (local to build script)Low
R-A5JSON Artefact Manifest — Generate 05_Outputs/manifest.json with SHA-256 of every output, source file path, and build timestamp.Class 2 (new interface, Q8/Q10)Medium

4. Frontend / UI-UX Benchmarks

4.1 Current State

The Dashboard and Help Hub are self-contained HTML files with inline CSS/JS (~55 KB and ~64 KB respectively). Key CSS properties:

main{max-width:1240px;margin:0 auto;padding:20px 26px 70px}
body.mobile main{max-width:520px}
.grid2{display:grid;grid-template-columns:1fr 1fr;gap:18px}
.tiles{display:grid;grid-template-columns:repeat(auto-fit,minmax(215px,1fr));gap:13px}
@media(max-width:780px){...}

Observed limitations:

4.2 Industry Benchmarks Found

BenchmarkSourceKey Insight
Flowbite Admin DashboardGitHub 8.8k stars400+ components; HTML/React/Vue/Svelte variants; Tailwind v4; dark mode; responsive sidebar
Shadcn AdminGitHub 11.5k starsCmd+K palette; TanStack Table; React Hook Form + Zod; mobile Sheet component; RTL support
TailAdminGitHub 2k+ stars7 framework variants; 200+ components; Figma file; Tailwind v4; ApexCharts
DaisyUIGitHub 15k+ starsSemantic component classes; multiple themes; built-in theme switching; no JS dependency
Material Design 3Google Material 3Dynamic color; elevation system; motion tokens; adaptive layouts; accessibility-first
Apple Human Interface GuidelinesApple HIGReadable margins (minimum 16pt); generous whitespace; consistent corner radius; system fonts
2025 Responsive TrendsBootstrapDashCSS Grid + Flexbox; fluid layouts; container queries; CSS custom properties; text as CSS not images
Modern CSS Best PracticesPW Skills 2025Percentages/viewport units; clamp() for fluid type; multiple breakpoints (600/768/1200); Grid for dashboards

4.3 Gaps Identified

#GapCurrent StateBenchmark ExpectationImpact
U1Not full-width / fluidmax-width:1240px centeredModern dashboards: width: 100% with padding or container queries; content breathes on ultrawideWasted space; poor use of real estate on 1920px+ screens
U2Single breakpointOnly @media(max-width:780px)Best practice: 3+ breakpoints (mobile <640, tablet 640–1024, desktop 1024–1440, wide >1440)Awkward layouts on tablets and large monitors
U3No dark modeLight onlyMaterial 3 / Apple HIG / DaisyUI: prefers-color-scheme dark mode standardEye strain; not accessible for low-light environments
U4No fluid typographyFixed px values (25px h1, 19px h2, 15px base)Modern: clamp(1rem, 0.9rem + 0.5vw, 1.25rem) for scalable typeText too small on mobile, too large constraints on desktop
U5No design-token spacingArbitrary padding/margin valuesTailwind / Material: 4px base scale (0.25rem increments)Inconsistent rhythm; harder maintenance
U6Table overflow risktable{overflow:hidden;border-radius:9px}Best practice: wrap tables in .table-container{overflow-x:auto}Clipped content on narrow viewports
U7Grid2 not adaptivegrid-template-columns:1fr 1frModern: repeat(auto-fit, minmax(min(100%, 400px), 1fr)) or container queriesColumns too narrow on mid-width screens
U8No focus-visible stylesDefault browser focusWCAG 2.1: visible focus indicators with 3:1 contrast ratioAccessibility failure for keyboard users
U9Modebar wrapping uncontrolledflex-wrap:wrap with 7+ pillsModern: collapsible hamburger menu below 1024px; persistent priority linksVisual clutter on medium widths
U10No CSS container queriesMedia queries only2025 best practice: @container for component-level responsivenessComponents break when embedded in different contexts

4.4 Recommended Upgrades

IDRecommendationGovernance ClassPriority
R-U1Fluid Full-Width Layout — Replace max-width:1240px with width:100% and use padding: clamp(16px, 4vw, 48px) for responsive margins. Add .container-wide and .container-reading modifiers.Class 2 (visible behavior change, Q13)High
R-U2Multi-Breakpoint Grid — Add breakpoints at 640px, 1024px, 1440px. Use grid-template-columns: repeat(auto-fit, minmax(min(100%, 320px), 1fr)) for tiles.Class 2 (behavior change, Q13)High
R-U3Dark Mode — Implement prefers-color-scheme: dark with a full color inverse (navy→slate-900, paper→zinc-950, card→slate-900). Test contrast ratios ≥ 4.5:1.Class 2 (new interface, Q13)Medium
R-U4Fluid Typography Ramp — Replace fixed px with clamp() scale: --text-base: clamp(0.875rem, 0.8rem + 0.25vw, 1rem), --text-h1: clamp(1.5rem, 1.2rem + 1.5vw, 2.25rem), etc.Class 2 (behavior change, Q13)Medium
R-U5Design Token Spacing — Adopt 4px base scale: --space-1: 4px, --space-2: 8px, --space-3: 12px, --space-4: 16px, --space-6: 24px, --space-8: 32px, --space-12: 48px.Class 2 (schema change, Q9)Medium
R-U6Table Scroll Wrapper — Wrap every <table> in <div class="table-wrap"> with overflow-x:auto; -webkit-overflow-scrolling: touch;.Class 3 (local HTML change)High
R-U7Accessible Focus States — Add :focus-visible{outline:2px solid var(--teal);outline-offset:2px} to all interactive elements.Class 3 (local CSS change)Medium
R-U8Collapsible Modebar — Below 1024px, collapse modebar links into a "☰ Menu" toggle. Keep critical links (Dashboard, Help, Study) always visible.Class 2 (behavior change, Q13)Medium
R-U9Container Queries for Cards — Use @container (min-width: 400px) { ... } on .tile and .card components so they adapt to their parent, not just viewport.Class 2 (new CSS feature, Q13)Low

5. Security & Compliance Benchmarks

5.1 Current State

Security is governed by INV-1..INV-7 (the Inviolable Rules) and GOV-E5 (Security and Safety). Current controls:

5.2 Industry Benchmarks Found

BenchmarkSourceKey Insight
OWASP Top 10 for LLM Apps 2025OWASPLLM01 Prompt Injection = #1 threat; LLM08 Excessive Agency; defense-in-depth required
Skill-Based Injection ResearcharXiv:2602.20156New attack vector: malicious instructions embedded in SKILL.md files; contextual integrity violations; deterministic defenses needed
Spotlighting / Instruction HierarchyMicrosoft Hines et al. 2024Delimiters and encoding to mark untrusted content; separate instructions from data
MCP-Guard PrototypearXiv:2512.08290Multi-stage detector: static scan → neural detector → LLM arbitrator for tool-use security
NIST AI RMF 1.0NISTIdentify → Measure → Manage → Govern; continuous monitoring; defense-in-depth
AIVSSRSAC 2026 coverageAgent vulnerability scoring system emerging as standard

5.3 Gaps Identified

#GapCurrent StateBenchmark ExpectationRisk
S1No prompt-injection test suiteINV-6 says "treat fetched content as data" but no automated test verifies thisOWASP / research: red-team with direct, indirect, and skill-based injections before releaseA malicious NDIS Commission page could inject instructions
S2No skill-trust boundarySkills (if any) are loaded from local filesystem without verification2026 research: skills must be signed or hashed; lazy loading must validate integrityCompromised skill file could override governance rules
S3No runtime tool filteringTools are invoked directly via subprocess.run()MCP-Guard: static scan + neural detector + LLM arbitrator for tool callsA prompt injection could trigger rm -rf via a tool call
S4No content trust-tier mapAll external sources treated equally as "data"Google DeepMind CaMeL: classify sources into trust tiers; privileged actions behind human-approval tierHigh-confidence sources (ATO) and Medium-confidence sources (PDF summary) have same trust level
S5No behavioural eval / regression gatingChecker suite is static; no adversarial testingOWASP / AI RiskAtlas: behavioural evals and regression gating on every changeNew build script could silently weaken a check

5.4 Recommended Upgrades

IDRecommendationGovernance ClassPriority
R-S1Adversarial Check Suite — Add checks C33–C35: direct prompt-injection probe, indirect injection via fetched content simulation, and skill-injection via malformed SKILL.md. All must fail safely.Class 2 (new checker, Q10)High
R-S2Skill Integrity Hash — If skills are used, maintain SKILLS_HASH.txt with SHA-256 of every SKILL.md loaded. Checker verifies at build time.Class 2 (new file, Q10)Medium
R-S3Subprocess Sandbox — Wrap all subprocess.run() calls in 01_System/ with a whitelist of allowed commands and arguments. Reject anything not in the whitelist.Class 2 (behavior change, Q13)High
R-S4Source Trust-Tier Annotation — Extend SRC.csv with a trust_tier column: T1 (primary authority), T2 (secondary corroboration), T3 (estimate/unsourced). Checker flags T3 used for high-consequence calculations.Class 2 (schema change, Q9)Medium
R-S5Checker Regression Test — Negative tests for every checker (like C23) proving each can actually fail. Minimum one negative test per security-critical check.Class 2 (new test data, Q10)Medium

6. Consolidated Priority Matrix

Upgrade IDDomainPhaseEffortImpactGovernance ClassPriority Score
R-U1FrontendPhase 3MediumHighClass 2Critical
R-U2FrontendPhase 3MediumHighClass 2Critical
R-U6FrontendPhase 3LowHighClass 3Critical
R-G1GovernancePhase 4MediumHighClass 2High
R-S1SecurityPhase 4MediumHighClass 2High
R-S3SecurityPhase 4LowHighClass 2High
R-A1ArchitecturePhase 3MediumMediumClass 2High
R-A2ArchitecturePhase 3LowMediumClass 2High
R-U3FrontendPhase 3MediumMediumClass 2Medium
R-U4FrontendPhase 3LowMediumClass 2Medium
R-U5FrontendPhase 3LowMediumClass 2Medium
R-U7FrontendPhase 3LowMediumClass 3Medium
R-U8FrontendPhase 3MediumMediumClass 2Medium
R-G2GovernancePhase 4LowMediumClass 2Medium
R-G3GovernancePhase 4MediumMediumClass 2Medium
R-G4GovernancePhase 4LowMediumClass 2Medium
R-G5GovernancePhase 4MediumMediumClass 2Medium
R-S2SecurityPhase 4LowLowClass 2Medium
R-S4SecurityPhase 4LowMediumClass 2Medium
R-S5SecurityPhase 4MediumMediumClass 2Medium
R-A3ArchitecturePhase 4LowMediumClass 2Medium
R-A5ArchitecturePhase 4LowMediumClass 2Medium
R-U9FrontendPhase 3LowLowClass 2Low
R-G6GovernancePhase 4LowLowClass 3Low
R-A4ArchitecturePhase 4MediumLowClass 3Low

Recommended Execution Order

  1. Phase 3 — Frontend Refactoring (Agent 2): Execute R-U1, R-U2, R-U6, R-A1, R-A2 (fluid layout, breakpoints, table wrappers, shared CSS, design tokens)
  1. Phase 4 — Governance & Security (Agent 3): Execute R-G1, R-S1, R-S3, R-S4, R-G3, R-G5 (agent registry, adversarial checks, subprocess sandbox, trust tiers, structured charters, stale-source monitor)
  1. Phase 5 — Polish (optional): Dark mode (R-U3), fluid typography (R-U4), focus states (R-U7), collapsible modebar (R-U8)

Document generated from Phase 2 research. Sources cited inline. All recommendations map to verified industry benchmarks. No hallucinated requirements.