⚠️ RETIRED / SUPERSEDED — Simon's decision, 19 Jul 2026. This narrative Session Log ended at Session 19 (22–23 June 2026). From ~27 June onward the per-task record moved to dated cc-report wiki pages (xzopia/cc-reports/<task>-<date>), now the permanent record format — there will be no further narrative session entries here, and July→today was deliberately NOT backfilled. Full rationale + real timeline: docs/session-protocol.md → the "Documentation record model (current + how it evolved)" section (commit 28c7ca4). The reference tables lower in this file (Server Architecture / Services per Server / DNS) remain live and current; only this narrative log is retired.
Entries below (through Session 19) are preserved as historical record — do not extend them.
¶ 22–23 June 2026 — Session 19 (TRMM Phase 2 pilot scaffold built; TV inventory reconcile; TV enrolment method locked)
TacticalRMM Phase 2 — pilot scaffold BUILT (23 Jun):
Phase 0 investigation: B2 onboarding trigger confirmed as task_type: "onboarding" — native TRMM type. Fires once per agent on first check-in when last_run=None and not RUNNING/COMPLETED. Agent minimum v2.6.0 (current bundled: v2.10.0 ✓). Verified from source in /rmm/api/tacticalrmm/tacticalrmm/constants.py + core/tasks.py + autotasks/models.py. /agents/onboarding/ endpoint returns 404 in 1.4.1 — does not exist. task_type: onboarding is the correct, native mechanism.
TRMM policy homework applied via API (commit 894f64b): global defaults set (policy 5 WS, policy 7 Srv); NMW client → policy 6; Task 4 bitmask bug fixed (64→1, Sat→Sun); 5 patch-compliance checks on policies 5-9; reboot/shutdown tasks 11-16 with Conditional Reboot (id=138) + Shutdown (id=139) scripts.
Installer PS1 generated: agent v2.10.0, client=64, site=66, 720h token (~23 Jul 2026 expiry). Staged in /tmp on S4. NOT committed — contains temporary install credential.
Lesson #96 added: PUT /clients/sites/{id}/ requires {"site": {...}} wrapper — same rule as lesson #92 for client PUT. Flat payload produces Django 500.
Pilot doc created: docs/migration/trmm-pilot-scaffold.md — Phase 0 B2 mechanism, Phase 1 scaffold state, Test 1 checklist, Phase 2 automation steps.
TRMM on 1.4.1 — do NOT apply 1.5.1 until post-pilot (snapshot S1 + check agent-version floor first).
Next (Simon's actions — Test 1 gate): Pick 1-2 internal low-risk test machines → paste installer PS1 as a NinjaOne script body → run it on those machines → verify 7-item Test 1 checklist (see docs/migration/trmm-pilot-scaffold.md).
TeamViewer inventory reconcile — completed 22 Jun (commit 7ea790a):
374 classic TV devices / 55 groups / 88 managed devices cross-referenced against TRMM 63 clients + NinjaOne 309 baseline.
WeBottle + SC-Eng have 0 classic TV devices (no TV installed). CGC and CRJ unidentified — CGC = Clandeboye Golf Club (13 devices online), CRJ = unknown.
Single PS1 deploy script: silent uninstall of existing TV + install branded custom Host + assign to _New Devices + capture ClientID from registry. One unavoidable "Allow" click (Premium licence ceiling — not a misconfiguration). Corporate licence removes the click (MSI + silent --assignment) but not needed for current scale.
Guard signal LOCKED: IsTechnician=true custom field — structural exclusion (task only on client-facing policies 5/7/8/9) + in-script guard (belt-and-braces). Internal Xzopia machines keep the FULL TV client; guard ensures they are never hit by the deploy script.
Click-to-connect TESTED: TeamViewer.exe -i <TeamViewer-ID> jumps straight to remote window, no password prompt. GAP: local TeamViewer-ID capture not yet established (registry gives ClientID, a different value from the 9-10 digit partner ID). TeamViewer-ID + click-to-connect gated on closing this gap.
⚠️ ROTATE the exposed TV API token (16140329-sV0EH0obou2zo1rvwHvE) — visible in a screenshot during this session.
Build prompt ready: cc-prompt-teamviewer-deploy.md. Run AFTER pilot completes.
Commit history:
db734bf — docs: add TRMM Phase 2 homework-apply plan
894f64b — fix(trmm): apply policy homework A-E via API
Status: TRMM Phase 2 pilot scaffold ✅ BUILT | Test 1 ⏸ AWAITING Simon (pick machines + run installer) | TV reconcile ✅ | TV enrolment method locked ✅ | TV deploy build prompt READY (post-pilot)
S6 was reimaged (corresponds to 06 Jun CIPP abandonment). Confirmed 18 Jun via SSH investigation.
Current state: blank Ubuntu 24.04.4 LTS, kernel 6.8.0-124, KVM/QEMU. Default hostname "ubuntu", no xzopia user, root SSH on port 22, no hardening.
lsblk: plain ext4 on vda — NO LUKS. No CIPP, no Docker, no nginx remnants.
The "S6 cleanup before Wazuh" gate (CIPP Docker teardown, AllowTcpForwarding revert) is MOOT — done by reimage.
Next move: standard 13-step server-prep checklist (xzopia user, hostname xzopiaserver6, SSH → 2222, ssh.socket → ssh.service, UFW), then Wazuh (still gated on sizing, id=25).
PM 18 Jun update: S6 went from refusing port 22 (AM) to timing out (PM) after the power outage that hit the WSL desktop. Timeout vs refused suggests the box may not have come back up cleanly. Confirm S6 is powered on via the Fasthosts panel before prep. Stage 0 is BLOCKED until S6 is reachable.
OS: CentOS 7 (EOL 30 Jun 2024), OpenVZ, uptime 212 days. Hosting: Virtualmin stack with one domain (downwholesale — stale WordPress, no traffic since April 2022). DNS (NS/A/MX) all point elsewhere (AWS/M365/external) — nothing live on this VPS.
Security finding: masscan present at ~/.cache/.masscan + three tiny random-named hidden dotfiles dated 8–9 Apr 2022. Consistent with botnet/scanner infection ~4 days after box was set up.
Action: Astutium panel snapshot → sign-off → shutdown -h now (reversible halt 18 Jun). ~1–2 week grace window then cancel VPS to stop billing. See FUTURE_DEVELOPMENT.md decommission record.
Root cause confirmed by VoiceHost: the 408 on re-register was a per-node IP block triggered by accumulated failed registration attempts, NOT a nonce/digest-auth incompatibility.
Fix: VoiceHost clear the block for 77.68.112.222 on node 185.91.41.30. Then enable SRTP and re-test with a live sofia SIP trace.
FusionPBX work PARKED until VoiceHost clears the block. 3CX on Basic (no renewal deadline).
Phase 1 DONE: 62 clients + 64 sites created in TacticalRMM. 61 created this run; 1 skipped (Ards Self Storage, prove-out). Zero errors, zero zero-site clients, zero diff. Ziggi Cig has 3 sites (Main Office, Ballymena, Newtownards). Report: docs/migration/phase1-clients-sites-report.md.
API quirks found and documented (LESSONS_LEARNED.md #87): no /api/v3 prefix; sites nested in client; prove-out-first methodology confirmed.
Phase 2 GATED on paid NinjaOne overlap. NinjaOne renewal end June — decide overlap window before renewal date. 309 agents to deploy via NinjaOne script.
¶ 15 June 2026 — Session 15 (Engineer Hub + training docs + CIPP hardening runbook)
Documentation batch committed (14 new files in docs/):
docs/engineer-hub.md — orientation hub for new engineers; house rules (incl. ITFlow/Twenty data-ownership boundary as rule 7); stack at a glance; onboarding order.
docs/cipp-hardening-runbook.md — CIPP hardening runbook for Stephen: GDAP least-privilege migration (67 GA relationships → 15 CIPP-recommended groups), SAM 6-secret cleanup, PS 7.4 EOL upgrade, RBAC review, service-account hardening, Azure resource hardening. PREP vs EXECUTE format; gated on Simon sign-off.
docs/engineer-training-{fusionpbx,wikijs,bookstack,vaultwarden,twenty-crm,mcp,backups,security-baseline,wazuh}.md (×9) — per-stack training pages with official docs, video walkthroughs, key concepts, Xzopia-specific gotchas, and a sign-off checklist. Wazuh is a stub pending S6 deployment.
docs/engineer-training-tacticalrmm-itflow.md — repo mirror of wiki id=18 (previously wiki-only, now repo-managed). ITFlow/Twenty data-ownership boundary added to ITFlow section.
docs/data-ownership.md — standalone ITFlow vs Twenty CRM boundary reference: what each system owns, lifecycle handoff at Won opportunity, overlap rule, one-line test.
¶ 12 June 2026 — Session 13 (Backup Timer Fix + Record Correction)
Backup timers — both were dead, now fixed:
⚠️ The Session 12 claim "backup system fully complete — no further action needed" was inaccurate. Both xzopia-backup.timer units were inactive (dead) — systemctl enable had been run but not enable --now, and neither box was rebooted. No scheduled backup had ever run on either server.
S5 timer: caught and fixed 2026-06-11 — active (waiting) ✅
S2 timer: caught and fixed 2026-06-12 after 3 days with no Vaultwarden/ITFlow/BookStack backup — active (waiting) ✅
S2 first unattended run: due 2026-06-13 ~01:01 — confirm tomorrow morning
Dead-man's switch tightened:
Both healthchecks.io checks: Period 24h → Grace 2h (was 1h — 25h window masked 3-day-dead timer via manual debug pings)
Alarm arrives by ~03:01 on any day a run is missed
New lessons: LESSONS_LEARNED.md #76 (timer enable vs enable --now, unattended-run gate) + #77 (grace window sizing).
TLS route: phones and trunk both SIP-TLS. Port 5061 open (dynamic softphones). Port 5080 geofenced to 3 trunk IPs. Blanket UK-wide drop on 5060:5091 NOT applied.
Trunk TLS at cutover: proxy = secure.st.sipconvergence.co.uk, transport = tls, SRTP on (chargeable — enable at VoiceHost first). GoDaddy SF Root G2 CA needed in FreeSWITCH.
Geofence resolved: TRUNK_IPS = 37.157.54.206, 185.91.41.29, 185.91.41.30 (secure.st.sipconvergence.co.uk). NOT empty, NOT just .30.
Test restore PASSED: ITFlow dump gunzip -t OK + all 5 core tables present; Vaultwarden PRAGMA integrity_check → ok (39 ciphers, 1 user); FusionPBX pg_dump -- PostgreSQL database dump complete marker confirmed. All temp targets shredded after verification.
Timers armed: Server 2 01:00, Server 5 01:30. Systemd persistent + randomised 120s delay.
Alerting: Mailgun EU proven end-to-end. Transport: smtp.eu.mailgun.org:465 implicit TLS; sending domain mg.xzopiasecure.com (SPF + DKIM verified); M365 transport rule sets SCL -1 for mg.xzopiasecure.com (required — Defender quarantined first sends). Credentials in /etc/msmtprc (chmod 600, root:root, not in repo).
Two script bugs found and fixed during live run: BookStack .env quoted password (Docker Compose strips quotes; cut did not); find \( ... ) shred syntax error (unescaped ) — fixed to \)).
Docs updated: docs/backup-solution.md (Mailgun section, status complete); docs/backup-restore-runbook.md (Scenario 2 step ordering fix — volumes before DB dumps to prevent Vaultwarden SQLite overwrite); docs/session-handoff.md.
Open follow-ups:
Confirm first unattended timer run tonight (check tomorrow: journalctl -u xzopia-backup, restic snapshots → expect count 2).
Dead-man's-switch healthcheck — on roadmap as backstop for total-failure scenarios.
Status: Backup system ✅ Complete | CIPP v10.5.0 update pending user action (sync forks + Azure Function App restart) | CIPP SAM wizard re-test pending
Custom SWA auth emulator created: CIPP is designed exclusively for Azure Static Web Apps which provides managed /.auth/* endpoints. Self-hosted has no equivalent. Built a Python3 SWA auth emulator (/opt/cipp/auth-service.py, systemd cipp-auth.service, port 7072) implementing /.auth/me, /.auth/login/aad, /authredirect (PKCE OAuth2 with server-side PKCE store), /.auth/logout, and /validate (nginx auth_request for x-ms-client-principal injection).
Authentication fully solved: Microsoft OAuth login worked end-to-end. User confirmed: logged in, CIPP dashboard visible.
Two fundamental blockers prevented production use:
Azurite API version lag: CIPP's Durable Functions SDK requested Storage API 2026-02-06; Azurite 3.35.0 rejected it with InvalidHeaderValue. cippapi containers failed to create their task hub — Azurite is not a complete Azure Storage stand-in for Durable Functions.
CIPP credential flow coupled to SAM wizard: Manual injection of ApplicationSecret + RefreshToken into .env hit AADSTS7000215: Invalid client secret and RefreshToken truncation (1580 chars vs ~1842 expected). No tenants could be loaded.
Decision: Abandon self-hosted Docker route. Deploy CIPP via official Azure route (SWA + Functions + Key Vault + Storage). See docs/cipp-azure-deployment.md.
Server 6 repurposed at this point: Hardened box (Lynis 81/100, LUKS2, Nginx 1.31.1, UFW) → Wazuh SIEM candidate. (Historical — see 18 Jun correction below.)
[18 Jun correction] S6 was subsequently REIMAGED (confirmed 18 Jun). The hardening, LUKS2, nginx, and all CIPP remnants described above no longer exist. Current state: blank Ubuntu 24.04.4 LTS, no LUKS, no hardening, root SSH on port 22. The "cleanup before Wazuh" gate is MOOT — done by reimage. Next: 13-step server-prep checklist, then Wazuh (gated on sizing).
/api/ proxy to 127.0.0.1:7071 preserved (cippapi via cipp-nginx container).
Verification: curl -sk https://cipp.xzopiasecure.com returns CIPP HTML with <title>Loading</title> and "Logging into CIPP" — not the Azure Functions default page.
Status: CIPP frontend LIVE ✅ | Entra ID app registration — still pending user action.
¶ 04 June 2026 — Session 8 (CIPP Entra ID — Planning)
Session summary: Reviewed session handoff and planned CIPP Entra ID app registration.
No infrastructure changes made this session — planning and documentation only.
Reviewed full Prompt 3 from docs/cipp-deployment-plan.md (Entra ID + CE compliance).
Documented step-by-step Entra ID app registration process in session-handoff.md:
app registration, client secret, API permissions (12 Graph permissions), Conditional
Access policy (MFA required — CE auto-fail if missing), service account creation,
Vaultwarden storage.
Server 6 infrastructure remains fully operational — no changes made.
Next action: user to complete Entra ID portal steps, then return for CIPP wizard + CE doc.
Status: CIPP Entra ID config — NOT STARTED. Pending user action in Microsoft portals.
API credentials generated and stored in Vaultwarden (priority #1 and #2 complete):
NinjaOne API — Dev: dedicated client_id + client_secret (eu.ninjarmm.com). In Vaultwarden "API Credentials".
Xero API — Live: Custom Connection client_id + client_secret + tenant_id. In Vaultwarden.
Giacom API — Live (billing): Azure APIM subscription key. In Vaultwarden.
Giacom API — Live (partner): Azure APIM subscription key. In Vaultwarden.
Remaining credentials — all generated and deployed in Session 5 ✅
Outstanding:
Giacom APIM endpoint paths — RESOLVED (02 Jun): paths corrected to official Billing Service API v1 base (cloudmarket-services.azure-api.net/Billing/v1) — /AccountTotals, / (BillingList, query-param driven), /SubscriptionsManagementReport — live HTTP 200 confirmed via MCP JSON-RPC. GIACOM_PATH_* vars written to server .env.
Xero and Giacom dev credentials are reusing Live keys — isolate when dedicated dev apps are created.
Own config/DB backed up via fleet Restic→B2 (repo server3/, daily 03:00, DMS armed, Lynis 83)
Client backup data → SEPARATE B2 bucket via Comet Storage Vault (not the internal xzopia-backup bucket)
REUSE-IN-PLACE of the old decommissioned MCP box (resized, not reimaged; reports id=120-125). Old MCP stack (ninjaone/xero/giacom + claudemcp) removed 14 Jul.
Server 4 (LIVE PRODUCTION MCP — 12 connectors; snapshot/backup before any change):
Backup: restic → B2 daily 01:30. Timer currently INACTIVE — fix before Phase 1 config starts.
VoIP security (voip-security.sh) not yet run. Run with TRUNK_IPS set (see build doc) before cutover.
Build plan: docs/server5-fusionpbx-build.md
Server 6 — Wazuh SIEM candidate (reimaged + blank as of 06 Jun 2026, confirmed 18 Jun):
Current state: blank Ubuntu 24.04.4 LTS (kernel 6.8.0-124, KVM/QEMU). Default hostname "ubuntu", no xzopia user, root SSH on port 22 — none of the old hardening, LUKS, nginx, or CIPP remnants exist.
The "S6 cleanup before Wazuh" gate (CIPP Docker teardown, AllowTcpForwarding revert) is MOOT — done by reimage. No cleanup needed.
⚠️ PM 18 Jun: S6 went from refusing port 22 (AM) to timing out (PM) after power outage. Confirm S6 is powered on via the Fasthosts panel before attempting prep.
Next: standard 13-step server-prep checklist (xzopia user, hostname xzopiaserver6, SSH → 2222, ssh.socket → ssh.service, UFW), then Wazuh (gated on sizing, design doc id=25).
Server 5 FusionPBX — fix backup timer — sudo systemctl enable --now xzopia-backup.timer on Server 5. Verify with systemctl status xzopia-backup.timer.
Server 5 FusionPBX — Phase 2 VoIP security — run scripts/server5/voip-security.sh with TRUNK_IPS=("37.157.54.206" "185.91.41.29" "185.91.41.30"). Geofences port 5080 to trunk IPs; leaves 5061 open.
Server 5 FusionPBX — Phase 1 PBX config — FusionPBX web UI: both domains, extensions, gateways (UDP for now; switch to TLS at cutover), IVRs, ring groups, dialplan. See docs/server5-fusionpbx-build.md checklist. Confirm ring group 808 OOH destination with Simon.
Server 5 FusionPBX — SAN cert expansion — add pm.xzopiasecure.com to LE cert after both FusionPBX domains configured.
GDAP execution — partner.microsoft.com overdue since 1 May; run runbook, set up 5 groups, apply 22-role template to all 70 clients (docs/gdap-runbook.md). Can run in parallel with FusionPBX Phase 1.
Server 6 cleanup — ✅ MOOT: S6 reimaged 06 Jun 2026. CIPP/LUKS/AllowTcpForwarding all wiped. Next: 13-step server-prep checklist once S6 confirmed powered on post-outage (check Fasthosts panel).
Wazuh deployment plan — produce full plan, then begin on Server 6 or new Server 7.
CIPP via Azure — ✅ Operational (v10.5.2, Session 14). Follow-ups: CPV refresh, GDAP 15-group migration, SAM 6-secret cleanup, PS 7.4 EOL Nov 2026.
API credentials generated and stored in Vaultwarden — all generated/tested/stored except FusionPBX (pending Server 5) and dev-isolation for Xero/Giacom
MCP services deployed on Server 4 — all 10 services deployed and verified ✅ (3001–3010: NinjaOne, Xero, Giacom, Wiki.js, M365, SentinelOne, TacticalRMM, ITFlow, BookStack, TwentyCRM)
Wazuh MCP + AI triage loop — Wazuh alert → Claude API → ITFlow auto-ticket
Nextcloud — self-hosted document storage (OneDrive/SharePoint analog, non-Microsoft), backed by the existing Backblaze B2 as S3-compatible external storage. Sequenced AFTER the S6/Wazuh build (Simon's call, 20 Jul 2026). Fills a real gap: there is currently no home for binary/signed documents — BookStack is text/markdown-only with no attachment-upload capability (first hit trying to file Momentum's DocuSign-executed DPA). Needs a server-placement decision + the standard Stage 0–5 pipeline; worth weighing against the Phase 4 SecureAgent portal (possible overlap). Alternatives considered: Seafile (lighter sync), or filebrowser-on-B2 (simpler, loses versioning/sharing).
Security service go-live — £3–5/device/month line item on all 70 managed contracts
SecureAgent portal (React + Node.js) — currently placeholder nginx
Portal pulls from TacticalRMM, ITFlow, Xero, Giacom, Twenty CRM
Microsoft Teams bot — ticket creation from Teams messages
WhatsApp Business integration
Predictive maintenance AI
Self-healing TacticalRMM automation
Client onboarding automation
Draytek VigorACS MCP
TP-Link Omada MCP
Vulnerability management — daily CVE feed
GDAP automation
Web hosting migration (Astutium → Fasthosts)
Evaluate ECC (Everything Claude Code) toolkit — selectively adopt skills/agents that
benefit the Xzopia code set (esp. security-reviewer / AgentShield scanning, MCP
conventions, deploy/verify workflow patterns). Use selective install (--with/--without),
NOT the full 249-skill catalogue. Security-review and version-pin any hooks/agents
before install — it's third-party config that executes in-harness. Trial on Server 4
(dev) only, never the live box.
Server 4 is the LIVE PRODUCTION MCP (12 connectors, mcp.xzopiasecure.com) — never run setup-server4-full against the running box. Manual changes only; snapshot/backup before every change; verify all 12 connectors after. The 12th is Companies House (mcp-companieshouse, port 3012, added 18 Jul). Server 3 is now the LIVE Comet Server (Xzopia BaaS, backup.xzopiasecure.com) — client-revenue infra; backup-first before changes.
Server 2 uses Apache2 (not Nginx or Caddy) as reverse proxy.
Server 3 (Comet Server / Xzopia BaaS) uses Caddy. Server 4 (live MCP) uses Nginx (1.31.1 mainline).
Mandatory pre-setup checklist for every new server (see docs/new-server-preparation.md):
SSH in as root
adduser xzopia — set password, Full Name: Simon Killen