Last edited 3 weeks ago
by Peter A. Smode

KitsNet Operations:Services:Docker:High Availability NAS:Monitoring: Difference between revisions

Peter A. Smode (talk | contribs)
Created page with "= KitsNet NAS HA Monitoring = This document describes the design, implementation, operation, and recovery of the monitoring for the KitsNet Docker NAS high-availability system. It covers the read-only collector installed on Wort and the associated Zabbix instrumentation in the '''KitsNet NAS Hypervisor''' template. The monitoring release documented here is '''nas-ha-monitoring-r1'''. It supplements the Stage 4 revision 4 NAS HA implementation. It does not control failo..."
 
Peter A. Smode (talk | contribs)
mNo edit summary
 
Line 816: Line 816:
* Monitoring release: <code>nas-ha-monitoring-r1</code>
* Monitoring release: <code>nas-ha-monitoring-r1</code>
* Stage 4 r4 release tag: <code>nas-ha-stage4-r4-production-r2</code>
* Stage 4 r4 release tag: <code>nas-ha-stage4-r4-production-r2</code>
[[Category:High Availability NAS]]

Latest revision as of 09:28, 9 September 2026

1 KitsNet NAS HA Monitoring[edit | edit source]

This document describes the design, implementation, operation, and recovery of the monitoring for the KitsNet Docker NAS high-availability system. It covers the read-only collector installed on Wort and the associated Zabbix instrumentation in the KitsNet NAS Hypervisor template.

The monitoring release documented here is nas-ha-monitoring-r1. It supplements the Stage 4 revision 4 NAS HA implementation. It does not control failover, renew leases, attach or detach storage, restart NAS services, or enable or disable fencing.

1.1 Current accepted state[edit | edit source]

The monitoring implementation was accepted on 9 September 2026 with the following state:

  • Wort was the libvirt hypervisor and storage arbiter.
  • Fox (libvirt domain uv059) was the active NAS owner and Keepalived MASTER.
  • Anchor (libvirt domain uv060) was the standby NAS and Keepalived BACKUP.
  • The production logical volume was live-attached only to Fox as vdb.
  • Neither NAS domain contained the production disk in its persistent XML.
  • Wort authority named Fox, generation 13, active.
  • The active lease was valid with a 30-second limit.
  • Production fencing was enabled through a valid /etc/kitsnet-nas/fencing-enabled control file.
  • The collector returned HEALTHY_FOX, completed in approximately 66–70 milliseconds, and was accessible through Zabbix Agent 2.
  • Wort was running Zabbix Agent 2 version 7.0.28.

The accepted Git commit is:

1d1904d1ca9eedfa076358928fdf61999b7bc7a5

The annotated release tag is:

nas-ha-monitoring-r1

1.2 Monitoring objectives[edit | edit source]

The monitoring has four distinct objectives:

  1. Show whether the complete HA system is internally coherent now.
  2. Show how close the current owner is to losing lease freshness.
  3. Detect storage-attachment states that could threaten data integrity.
  4. Surface hypervisor scheduling pressure that could delay the owner guest and its heartbeat.

The lease metrics are the most direct early warning of an approaching takeover opportunity. vCPU scheduler delay is supporting evidence: it can explain why a guest is responding slowly, but scheduler delay alone does not initiate failover.

1.3 Architecture[edit | edit source]

The monitoring uses one read-only JSON master item and Zabbix dependent items.

Wort state files and libvirt XML
              |
              v
/usr/local/sbin/kitsnet-nas-ha-status
              |
              v
Zabbix Agent 2 item: kitsnet.nas.ha.status
              |
              v
Zabbix server JSONPath preprocessing
              |
              v
Dependent metrics, triggers, and lease graph

Every five seconds, Zabbix requests one JSON document from Wort. Zabbix then extracts all individual HA values from that document. This provides a time-consistent snapshot and avoids running a separate privileged command for every metric.

The pre-existing vCPU scheduler-delay collectors remain separate because they maintain per-domain counters and calculate change over time.

1.4 Data sources on Wort[edit | edit source]

Source Information used Persistence
/var/lib/kitsnet-nas-arbiter/owner.json Authoritative owner, generation, and active status Durable
/var/lib/kitsnet-nas-arbiter/fox.json Last published Fox intent and generation Durable
/var/lib/kitsnet-nas-arbiter/anchor.json Last published Anchor intent and generation Durable
/run/kitsnet-nas-arbiter/lease.json Current owner lease, generation, boot identifier, and monotonic timestamp Runtime; recreated after boot
/proc/sys/kernel/random/boot_id Confirms that the lease belongs to the current Wort boot Runtime
/proc/uptime through Python monotonic time Calculates lease age without wall-clock dependence Runtime
/etc/kitsnet-nas/fencing-enabled Confirms that production fencing is enabled with the required content, type, mode, and ownership Durable
libvirt live domain XML Shows which running domain owns the production LV as vdb Runtime
libvirt inactive domain XML Detects any prohibited persistent production-vdb definition Durable
libvirt domain state Confirms whether Fox and Anchor are running Runtime

The production source matched by the collector is:

/dev/T1/uv059-dkrnfs1

The collector recognizes only these owner-to-domain mappings:

HA owner Libvirt domain
Fox uv059
Anchor uv060

1.5 Collector design[edit | edit source]

1.5.1 Read-only behavior[edit | edit source]

kitsnet-nas-ha-status performs only file reads and these libvirt queries:

  • virsh domstate
  • virsh dumpxml
  • virsh dumpxml --inactive

It does not call kitsnet-nas-storage, kitsnet-nas-intent, virsh destroy, virsh attach-disk, or virsh detach-disk.

The collector always attempts to return valid JSON. If collection raises an exception, it returns collection_ok=0 and a bounded error string rather than emitting partial normal data.

1.5.2 Lease calculation[edit | edit source]

For an active authority, the configured lease limit is 30 seconds. During acquisition, before the authority becomes active, the limit is 60 seconds.

lease age = Wort monotonic time - stored lease monotonic time
lease headroom = applicable lease limit - lease age

A lease is valid only when all of the following are true:

  • The authority owner is Fox or Anchor.
  • The lease owner matches the authority owner.
  • The lease generation matches the authority generation.
  • The lease boot identifier matches the current Wort boot identifier.
  • The calculated age is between -1 second and the applicable lease limit.

The one-second negative allowance tolerates insignificant numeric timing effects. It does not permit a lease from another boot or generation.

1.5.3 Coherence classification[edit | edit source]

Code State Meaning
0 HEALTHY_FOX or HEALTHY_ANCHOR Authority, lease, intent, domain, fencing, and disk attachment agree.
1 TRANSITION_NO_OWNER or TRANSITION_ACQUIRING A normally brief interval exists between release and activation.
2 INCONSISTENT or DEGRADED_FENCING_CONTROL The state cannot be accepted as fully coherent, or the fencing control is invalid.
3 UNSAFE_ATTACHMENT The production LV appears live on more than one NAS domain, or appears in persistent domain XML.

The collector evaluates unsafe attachment before all other classifications. An invalid fencing control is evaluated next. A healthy result requires:

  • exactly one active authority owner;
  • a valid current lease;
  • exactly one live production-disk attachment on that same owner;
  • a matching MASTER intent and generation for the owner;
  • a BACKUP intent for the peer;
  • both NAS domains running;
  • valid production fencing control; and
  • no persistent production-disk attachment.

1.6 Installed files and Git mapping[edit | edit source]

Purpose Canonical Git path Deployment mirror or live path
Complete monitoring release src/nas-ha-monitoring/ Not deployed as a directory
JSON collector src/nas-ha-monitoring/kitsnet-nas-ha-status Git mirror: scripts/kitsnet-nas-ha-status
Wort: /usr/local/sbin/kitsnet-nas-ha-status
Zabbix user parameter src/nas-ha-monitoring/kitsnet-nas-ha-status.conf Git mirror: configs/nas/kitsnet-nas-ha-status.conf
Wort: /etc/zabbix/zabbix_agent2.d/kitsnet-nas-ha-status.conf
Restricted sudo rule src/nas-ha-monitoring/zabbix-kitsnet-nas-ha-status.sudoers Git mirror: configs/sudoers/zabbix-kitsnet-nas-ha-status
Wort: /etc/sudoers.d/zabbix-kitsnet-nas-ha-status
Zabbix template src/nas-ha-monitoring/KitsNet-NAS-Hypervisor-zabbix-7.0.yaml Git mirror: configs/nas/KitsNet-NAS-Hypervisor-zabbix-7.0.yaml
Imported into Zabbix
Deployment controller src/nas-ha-monitoring/install-monitoring-on-wort.sh Run from mgr1 as psmode
Release notes src/nas-ha-monitoring/README.txt Documentation within the Git release

The MediaWiki documentation is intentionally not stored in Git. MediaWiki page history provides document versioning.

1.6.1 Accepted file hashes[edit | edit source]

File SHA-256
kitsnet-nas-ha-status e3760679256dd282171c0cbc384d8d886e994ff83dcb3382e75abf4cf694c13d
kitsnet-nas-ha-status.conf c1a358b68d7b736907fe6a8c616cf8771ed8b0af0906f1183e01dcabdc6d4299
zabbix-kitsnet-nas-ha-status.sudoers f346f2ad0d9f2a30b9917130d15f3e496722c280fa9145cc3b9eebef6d9b183d
install-monitoring-on-wort.sh eec85001bd3496022bf98b1e8405adb6629805733144aef0d1aa1a82c2b2d7f3
README.txt source 4784f4faf7fe29be97b453e8690b8398ac4a80e5eb5ac6d5f5845b26d8c2c606
KitsNet-NAS-Hypervisor-zabbix-7.0.yaml 7ad02ed3c5fae21040dc2dba6f09dc31bd7837803cc86dd7d277e310d1bbf372

1.7 Zabbix implementation[edit | edit source]

1.7.1 Master item[edit | edit source]

Item Key Interval Type
NAS HA status JSON kitsnet.nas.ha.status 5 seconds Zabbix agent passive item, text

The user parameter is:

UserParameter=kitsnet.nas.ha.status,/usr/bin/sudo -n /usr/local/sbin/kitsnet-nas-ha-status

The corresponding sudo rule permits the Zabbix account to execute only the read-only collector as root:

zabbix ALL=(root) NOPASSWD: /usr/local/sbin/kitsnet-nas-ha-status

Root access is required because the durable arbiter state and fencing control are intentionally protected from ordinary users. The collector does not need SSH access to Fox or Anchor.

1.7.2 Dependent items[edit | edit source]

All following items are populated from the master JSON using JSONPath preprocessing.

Zabbix item key Value Normal stable value
kitsnet.nas.ha.collection.ok Whether the complete collection completed 1
kitsnet.nas.ha.collection.duration Collection duration in milliseconds Normally well below 1,000 ms
kitsnet.nas.ha.coherence.code Numeric overall classification 0
kitsnet.nas.ha.coherence.state Human-readable overall classification HEALTHY_FOX or HEALTHY_ANCHOR
kitsnet.nas.ha.authority.owner Wort authoritative owner fox or anchor
kitsnet.nas.ha.authority.generation Current authority generation Positive integer; must match the owner intent
kitsnet.nas.ha.authority.active Whether the authority has completed activation 1
kitsnet.nas.ha.lease.age Seconds since the last accepted heartbeat refresh Normally near 0–3 seconds
kitsnet.nas.ha.lease.limit Applicable lease limit 30 while active; 60 while acquiring
kitsnet.nas.ha.lease.headroom Seconds remaining before the applicable limit Normally near 27–30 seconds while active
kitsnet.nas.ha.lease.valid Owner, generation, boot, and age validation result 1
kitsnet.nas.ha.fencing.valid Strict fencing-control validation 1
kitsnet.nas.ha.intent.fox.state Last durable Fox intent MASTER or BACKUP, according to ownership
kitsnet.nas.ha.intent.fox.generation Fox intent generation Matches authority generation when Fox owns storage
kitsnet.nas.ha.intent.anchor.state Last durable Anchor intent MASTER or BACKUP, according to ownership
kitsnet.nas.ha.intent.anchor.generation Anchor intent generation Matches authority generation when Anchor owns storage
kitsnet.nas.ha.domain.fox.running Fox libvirt running state 1
kitsnet.nas.ha.domain.anchor.running Anchor libvirt running state 1
kitsnet.nas.ha.disk.live.count Total live production-vdb attachments 1
kitsnet.nas.ha.disk.live.owner Domain owner of the sole live production disk Same as authority owner
kitsnet.nas.ha.disk.fox.live Production-vdb count in Fox live XML 1 when Fox owns; otherwise 0
kitsnet.nas.ha.disk.anchor.live Production-vdb count in Anchor live XML 1 when Anchor owns; otherwise 0
kitsnet.nas.ha.disk.persistent.count Total production-vdb entries in inactive domain XML 0

1.7.3 Existing scheduler-delay items[edit | edit source]

Guest Zabbix key Interval
Fox / uv059 kitsnet.libvirt.vcpu.delay[uv059] 30 seconds
Anchor / uv060 kitsnet.libvirt.vcpu.delay[uv060] 30 seconds

The helper /usr/local/sbin/kitsnet-vcpu-delay-percent calculates the increase in cumulative libvirt vCPU delay divided by elapsed time and the number of vCPUs. The result estimates the percentage of requested vCPU execution time spent waiting for host scheduling.

The first sample after state creation returns zero because no prior counter exists. Counter rollback, non-positive elapsed time, or an invalid vCPU count also returns zero rather than a misleading negative value.

1.8 Trigger thresholds[edit | edit source]

1.8.1 HA triggers[edit | edit source]

Condition Persistence Severity Operational meaning
No master JSON data 30 seconds High Agent, user parameter, collector, or communication path is unavailable.
collection_ok=0 15 seconds High The collector is returning controlled failure snapshots.
Collection duration above 1,000 ms 5 minutes Warning Wort or libvirt query latency is elevated.
Coherence code 1 Continuously for 20 seconds Warning A release or acquisition transition is taking too long.
Coherence code 2 Continuously for 15 seconds High Authority, lease, intent, domain, disk, or fencing state is inconsistent.
Coherence code 3 Immediate Disaster Storage attachment is unsafe.
Lease age above 10 seconds Continuously for 15 seconds Warning Heartbeat delivery is delayed well beyond its normal two-second interval.
Lease age above 20 seconds Continuously for 10 seconds High The owner has little active-lease headroom remaining.
Lease age above 25 seconds Immediate Disaster The active lease is within five seconds of its 30-second limit.
Lease invalid Continuously for 10 seconds High Owner, generation, boot identifier, or age validation failed.
Fencing control invalid or disabled Immediate Disaster Automatic fencing cannot be considered fully operational.
Fox or Anchor domain not running Continuously for 15 seconds High A NAS member is unavailable, including a node deliberately destroyed during fencing.
More than one live production-disk attachment Immediate Disaster Possible double attachment; treat as a data-integrity emergency.
No live production-disk attachment Continuously for 20 seconds High No NAS node has served as production storage owner for too long.
Any persistent production-disk attachment Immediate Disaster The deployment violates the live-attachment-only design.

The lease-age triggers intentionally overlap. During a worsening event, Zabbix may show Warning, High, and Disaster problems together. The highest severity describes the current urgency; the earlier events preserve the escalation timeline.

1.8.2 Scheduler-delay triggers[edit | edit source]

The same thresholds apply separately to Fox and Anchor.

Delay Persistence Severity
Above 5 percent 5 minutes Warning
Above 15 percent 2 minutes High
Above 30 percent 1 minute Disaster

Scheduler delay is a pressure indicator, not proof that a failover is imminent. Interpret it alongside lease age, collector duration, guest service monitoring, and host resource pressure.

1.9 Graphs[edit | edit source]

The template contains two graphs:

  • KitsNet NAS HA Lease — lease age, lease headroom, and the applicable limit.
  • KitsNet NAS vCPU Scheduler Delay — Fox and Anchor scheduler-delay percentages.

The lease graph is the clearest historical view of proximity to a stale-owner condition. Under normal operation, lease age remains near the bottom of the graph and headroom remains near the 30-second limit.

1.10 Initial deployment[edit | edit source]

1.10.1 Preconditions[edit | edit source]

Before installation:

  • Stage 4 r4 NAS HA must already be installed and accepted.
  • Wort must have Zabbix Agent 2, Python 3, sudo, and libvirt tools.
  • The psmode identity on mgr1 must have a working SSH agent identity for Wort.
  • Do not use IdentitiesOnly=yes.
  • Connect to Wort as psmode and use sudo on Wort; do not open inbound root SSH.
  • Confirm that the NAS state is stable before using installation acceptance checks.

1.10.2 Install or reinstall from Git[edit | edit source]

Run this on mgr1 as psmode:

for i in {1..10}; do echo; done
repo=/srv/git/KNdkr
report=/tmp/nas-ha-monitoring-install.log

install_nas_monitoring()
(
    set -Eeuo pipefail
    cd "$repo"
    test "$(git branch --show-current)" = master
    git fetch --quiet origin master
    test "$(git rev-parse HEAD)" = "$(git rev-parse origin/master)"
    git tag -v nas-ha-monitoring-r1 2>/dev/null ||
        git tag -n99 nas-ha-monitoring-r1
    src/nas-ha-monitoring/install-monitoring-on-wort.sh
)

install_nas_monitoring 2>&1 | tee "$report"
install_rc=${PIPESTATUS[0]}
echo "install_rc=$install_rc"
echo "saved_output=$report"

The installer:

  1. copies the three Wort deployment files through an SSH session initiated as psmode;
  2. validates Python syntax and the sudoers fragment before installation;
  3. refuses to overwrite a different existing deployment;
  4. installs the collector as mode 0755;
  5. installs the Zabbix configuration as mode 0644;
  6. installs the sudoers fragment as mode 0440;
  7. restores SELinux file contexts when restorecon is available;
  8. runs the collector as the Zabbix account;
  9. requires a healthy, valid, single-owner production state; and
  10. restarts only zabbix-agent2.service.

It does not restart Keepalived, NFS, Veeam, heartbeat, or transition services. It does not alter fencing or disk attachments.

1.10.3 Import the Zabbix template[edit | edit source]

In Zabbix:

  1. Open Data collection → Templates.
  2. Select Import.
  3. Select src/nas-ha-monitoring/KitsNet-NAS-Hypervisor-zabbix-7.0.yaml, or a verified copy of that file.
  4. Enable Create new.
  5. Enable Update existing.
  6. Leave Delete missing disabled.
  7. Review the proposed changes and import.

The release preserves the existing template UUID, vCPU item UUIDs, vCPU trigger UUIDs, and scheduler-delay graph UUID. The existing template group name Templates/Montioring, including its historical spelling, is also preserved.

If the KitsNet NAS Hypervisor template is already linked to the Wort host, the new items are inherited automatically. Otherwise, link the template to Wort after import.

1.11 Verification[edit | edit source]

1.11.1 Wort collector verification[edit | edit source]

Run from mgr1 as psmode:

for i in {1..10}; do echo; done
report=/tmp/nas-ha-monitoring-verification.log

verify_nas_monitoring()
(
    set -Eeuo pipefail
    ssh -o BatchMode=yes psmode@wort.lan.kitsnet.us 'sudo -n bash -s' <<'REMOTE'
sha256sum \
    /usr/local/sbin/kitsnet-nas-ha-status \
    /etc/zabbix/zabbix_agent2.d/kitsnet-nas-ha-status.conf \
    /etc/sudoers.d/zabbix-kitsnet-nas-ha-status
systemctl is-active zabbix-agent2.service
sudo -u zabbix sudo -n /usr/local/sbin/kitsnet-nas-ha-status |
    python3 -m json.tool
zabbix_agent2 -t kitsnet.nas.ha.status
REMOTE
)

verify_nas_monitoring 2>&1 | tee "$report"
verification_rc=${PIPESTATUS[0]}
echo "verification_rc=$verification_rc"
echo "saved_output=$report"

A normal stable result has:

collection_ok=1
coherence_code=0
fencing_control_valid=1
lease_valid=1
live_production_vdb_count=1
persistent_production_vdb_count=0

The owner may legitimately be Fox or Anchor. Do not encode generation 13 or Fox ownership as permanent monitoring expectations.

1.11.2 Zabbix verification[edit | edit source]

After at least 30 seconds:

  1. Open Monitoring → Latest data.
  2. Select the Wort host.
  3. Filter the item name by NAS HA.
  4. Confirm that the master item and every dependent item are supported and have current timestamps.
  5. Confirm that the normal stable values listed above are present.
  6. Open Monitoring → Hosts → Graphs for Wort and inspect KitsNet NAS HA Lease.
  7. Confirm that no new HA problem is open under Monitoring → Problems.

If the master item has a value but dependent items do not, inspect preprocessing errors on the dependent items. If the master item is unsupported, test the user parameter directly on Wort before changing the template.

1.12 Operational interpretation[edit | edit source]

1.12.1 Normal owner heartbeat[edit | edit source]

The guest heartbeat normally refreshes Wort approximately every two seconds. A typical observed lease age is therefore a small number of seconds. Occasional modest increases are not themselves a failover event.

The validated load-and-delay test produced 15 successful refreshes and a maximum lease age of 8.662 seconds under simultaneous CPU pressure, memory pressure, NFS I/O, a competing MASTER request, and 500 ms per-packet SSH delay. The active lease limit remained 30 seconds, Fox was not fenced, and Anchor did not acquire the disk.

1.12.2 What approaching failover looks like[edit | edit source]

A likely developing owner-path problem normally appears in this order:

  1. Lease age rises above its usual range.
  2. Lease headroom falls by the same amount.
  3. Scheduler delay, collector duration, guest service alarms, or network alarms may provide an explanation.
  4. The 10-second lease-age Warning opens if the condition persists.
  5. The 20-second High alarm opens if the lease continues aging.
  6. The 25-second Disaster alarm indicates less than five seconds before the active lease limit.
  7. Lease validity becomes false after expiry.

Lease expiry by itself does not attach the disk to the peer. The fencing path is exercised when a competing eligible MASTER requests takeover and the complete authority, lease, VM, and attachment checks permit it.

1.12.3 Successful failover[edit | edit source]

During an orderly or automatic handoff, a brief coherence code 1 and a brief zero live-disk count can be normal. The alert persistence periods allow those transitions to complete without immediately opening a problem.

After completion, all of the following should change coherently:

  • authority owner;
  • authority generation;
  • owner MASTER intent;
  • peer BACKUP intent;
  • live disk owner; and
  • coherence state.

The final state should be HEALTHY_ANCHOR or HEALTHY_FOX with a freshly renewed lease.

1.13 Alert response[edit | edit source]

Alert First response Prohibited shortcut
Lease age Warning or High Compare lease age with vCPU delay, guest reachability, heartbeat service status, Wort load, and SSH latency. Preserve evidence before intervening. Do not force a failover solely because lease age briefly increased.
Lease within five seconds of expiry Treat as urgent. Determine whether the owner guest is running and whether its heartbeat/control path has stalled. Watch authority and attachment state continuously. Do not manually attach the production LV to the peer.
Coherence inconsistent Compare owner, generation, intents, lease validity, domain states, and live owner from the same JSON snapshot. Do not restart both NAS nodes together.
Unsafe attachment Stop ordinary recovery activity and inspect live and inactive XML on both domains. Protect data integrity first. Do not mount or start NFS on either node until exclusive ownership is proven.
Fencing control invalid Inspect file existence, regular-file type, mode 0600, root ownership, and content 1. Determine whether disablement was intentional maintenance. Do not recreate the flag during a controlled rebuild or shutdown procedure that requires fencing disabled.
NAS domain not running Determine whether it was deliberately fenced, administratively stopped, or failed. Confirm where the disk and authority ended. Do not restart a destroyed former owner until storage ownership is known.
Collector failure or no data Check Zabbix Agent 2, the user parameter, sudo validation, SELinux denials, and libvirt query access. Do not weaken the sudo rule to unrestricted commands.

1.14 Wort major-version rebuild[edit | edit source]

Monitoring restoration occurs only after the base hypervisor, libvirt configuration, Stage 4 r4 code, authority-state recovery, and controlled HA acceptance are complete.

1.14.1 Rebuild sequence[edit | edit source]

  1. Keep production fencing disabled during the controlled rebuild and recovery procedure, as required by the main NAS HA rebuild runbook.
  2. Reinstall Zabbix Agent 2 from the approved Zabbix 7.0 repository.
  3. Restore the libvirt domains, network, pools, and live-attachment-only storage design.
  4. Restore the Stage 4 r4 Wort scripts and arbiter state according to the NAS HA rebuild procedure.
  5. Complete the fencing-disabled HA acceptance and establish a coherent owner.
  6. Restore the monitoring collector from Git by running src/nas-ha-monitoring/install-monitoring-on-wort.sh on mgr1 as psmode.
  7. Confirm that the collector works through Zabbix Agent 2.
  8. Re-enable production fencing only through the separately validated HA enablement procedure.
  9. Confirm fencing_control_valid=1 and coherence_code=0.
  10. Import the template only if the Zabbix server configuration was also lost or rebuilt. Ordinarily the existing template remains intact.
  11. Verify Latest data, graphs, and Problems.

The monitoring installer expects the production HA system to be coherent and fencing to be valid at its final acceptance check. Therefore, during a rebuild, run it after the HA solution has returned to its final accepted operational state.

1.15 Zabbix server or template recovery[edit | edit source]

If the Zabbix server loses the template:

  1. Obtain configs/nas/KitsNet-NAS-Hypervisor-zabbix-7.0.yaml from commit 1d1904d1ca9eedfa076358928fdf61999b7bc7a5 or tag nas-ha-monitoring-r1.
  2. Verify its SHA-256 hash.
  3. Import with Create new and Update existing enabled.
  4. Keep Delete missing disabled unless a separately reviewed change explicitly requires deletion.
  5. Link KitsNet NAS Hypervisor to Wort if the linkage was not restored with the Zabbix database.
  6. Allow at least 30 seconds for initial collection and dependent-item processing.
  7. Perform the Zabbix verification procedure above.

1.16 Security and secrets[edit | edit source]

No private SSH key, password, keytab, Zabbix TLS secret, or other private credential is stored in the monitoring release.

The Zabbix account receives only this additional sudo permission:

/usr/local/sbin/kitsnet-nas-ha-status

The collector has no command-line arguments, rejects no user-supplied domain name because none is accepted, and uses fixed state paths, domain names, and production source paths. This prevents the Zabbix item key from being used as a general libvirt or shell-command interface.

The deployment controller follows KitsNet SSH policy:

  • connections originate on mgr1 as psmode;
  • the working identity comes from the SSH agent;
  • IdentitiesOnly=yes is not used; and
  • privilege escalation occurs through sudo on Wort.

1.17 Known scope boundaries[edit | edit source]

The r1 collector measures the authoritative Wort-side HA control and attachment state. It does not directly query inside Fox or Anchor for:

  • Keepalived process state or advertisement age;
  • NFS service state or client latency;
  • Veeam service state;
  • guest filesystem mount state;
  • guest heartbeat unit restart count; or
  • guest network packet loss.

Those guest and service metrics should remain on the Fox and Anchor host templates and can be presented beside the Wort HA metrics in a Zabbix dashboard. Wort is deliberately not given broad remote credentials merely to collect them.

1.18 Change control[edit | edit source]

Changes should begin in src/nas-ha-monitoring/. After testing, update the deployment mirrors in scripts/, configs/nas/, and configs/sudoers/. Validate exact byte equality between canonical files, mirrors, and deployed files before committing.

Stage only the monitoring paths. The repository may contain unrelated modified or untracked work that must not enter a monitoring commit.

For every release:

  1. validate Python and shell syntax;
  2. validate the sudoers fragment with visudo -cf;
  3. import and validate the Zabbix YAML against the deployed Zabbix major version;
  4. test the collector as the Zabbix account;
  5. verify healthy HA state without initiating a failover;
  6. review the staged path list and staged diff;
  7. commit;
  8. create an annotated monitoring release tag;
  9. push both the branch and tag; and
  10. verify the remote branch and peeled tag commit IDs.

1.19 References[edit | edit source]