Created page with "= KitsNet NAS HA Monitoring = This document describes the design, implementation, operation, and recovery of the monitoring for the KitsNet Docker NAS high-availability system. It covers the read-only collector installed on Wort and the associated Zabbix instrumentation in the '''KitsNet NAS Hypervisor''' template. The monitoring release documented here is '''nas-ha-monitoring-r1'''. It supplements the Stage 4 revision 4 NAS HA implementation. It does not control failo..." |
mNo edit summary |
||
| Line 816: | Line 816: | ||
* Monitoring release: <code>nas-ha-monitoring-r1</code> | * Monitoring release: <code>nas-ha-monitoring-r1</code> | ||
* Stage 4 r4 release tag: <code>nas-ha-stage4-r4-production-r2</code> | * Stage 4 r4 release tag: <code>nas-ha-stage4-r4-production-r2</code> | ||
[[Category:High Availability NAS]] | |||
Latest revision as of 09:28, 9 September 2026
1 KitsNet NAS HA Monitoring[edit | edit source]
This document describes the design, implementation, operation, and recovery of the monitoring for the KitsNet Docker NAS high-availability system. It covers the read-only collector installed on Wort and the associated Zabbix instrumentation in the KitsNet NAS Hypervisor template.
The monitoring release documented here is nas-ha-monitoring-r1. It supplements the Stage 4 revision 4 NAS HA implementation. It does not control failover, renew leases, attach or detach storage, restart NAS services, or enable or disable fencing.
1.1 Current accepted state[edit | edit source]
The monitoring implementation was accepted on 9 September 2026 with the following state:
- Wort was the libvirt hypervisor and storage arbiter.
- Fox (libvirt domain
uv059) was the active NAS owner and Keepalived MASTER. - Anchor (libvirt domain
uv060) was the standby NAS and Keepalived BACKUP. - The production logical volume was live-attached only to Fox as
vdb. - Neither NAS domain contained the production disk in its persistent XML.
- Wort authority named Fox, generation 13, active.
- The active lease was valid with a 30-second limit.
- Production fencing was enabled through a valid
/etc/kitsnet-nas/fencing-enabledcontrol file. - The collector returned
HEALTHY_FOX, completed in approximately 66–70 milliseconds, and was accessible through Zabbix Agent 2. - Wort was running Zabbix Agent 2 version 7.0.28.
The accepted Git commit is:
1d1904d1ca9eedfa076358928fdf61999b7bc7a5
The annotated release tag is:
nas-ha-monitoring-r1
1.2 Monitoring objectives[edit | edit source]
The monitoring has four distinct objectives:
- Show whether the complete HA system is internally coherent now.
- Show how close the current owner is to losing lease freshness.
- Detect storage-attachment states that could threaten data integrity.
- Surface hypervisor scheduling pressure that could delay the owner guest and its heartbeat.
The lease metrics are the most direct early warning of an approaching takeover opportunity. vCPU scheduler delay is supporting evidence: it can explain why a guest is responding slowly, but scheduler delay alone does not initiate failover.
1.3 Architecture[edit | edit source]
The monitoring uses one read-only JSON master item and Zabbix dependent items.
Wort state files and libvirt XML
|
v
/usr/local/sbin/kitsnet-nas-ha-status
|
v
Zabbix Agent 2 item: kitsnet.nas.ha.status
|
v
Zabbix server JSONPath preprocessing
|
v
Dependent metrics, triggers, and lease graph
Every five seconds, Zabbix requests one JSON document from Wort. Zabbix then extracts all individual HA values from that document. This provides a time-consistent snapshot and avoids running a separate privileged command for every metric.
The pre-existing vCPU scheduler-delay collectors remain separate because they maintain per-domain counters and calculate change over time.
1.4 Data sources on Wort[edit | edit source]
| Source | Information used | Persistence |
|---|---|---|
/var/lib/kitsnet-nas-arbiter/owner.json
|
Authoritative owner, generation, and active status | Durable |
/var/lib/kitsnet-nas-arbiter/fox.json
|
Last published Fox intent and generation | Durable |
/var/lib/kitsnet-nas-arbiter/anchor.json
|
Last published Anchor intent and generation | Durable |
/run/kitsnet-nas-arbiter/lease.json
|
Current owner lease, generation, boot identifier, and monotonic timestamp | Runtime; recreated after boot |
/proc/sys/kernel/random/boot_id
|
Confirms that the lease belongs to the current Wort boot | Runtime |
/proc/uptime through Python monotonic time
|
Calculates lease age without wall-clock dependence | Runtime |
/etc/kitsnet-nas/fencing-enabled
|
Confirms that production fencing is enabled with the required content, type, mode, and ownership | Durable |
| libvirt live domain XML | Shows which running domain owns the production LV as vdb
|
Runtime |
| libvirt inactive domain XML | Detects any prohibited persistent production-vdb definition
|
Durable |
| libvirt domain state | Confirms whether Fox and Anchor are running | Runtime |
The production source matched by the collector is:
/dev/T1/uv059-dkrnfs1
The collector recognizes only these owner-to-domain mappings:
| HA owner | Libvirt domain |
|---|---|
| Fox | uv059
|
| Anchor | uv060
|
1.5 Collector design[edit | edit source]
1.5.1 Read-only behavior[edit | edit source]
kitsnet-nas-ha-status performs only file reads and these libvirt queries:
virsh domstatevirsh dumpxmlvirsh dumpxml --inactive
It does not call kitsnet-nas-storage, kitsnet-nas-intent, virsh destroy, virsh attach-disk, or virsh detach-disk.
The collector always attempts to return valid JSON. If collection raises an exception, it returns collection_ok=0 and a bounded error string rather than emitting partial normal data.
1.5.2 Lease calculation[edit | edit source]
For an active authority, the configured lease limit is 30 seconds. During acquisition, before the authority becomes active, the limit is 60 seconds.
lease age = Wort monotonic time - stored lease monotonic time lease headroom = applicable lease limit - lease age
A lease is valid only when all of the following are true:
- The authority owner is Fox or Anchor.
- The lease owner matches the authority owner.
- The lease generation matches the authority generation.
- The lease boot identifier matches the current Wort boot identifier.
- The calculated age is between -1 second and the applicable lease limit.
The one-second negative allowance tolerates insignificant numeric timing effects. It does not permit a lease from another boot or generation.
1.5.3 Coherence classification[edit | edit source]
| Code | State | Meaning |
|---|---|---|
| 0 | HEALTHY_FOX or HEALTHY_ANCHOR
|
Authority, lease, intent, domain, fencing, and disk attachment agree. |
| 1 | TRANSITION_NO_OWNER or TRANSITION_ACQUIRING
|
A normally brief interval exists between release and activation. |
| 2 | INCONSISTENT or DEGRADED_FENCING_CONTROL
|
The state cannot be accepted as fully coherent, or the fencing control is invalid. |
| 3 | UNSAFE_ATTACHMENT
|
The production LV appears live on more than one NAS domain, or appears in persistent domain XML. |
The collector evaluates unsafe attachment before all other classifications. An invalid fencing control is evaluated next. A healthy result requires:
- exactly one active authority owner;
- a valid current lease;
- exactly one live production-disk attachment on that same owner;
- a matching MASTER intent and generation for the owner;
- a BACKUP intent for the peer;
- both NAS domains running;
- valid production fencing control; and
- no persistent production-disk attachment.
1.6 Installed files and Git mapping[edit | edit source]
| Purpose | Canonical Git path | Deployment mirror or live path |
|---|---|---|
| Complete monitoring release | src/nas-ha-monitoring/
|
Not deployed as a directory |
| JSON collector | src/nas-ha-monitoring/kitsnet-nas-ha-status
|
Git mirror: scripts/kitsnet-nas-ha-statusWort: /usr/local/sbin/kitsnet-nas-ha-status
|
| Zabbix user parameter | src/nas-ha-monitoring/kitsnet-nas-ha-status.conf
|
Git mirror: configs/nas/kitsnet-nas-ha-status.confWort: /etc/zabbix/zabbix_agent2.d/kitsnet-nas-ha-status.conf
|
| Restricted sudo rule | src/nas-ha-monitoring/zabbix-kitsnet-nas-ha-status.sudoers
|
Git mirror: configs/sudoers/zabbix-kitsnet-nas-ha-statusWort: /etc/sudoers.d/zabbix-kitsnet-nas-ha-status
|
| Zabbix template | src/nas-ha-monitoring/KitsNet-NAS-Hypervisor-zabbix-7.0.yaml
|
Git mirror: configs/nas/KitsNet-NAS-Hypervisor-zabbix-7.0.yamlImported into Zabbix |
| Deployment controller | src/nas-ha-monitoring/install-monitoring-on-wort.sh
|
Run from mgr1 as psmode
|
| Release notes | src/nas-ha-monitoring/README.txt
|
Documentation within the Git release |
The MediaWiki documentation is intentionally not stored in Git. MediaWiki page history provides document versioning.
1.6.1 Accepted file hashes[edit | edit source]
| File | SHA-256 |
|---|---|
kitsnet-nas-ha-status
|
e3760679256dd282171c0cbc384d8d886e994ff83dcb3382e75abf4cf694c13d
|
kitsnet-nas-ha-status.conf
|
c1a358b68d7b736907fe6a8c616cf8771ed8b0af0906f1183e01dcabdc6d4299
|
zabbix-kitsnet-nas-ha-status.sudoers
|
f346f2ad0d9f2a30b9917130d15f3e496722c280fa9145cc3b9eebef6d9b183d
|
install-monitoring-on-wort.sh
|
eec85001bd3496022bf98b1e8405adb6629805733144aef0d1aa1a82c2b2d7f3
|
README.txt source
|
4784f4faf7fe29be97b453e8690b8398ac4a80e5eb5ac6d5f5845b26d8c2c606
|
KitsNet-NAS-Hypervisor-zabbix-7.0.yaml
|
7ad02ed3c5fae21040dc2dba6f09dc31bd7837803cc86dd7d277e310d1bbf372
|
1.7 Zabbix implementation[edit | edit source]
1.7.1 Master item[edit | edit source]
| Item | Key | Interval | Type |
|---|---|---|---|
| NAS HA status JSON | kitsnet.nas.ha.status
|
5 seconds | Zabbix agent passive item, text |
The user parameter is:
UserParameter=kitsnet.nas.ha.status,/usr/bin/sudo -n /usr/local/sbin/kitsnet-nas-ha-status
The corresponding sudo rule permits the Zabbix account to execute only the read-only collector as root:
zabbix ALL=(root) NOPASSWD: /usr/local/sbin/kitsnet-nas-ha-status
Root access is required because the durable arbiter state and fencing control are intentionally protected from ordinary users. The collector does not need SSH access to Fox or Anchor.
1.7.2 Dependent items[edit | edit source]
All following items are populated from the master JSON using JSONPath preprocessing.
| Zabbix item key | Value | Normal stable value |
|---|---|---|
kitsnet.nas.ha.collection.ok
|
Whether the complete collection completed | 1
|
kitsnet.nas.ha.collection.duration
|
Collection duration in milliseconds | Normally well below 1,000 ms |
kitsnet.nas.ha.coherence.code
|
Numeric overall classification | 0
|
kitsnet.nas.ha.coherence.state
|
Human-readable overall classification | HEALTHY_FOX or HEALTHY_ANCHOR
|
kitsnet.nas.ha.authority.owner
|
Wort authoritative owner | fox or anchor
|
kitsnet.nas.ha.authority.generation
|
Current authority generation | Positive integer; must match the owner intent |
kitsnet.nas.ha.authority.active
|
Whether the authority has completed activation | 1
|
kitsnet.nas.ha.lease.age
|
Seconds since the last accepted heartbeat refresh | Normally near 0–3 seconds |
kitsnet.nas.ha.lease.limit
|
Applicable lease limit | 30 while active; 60 while acquiring
|
kitsnet.nas.ha.lease.headroom
|
Seconds remaining before the applicable limit | Normally near 27–30 seconds while active |
kitsnet.nas.ha.lease.valid
|
Owner, generation, boot, and age validation result | 1
|
kitsnet.nas.ha.fencing.valid
|
Strict fencing-control validation | 1
|
kitsnet.nas.ha.intent.fox.state
|
Last durable Fox intent | MASTER or BACKUP, according to ownership
|
kitsnet.nas.ha.intent.fox.generation
|
Fox intent generation | Matches authority generation when Fox owns storage |
kitsnet.nas.ha.intent.anchor.state
|
Last durable Anchor intent | MASTER or BACKUP, according to ownership
|
kitsnet.nas.ha.intent.anchor.generation
|
Anchor intent generation | Matches authority generation when Anchor owns storage |
kitsnet.nas.ha.domain.fox.running
|
Fox libvirt running state | 1
|
kitsnet.nas.ha.domain.anchor.running
|
Anchor libvirt running state | 1
|
kitsnet.nas.ha.disk.live.count
|
Total live production-vdb attachments
|
1
|
kitsnet.nas.ha.disk.live.owner
|
Domain owner of the sole live production disk | Same as authority owner |
kitsnet.nas.ha.disk.fox.live
|
Production-vdb count in Fox live XML
|
1 when Fox owns; otherwise 0
|
kitsnet.nas.ha.disk.anchor.live
|
Production-vdb count in Anchor live XML
|
1 when Anchor owns; otherwise 0
|
kitsnet.nas.ha.disk.persistent.count
|
Total production-vdb entries in inactive domain XML
|
0
|
1.7.3 Existing scheduler-delay items[edit | edit source]
| Guest | Zabbix key | Interval |
|---|---|---|
Fox / uv059
|
kitsnet.libvirt.vcpu.delay[uv059]
|
30 seconds |
Anchor / uv060
|
kitsnet.libvirt.vcpu.delay[uv060]
|
30 seconds |
The helper /usr/local/sbin/kitsnet-vcpu-delay-percent calculates the increase in cumulative libvirt vCPU delay divided by elapsed time and the number of vCPUs. The result estimates the percentage of requested vCPU execution time spent waiting for host scheduling.
The first sample after state creation returns zero because no prior counter exists. Counter rollback, non-positive elapsed time, or an invalid vCPU count also returns zero rather than a misleading negative value.
1.8 Trigger thresholds[edit | edit source]
1.8.1 HA triggers[edit | edit source]
| Condition | Persistence | Severity | Operational meaning |
|---|---|---|---|
| No master JSON data | 30 seconds | High | Agent, user parameter, collector, or communication path is unavailable. |
collection_ok=0
|
15 seconds | High | The collector is returning controlled failure snapshots. |
| Collection duration above 1,000 ms | 5 minutes | Warning | Wort or libvirt query latency is elevated. |
| Coherence code 1 | Continuously for 20 seconds | Warning | A release or acquisition transition is taking too long. |
| Coherence code 2 | Continuously for 15 seconds | High | Authority, lease, intent, domain, disk, or fencing state is inconsistent. |
| Coherence code 3 | Immediate | Disaster | Storage attachment is unsafe. |
| Lease age above 10 seconds | Continuously for 15 seconds | Warning | Heartbeat delivery is delayed well beyond its normal two-second interval. |
| Lease age above 20 seconds | Continuously for 10 seconds | High | The owner has little active-lease headroom remaining. |
| Lease age above 25 seconds | Immediate | Disaster | The active lease is within five seconds of its 30-second limit. |
| Lease invalid | Continuously for 10 seconds | High | Owner, generation, boot identifier, or age validation failed. |
| Fencing control invalid or disabled | Immediate | Disaster | Automatic fencing cannot be considered fully operational. |
| Fox or Anchor domain not running | Continuously for 15 seconds | High | A NAS member is unavailable, including a node deliberately destroyed during fencing. |
| More than one live production-disk attachment | Immediate | Disaster | Possible double attachment; treat as a data-integrity emergency. |
| No live production-disk attachment | Continuously for 20 seconds | High | No NAS node has served as production storage owner for too long. |
| Any persistent production-disk attachment | Immediate | Disaster | The deployment violates the live-attachment-only design. |
The lease-age triggers intentionally overlap. During a worsening event, Zabbix may show Warning, High, and Disaster problems together. The highest severity describes the current urgency; the earlier events preserve the escalation timeline.
1.8.2 Scheduler-delay triggers[edit | edit source]
The same thresholds apply separately to Fox and Anchor.
| Delay | Persistence | Severity |
|---|---|---|
| Above 5 percent | 5 minutes | Warning |
| Above 15 percent | 2 minutes | High |
| Above 30 percent | 1 minute | Disaster |
Scheduler delay is a pressure indicator, not proof that a failover is imminent. Interpret it alongside lease age, collector duration, guest service monitoring, and host resource pressure.
1.9 Graphs[edit | edit source]
The template contains two graphs:
- KitsNet NAS HA Lease — lease age, lease headroom, and the applicable limit.
- KitsNet NAS vCPU Scheduler Delay — Fox and Anchor scheduler-delay percentages.
The lease graph is the clearest historical view of proximity to a stale-owner condition. Under normal operation, lease age remains near the bottom of the graph and headroom remains near the 30-second limit.
1.10 Initial deployment[edit | edit source]
1.10.1 Preconditions[edit | edit source]
Before installation:
- Stage 4 r4 NAS HA must already be installed and accepted.
- Wort must have Zabbix Agent 2, Python 3, sudo, and libvirt tools.
- The
psmodeidentity onmgr1must have a working SSH agent identity for Wort. - Do not use
IdentitiesOnly=yes. - Connect to Wort as
psmodeand use sudo on Wort; do not open inbound root SSH. - Confirm that the NAS state is stable before using installation acceptance checks.
1.10.2 Install or reinstall from Git[edit | edit source]
Run this on mgr1 as psmode:
for i in {1..10}; do echo; done
repo=/srv/git/KNdkr
report=/tmp/nas-ha-monitoring-install.log
install_nas_monitoring()
(
set -Eeuo pipefail
cd "$repo"
test "$(git branch --show-current)" = master
git fetch --quiet origin master
test "$(git rev-parse HEAD)" = "$(git rev-parse origin/master)"
git tag -v nas-ha-monitoring-r1 2>/dev/null ||
git tag -n99 nas-ha-monitoring-r1
src/nas-ha-monitoring/install-monitoring-on-wort.sh
)
install_nas_monitoring 2>&1 | tee "$report"
install_rc=${PIPESTATUS[0]}
echo "install_rc=$install_rc"
echo "saved_output=$report"
The installer:
- copies the three Wort deployment files through an SSH session initiated as
psmode; - validates Python syntax and the sudoers fragment before installation;
- refuses to overwrite a different existing deployment;
- installs the collector as mode 0755;
- installs the Zabbix configuration as mode 0644;
- installs the sudoers fragment as mode 0440;
- restores SELinux file contexts when
restoreconis available; - runs the collector as the Zabbix account;
- requires a healthy, valid, single-owner production state; and
- restarts only
zabbix-agent2.service.
It does not restart Keepalived, NFS, Veeam, heartbeat, or transition services. It does not alter fencing or disk attachments.
1.10.3 Import the Zabbix template[edit | edit source]
In Zabbix:
- Open Data collection → Templates.
- Select Import.
- Select
src/nas-ha-monitoring/KitsNet-NAS-Hypervisor-zabbix-7.0.yaml, or a verified copy of that file. - Enable Create new.
- Enable Update existing.
- Leave Delete missing disabled.
- Review the proposed changes and import.
The release preserves the existing template UUID, vCPU item UUIDs, vCPU trigger UUIDs, and scheduler-delay graph UUID. The existing template group name Templates/Montioring, including its historical spelling, is also preserved.
If the KitsNet NAS Hypervisor template is already linked to the Wort host, the new items are inherited automatically. Otherwise, link the template to Wort after import.
1.11 Verification[edit | edit source]
1.11.1 Wort collector verification[edit | edit source]
Run from mgr1 as psmode:
for i in {1..10}; do echo; done
report=/tmp/nas-ha-monitoring-verification.log
verify_nas_monitoring()
(
set -Eeuo pipefail
ssh -o BatchMode=yes psmode@wort.lan.kitsnet.us 'sudo -n bash -s' <<'REMOTE'
sha256sum \
/usr/local/sbin/kitsnet-nas-ha-status \
/etc/zabbix/zabbix_agent2.d/kitsnet-nas-ha-status.conf \
/etc/sudoers.d/zabbix-kitsnet-nas-ha-status
systemctl is-active zabbix-agent2.service
sudo -u zabbix sudo -n /usr/local/sbin/kitsnet-nas-ha-status |
python3 -m json.tool
zabbix_agent2 -t kitsnet.nas.ha.status
REMOTE
)
verify_nas_monitoring 2>&1 | tee "$report"
verification_rc=${PIPESTATUS[0]}
echo "verification_rc=$verification_rc"
echo "saved_output=$report"
A normal stable result has:
collection_ok=1 coherence_code=0 fencing_control_valid=1 lease_valid=1 live_production_vdb_count=1 persistent_production_vdb_count=0
The owner may legitimately be Fox or Anchor. Do not encode generation 13 or Fox ownership as permanent monitoring expectations.
1.11.2 Zabbix verification[edit | edit source]
After at least 30 seconds:
- Open Monitoring → Latest data.
- Select the Wort host.
- Filter the item name by
NAS HA. - Confirm that the master item and every dependent item are supported and have current timestamps.
- Confirm that the normal stable values listed above are present.
- Open Monitoring → Hosts → Graphs for Wort and inspect KitsNet NAS HA Lease.
- Confirm that no new HA problem is open under Monitoring → Problems.
If the master item has a value but dependent items do not, inspect preprocessing errors on the dependent items. If the master item is unsupported, test the user parameter directly on Wort before changing the template.
1.12 Operational interpretation[edit | edit source]
1.12.1 Normal owner heartbeat[edit | edit source]
The guest heartbeat normally refreshes Wort approximately every two seconds. A typical observed lease age is therefore a small number of seconds. Occasional modest increases are not themselves a failover event.
The validated load-and-delay test produced 15 successful refreshes and a maximum lease age of 8.662 seconds under simultaneous CPU pressure, memory pressure, NFS I/O, a competing MASTER request, and 500 ms per-packet SSH delay. The active lease limit remained 30 seconds, Fox was not fenced, and Anchor did not acquire the disk.
1.12.2 What approaching failover looks like[edit | edit source]
A likely developing owner-path problem normally appears in this order:
- Lease age rises above its usual range.
- Lease headroom falls by the same amount.
- Scheduler delay, collector duration, guest service alarms, or network alarms may provide an explanation.
- The 10-second lease-age Warning opens if the condition persists.
- The 20-second High alarm opens if the lease continues aging.
- The 25-second Disaster alarm indicates less than five seconds before the active lease limit.
- Lease validity becomes false after expiry.
Lease expiry by itself does not attach the disk to the peer. The fencing path is exercised when a competing eligible MASTER requests takeover and the complete authority, lease, VM, and attachment checks permit it.
1.12.3 Successful failover[edit | edit source]
During an orderly or automatic handoff, a brief coherence code 1 and a brief zero live-disk count can be normal. The alert persistence periods allow those transitions to complete without immediately opening a problem.
After completion, all of the following should change coherently:
- authority owner;
- authority generation;
- owner MASTER intent;
- peer BACKUP intent;
- live disk owner; and
- coherence state.
The final state should be HEALTHY_ANCHOR or HEALTHY_FOX with a freshly renewed lease.
1.13 Alert response[edit | edit source]
| Alert | First response | Prohibited shortcut |
|---|---|---|
| Lease age Warning or High | Compare lease age with vCPU delay, guest reachability, heartbeat service status, Wort load, and SSH latency. Preserve evidence before intervening. | Do not force a failover solely because lease age briefly increased. |
| Lease within five seconds of expiry | Treat as urgent. Determine whether the owner guest is running and whether its heartbeat/control path has stalled. Watch authority and attachment state continuously. | Do not manually attach the production LV to the peer. |
| Coherence inconsistent | Compare owner, generation, intents, lease validity, domain states, and live owner from the same JSON snapshot. | Do not restart both NAS nodes together. |
| Unsafe attachment | Stop ordinary recovery activity and inspect live and inactive XML on both domains. Protect data integrity first. | Do not mount or start NFS on either node until exclusive ownership is proven. |
| Fencing control invalid | Inspect file existence, regular-file type, mode 0600, root ownership, and content 1. Determine whether disablement was intentional maintenance.
|
Do not recreate the flag during a controlled rebuild or shutdown procedure that requires fencing disabled. |
| NAS domain not running | Determine whether it was deliberately fenced, administratively stopped, or failed. Confirm where the disk and authority ended. | Do not restart a destroyed former owner until storage ownership is known. |
| Collector failure or no data | Check Zabbix Agent 2, the user parameter, sudo validation, SELinux denials, and libvirt query access. | Do not weaken the sudo rule to unrestricted commands. |
1.14 Wort major-version rebuild[edit | edit source]
Monitoring restoration occurs only after the base hypervisor, libvirt configuration, Stage 4 r4 code, authority-state recovery, and controlled HA acceptance are complete.
1.14.1 Rebuild sequence[edit | edit source]
- Keep production fencing disabled during the controlled rebuild and recovery procedure, as required by the main NAS HA rebuild runbook.
- Reinstall Zabbix Agent 2 from the approved Zabbix 7.0 repository.
- Restore the libvirt domains, network, pools, and live-attachment-only storage design.
- Restore the Stage 4 r4 Wort scripts and arbiter state according to the NAS HA rebuild procedure.
- Complete the fencing-disabled HA acceptance and establish a coherent owner.
- Restore the monitoring collector from Git by running
src/nas-ha-monitoring/install-monitoring-on-wort.shonmgr1aspsmode. - Confirm that the collector works through Zabbix Agent 2.
- Re-enable production fencing only through the separately validated HA enablement procedure.
- Confirm
fencing_control_valid=1andcoherence_code=0. - Import the template only if the Zabbix server configuration was also lost or rebuilt. Ordinarily the existing template remains intact.
- Verify Latest data, graphs, and Problems.
The monitoring installer expects the production HA system to be coherent and fencing to be valid at its final acceptance check. Therefore, during a rebuild, run it after the HA solution has returned to its final accepted operational state.
1.15 Zabbix server or template recovery[edit | edit source]
If the Zabbix server loses the template:
- Obtain
configs/nas/KitsNet-NAS-Hypervisor-zabbix-7.0.yamlfrom commit1d1904d1ca9eedfa076358928fdf61999b7bc7a5or tagnas-ha-monitoring-r1. - Verify its SHA-256 hash.
- Import with Create new and Update existing enabled.
- Keep Delete missing disabled unless a separately reviewed change explicitly requires deletion.
- Link KitsNet NAS Hypervisor to Wort if the linkage was not restored with the Zabbix database.
- Allow at least 30 seconds for initial collection and dependent-item processing.
- Perform the Zabbix verification procedure above.
1.16 Security and secrets[edit | edit source]
No private SSH key, password, keytab, Zabbix TLS secret, or other private credential is stored in the monitoring release.
The Zabbix account receives only this additional sudo permission:
/usr/local/sbin/kitsnet-nas-ha-status
The collector has no command-line arguments, rejects no user-supplied domain name because none is accepted, and uses fixed state paths, domain names, and production source paths. This prevents the Zabbix item key from being used as a general libvirt or shell-command interface.
The deployment controller follows KitsNet SSH policy:
- connections originate on
mgr1aspsmode; - the working identity comes from the SSH agent;
IdentitiesOnly=yesis not used; and- privilege escalation occurs through sudo on Wort.
1.17 Known scope boundaries[edit | edit source]
The r1 collector measures the authoritative Wort-side HA control and attachment state. It does not directly query inside Fox or Anchor for:
- Keepalived process state or advertisement age;
- NFS service state or client latency;
- Veeam service state;
- guest filesystem mount state;
- guest heartbeat unit restart count; or
- guest network packet loss.
Those guest and service metrics should remain on the Fox and Anchor host templates and can be presented beside the Wort HA metrics in a Zabbix dashboard. Wort is deliberately not given broad remote credentials merely to collect them.
1.18 Change control[edit | edit source]
Changes should begin in src/nas-ha-monitoring/. After testing, update the deployment mirrors in scripts/, configs/nas/, and configs/sudoers/. Validate exact byte equality between canonical files, mirrors, and deployed files before committing.
Stage only the monitoring paths. The repository may contain unrelated modified or untracked work that must not enter a monitoring commit.
For every release:
- validate Python and shell syntax;
- validate the sudoers fragment with
visudo -cf; - import and validate the Zabbix YAML against the deployed Zabbix major version;
- test the collector as the Zabbix account;
- verify healthy HA state without initiating a failover;
- review the staged path list and staged diff;
- commit;
- create an annotated monitoring release tag;
- push both the branch and tag; and
- verify the remote branch and peeled tag commit IDs.
1.19 References[edit | edit source]
- Zabbix 7.0 dependent items: https://www.zabbix.com/documentation/7.0/en/manual/config/items/itemtypes/dependent_items
- Zabbix 7.0 template export and import: https://www.zabbix.com/documentation/7.0/en/manual/xml_export_import/templates
- Git repository:
/srv/git/KNdkronmgr1 - Monitoring release:
nas-ha-monitoring-r1 - Stage 4 r4 release tag:
nas-ha-stage4-r4-production-r2