| (3 intermediate revisions by the same user not shown) | |||
| Line 245: | Line 245: | ||
<syntaxhighlight lang="bash"> | <syntaxhighlight lang="bash"> | ||
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off | sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off | ||
</syntaxhighlight> | </syntaxhighlight> | ||
| Line 252: | Line 251: | ||
<syntaxhighlight lang="bash"> | <syntaxhighlight lang="bash"> | ||
sudo /usr/local/sbin/kitsnet-vm-startup-control disable-next-boot | sudo /usr/local/sbin/kitsnet-vm-startup-control disable-next-boot | ||
</syntaxhighlight> | </syntaxhighlight> | ||
| Line 259: | Line 257: | ||
<syntaxhighlight lang="bash"> | <syntaxhighlight lang="bash"> | ||
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot manual | sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot manual | ||
</syntaxhighlight> | </syntaxhighlight> | ||
| Line 266: | Line 263: | ||
<syntaxhighlight lang="bash"> | <syntaxhighlight lang="bash"> | ||
sudo /usr/local/sbin/kitsnet-vm-startup-control manual-next-boot | sudo /usr/local/sbin/kitsnet-vm-startup-control manual-next-boot | ||
</syntaxhighlight> | </syntaxhighlight> | ||
| Line 488: | Line 484: | ||
<syntaxhighlight lang="bash"> | <syntaxhighlight lang="bash"> | ||
sudo /usr/local/sbin/kitsnet-vm-startup-control release-cold-start | sudo /usr/local/sbin/kitsnet-vm-startup-control release-cold-start | ||
sudo /usr/local/sbin/kitsnet-cold-start | sudo /usr/local/sbin/kitsnet-cold-start | ||
| Line 534: | Line 529: | ||
<syntaxhighlight lang="bash"> | <syntaxhighlight lang="bash"> | ||
sudo /usr/local/sbin/kitsnet-vm-startup-control status | sudo /usr/local/sbin/kitsnet-vm-startup-control status | ||
</syntaxhighlight> | </syntaxhighlight> | ||
| Line 541: | Line 535: | ||
<syntaxhighlight lang="bash"> | <syntaxhighlight lang="bash"> | ||
systemctl status kitsnet-vm-auto-start.service --no-pager -l | systemctl status kitsnet-vm-auto-start.service --no-pager -l | ||
sudo journalctl -b -u kitsnet-vm-auto-start.service --no-pager | sudo journalctl -b -u kitsnet-vm-auto-start.service --no-pager | ||
| Line 549: | Line 542: | ||
<syntaxhighlight lang="bash"> | <syntaxhighlight lang="bash"> | ||
sudo /usr/local/sbin/kitsnet-nas-ha-status | sudo /usr/local/sbin/kitsnet-nas-ha-status | ||
</syntaxhighlight> | </syntaxhighlight> | ||
| Line 556: | Line 548: | ||
<syntaxhighlight lang="bash"> | <syntaxhighlight lang="bash"> | ||
echo "running_guests=$(sudo virsh list --state-running --name | sed '/^[[:space:]]*$/d' | wc -l)" | echo "running_guests=$(sudo virsh list --state-running --name | sed '/^[[:space:]]*$/d' | wc -l)" | ||
echo "autostart_links=$(sudo find /etc/libvirt/qemu/autostart -maxdepth 1 -type l 2>/dev/null | wc -l)" | echo "autostart_links=$(sudo find /etc/libvirt/qemu/autostart -maxdepth 1 -type l 2>/dev/null | wc -l)" | ||
| Line 575: | Line 566: | ||
<syntaxhighlight lang="bash"> | <syntaxhighlight lang="bash"> | ||
systemctl status kitsnet-vm-startup-complete.target --no-pager | systemctl status kitsnet-vm-startup-complete.target --no-pager | ||
cat /proc/sys/kernel/random/boot_id | cat /proc/sys/kernel/random/boot_id | ||
Latest revision as of 13:47, 11 September 2026
Status: Production implementation, revision 4, 10 September 2026
Production Git revision: 751d6af47385794e6c1772eef761f9aaf3370ee8 — Automate Wort controlled VM startup and shutdown
Purpose: Define and operate the production Wort VM shutdown/startup lifecycle so that NAS ownership, Docker Swarm dependencies, ordinary libvirt managed-save recovery, and remote-console listeners recover automatically and safely after an orderly Wort reboot or a true cold start.
1 Scope[edit | edit source]
Wort is the KVM/libvirt hypervisor hosting the KitsNet infrastructure VMs.
Dependency-sensitive infrastructure:
uv059—fox, preferred KitsNet HA NAS VMuv060—anchor, supported HA NAS fallback VMuv061—mgr1, Docker Swarm manageruv062—wrk1, Docker Swarm workeruv063—wrk2, Docker Swarm worker
Ordinary autostart/managed-save guests:
av001av002uv045uv047uv048uv049uv050uv052uv053uv054uv055uv056uv057uv058uv061uv062uv063
The expected ordinary libvirt autostart count is therefore 17.
The NAS pair is deliberately excluded from ordinary libvirt autostart and managed-save:
uv059Fox —Autostart: disableuv060Anchor —Autostart: disable
The controlled Wort startup implementation owns NAS VM startup exclusively.
2 Governing principles[edit | edit source]
The production policy is:
Preserve state for ordinary guests. Treat the NAS pair specially. Prefer Fox, require a healthy NAS, tolerate a healthy Anchor, never automatically fail back from a healthy Anchor, and never permit dependent infrastructure to outrun its prerequisites.
The Wort controller does not implement a second NAS failover mechanism. It relies on the existing NAS HA authority, lease, fencing, generation, attachment, VRRP, and transition mechanisms.
The controller must never:
- attach or detach the production LV itself;
- invoke the NAS transition script directly;
- manipulate the NFS VIP directly;
- mount the production XFS filesystem itself;
- remove
nopreempt; - force an Anchor-to-Fox failback; or
- treat VRRP MASTER state alone as proof of safe NAS service.
3 NAS HA behavior used by Wort[edit | edit source]
The production NAS rules are:
- Fox is preferred.
- Anchor is a fully supported fallback.
- Both Keepalived nodes have initial state
BACKUP. - Fox VRRP priority is 150.
- Anchor VRRP priority is 100.
- Both nodes use
nopreempt. - If both nodes participate in a fresh election, Fox is expected to win.
- If Fox is unavailable, Anchor may become MASTER and acquire production storage through the guarded Wort authority path.
- If Fox later returns while Anchor remains healthy MASTER, Fox remains BACKUP and diskless.
- Returning service from healthy Anchor to Fox is always an explicit planned failback operation.
- Wort remains the exclusive storage authority.
Accepted coherent production outcomes are:
3.1 Normal NAS state[edit | edit source]
authority_owner = fox coherence_state = HEALTHY_FOX fox intent = MASTER anchor intent = BACKUP live disk owner = fox
The environment label is NORMAL_NAS_FOX.
3.2 Degraded but operational NAS state[edit | edit source]
authority_owner = anchor coherence_state = HEALTHY_ANCHOR anchor intent = MASTER fox = unavailable or BACKUP live disk owner = anchor
The environment label is DEGRADED_NAS_ANCHOR.
In HEALTHY_ANCHOR:
- production NFS/container storage is allowed to serve;
- Docker Swarm startup may continue;
- ordinary guests may recover after their normal dependency gates;
- Veeam remains unavailable because it is intentionally tied to Fox; and
- Fox must not automatically reclaim service.
4 Production shutdown lifecycle[edit | edit source]
An orderly Wort shutdown or reboot uses a mixed shutdown model.
4.1 Ordinary guests[edit | edit source]
The 17 ordinary autostart guests use the customized libvirt managed-save implementation:
/usr/local/sbin/libvirt-guests-parallel.sh
Managed saves are bounded in parallel according to the configured PARALLEL_SUSPEND value.
4.2 NAS guests[edit | edit source]
Fox and Anchor are excluded from managed-save.
After all ordinary managed saves complete successfully, the wrapper invokes:
/usr/local/sbin/kitsnet-nas-host-shutdown
The helper determines the current production storage owner.
It then shuts down:
- the non-owner NAS VM first;
- the current storage-owning NAS VM last; and
- waits for production storage ownership and lease state to quiesce.
This ordering avoids deliberately provoking an HA failover during a Wort host shutdown.
For the normal HEALTHY_FOX state, the expected order is:
uv060 (Anchor, non-owner) uv059 (Fox, owner)
For a legitimate HEALTHY_ANCHOR state, the order reverses:
uv059 (Fox, non-owner) uv060 (Anchor, owner)
The helper does not force-destroy a NAS VM if graceful shutdown fails. A failure stops the shutdown helper with an error rather than deliberately bypassing storage safety.
4.3 Why the NAS pair must not be managed-saved[edit | edit source]
During initial automatic-reboot testing, normal libvirt managed-save captured Fox while its live-only production vdb was attached. The saved QEMU state therefore depended on:
/dev/T1/uv059-dkrnfs1
At the next boot, libvirt attempted to restore that saved state before the T1 volume group was active. The restore failed.
This established the production rule:
Fox and Anchor are special lifecycle guests. They are cleanly shut down and cold-started; they are never ordinary managed-save/autostart guests.
5 Production boot protection[edit | edit source]
Every fresh Wort boot begins with virtualization startup protected by:
/etc/systemd/system-generators/kitsnet-vm-killswitch-generator
The generator masks the relevant libvirt/QEMU automatic startup paths until the controlled startup implementation releases them for the current Wort boot.
This prevents:
- ordinary QEMU autostart from racing ahead of NAS;
- libvirt managed-save restore from bypassing dependency sequencing; and
- console listeners from binding against VMs that do not yet exist.
The generator recognizes a current-boot release marker under /run. Because /run is volatile, a new Wort boot always begins protected.
6 Startup modes[edit | edit source]
The supported next-boot modes are:
| Mode | Persistence | Behavior |
|---|---|---|
auto
|
Default | Fully automatic protected startup. No operator command is required after Wort boots. |
off
|
One boot | Leaves virtualization protected and starts no VMs automatically. |
manual
|
One boot | Leaves virtualization protected and waits for the operator to execute the controlled release/start/finish sequence. |
off and manual are one-shot requests. After they are consumed by a boot, the default returns to auto unless another request is set.
The absence of a next-boot request file represents normal auto mode.
7 Startup-mode commands[edit | edit source]
7.1 Show current and next-boot state[edit | edit source]
sudo /usr/local/sbin/kitsnet-vm-startup-control status
7.2 Normal automatic startup on the next boot[edit | edit source]
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot auto
Compatibility alias:
sudo /usr/local/sbin/kitsnet-vm-startup-control enable-next-boot
7.3 Keep all VMs off on the next boot[edit | edit source]
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
Compatibility alias:
sudo /usr/local/sbin/kitsnet-vm-startup-control disable-next-boot
7.4 Manual controlled startup on the next boot[edit | edit source]
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot manual
Compatibility alias:
sudo /usr/local/sbin/kitsnet-vm-startup-control manual-next-boot
8 Automatic startup service[edit | edit source]
Normal startup is executed by:
/etc/systemd/system/kitsnet-vm-auto-start.service /usr/local/sbin/kitsnet-vm-auto-start
The service is:
Type=oneshot;- enabled under
multi-user.target; - configured with
TimeoutStartSec=0; and - governed by the controller's own bounded readiness timeouts rather than a generic systemd service timeout.
A normal auto boot requires no operator command.
9 Host storage readiness gate[edit | edit source]
Automatic startup does not release virtualization until the shared production LV path exists as a block device:
/dev/T1/uv059-dkrnfs1
The controller waits for the device before calling release-cold-start.
This gate exists because Wort networking and multi-user.target can become available before the T1 volume group finishes activation.
During accepted production testing, the controller began at 23:06:23 and waited 28 seconds before the production LV became ready at 23:06:51.
If the block device does not become ready within the configured bound:
- automatic startup stops;
- virtualization remains protected; and
- Anchor is not started merely because the shared host storage is unavailable.
The absence of the shared host LV is an infrastructure-readiness failure, not evidence that Fox itself has failed.
10 Automatic dependency order[edit | edit source]
The normal production path is:
Wort boot
|
v
generator protects libvirt/QEMU startup
|
v
kitsnet-vm-auto-start.service
|
v
wait for /dev/T1/uv059-dkrnfs1
|
v
release virtualization management
|
v
start Fox
|
+-- preferred NAS state established ----------+
| |
+-- Fox unavailable / preference expires -----+--> allow Anchor path
|
v
HEALTHY_FOX or HEALTHY_ANCHOR
|
v
real NAS/NFS service path ready
|
v
mgr1
|
v
Swarm manager usable
|
v
wrk1 + wrk2
|
v
Swarm workers ready
|
v
recover remaining ordinary autostart guests
|
v
restore libvirt-guests lifecycle
|
v
VM-startup-complete target active
|
v
socat remote consoles start
There is no separate AD-before-Swarm or AD-before-ordinary-VM startup tier in this design.
KNADA domain controllers are ordinary guests for Wort startup purposes. They are important for interactive authentication and directory-dependent applications, but they are not a prerequisite for Docker Swarm infrastructure recovery.
11 Fox preference window[edit | edit source]
Fox is started first.
The production configuration uses a bounded preference opportunity, for example:
NAS_PRIMARY_PREFERENCE_TIMEOUT=180
This is a preference window, not a requirement that Fox must become healthy.
If Fox cannot be started, the controller enables the Anchor path immediately.
If Fox starts but does not establish the preferred safe state before the preference interval expires, the controller starts Anchor and lets the existing NAS HA implementation determine safe ownership.
The Wort startup controller does not stop Fox merely because the preference window expires.
Startup continues only after the final NAS readiness predicate establishes a coherent accepted state:
HEALTHY_FOX; orHEALTHY_ANCHOR.
12 NAS service and Swarm readiness[edit | edit source]
After a safe NAS owner exists, the controller starts mgr1.
It then validates the actual NAS/NFS service path from the manager before accepting NAS service readiness.
Next it validates that mgr1 is a usable Docker Swarm manager.
Only then are wrk1 and wrk2 started.
Startup does not continue to the ordinary-guest completion phase until the required workers reach the accepted Swarm-ready condition.
13 Ordinary guest recovery[edit | edit source]
After infrastructure recovery succeeds, finish-cold-start restores the 17 ordinary autostart definitions and starts/resumes those guests with bounded parallelism.
The production default uses parallelism 6.
The normal virsh start path is used rather than --force-boot. This preserves ordinary libvirt managed-save resume semantics.
Fox and Anchor are absent from the ordinary autostart set and therefore cannot be accidentally restarted by finish-cold-start after a legitimate Anchor-owned recovery.
Expected ordinary autostart count:
17
Expected total normally running guest count after complete recovery:
19
14 Remote console / socat ordering[edit | edit source]
Remote KVM console listeners use the template:
/etc/systemd/system/socat-kvm@.service
They are not enabled directly under multi-user.target.
The enabled instances are instead attached to:
kitsnet-vm-startup-complete.target
The target is activated by finish-cold-start only after:
- controlled NAS/Swarm startup has succeeded;
- ordinary guest recovery has completed;
- normal autostart definitions have been restored; and
- the
libvirt-guestsmanaged-save lifecycle has been re-enabled.
A volatile marker is written at:
/run/kitsnet-vm-startup/vm-startup-complete
It contains the current Wort boot ID.
The socat template also checks that this marker exists. Because /run is cleared at each boot, a previous boot cannot satisfy the condition accidentally.
14.1 Why this ordering exists[edit | edit source]
Before the completion-target integration, socat instances attempted to start under multi-user.target while libvirt was still deliberately masked by the startup protection.
In the observed failure:
- socat attempted startup at 22:50:22;
- the controlled VM startup service did not begin until 22:50:36; and
- VM recovery did not finish until 22:54:08.
The socat units consequently failed and attempted restarts against masked libvirt.
After the completion-target change, accepted testing showed:
kitsnet-vm-startup-complete.targetreached at 23:10:15;- socat startup began only afterward; and
- all 16 configured socat instances reached
activeautomatically.
No manual socat restart is required during a normal boot.
15 Manual-mode recovery[edit | edit source]
If the next boot was deliberately set to manual, Wort remains protected.
After validating the host, execute:
sudo /usr/local/sbin/kitsnet-vm-startup-control release-cold-start
sudo /usr/local/sbin/kitsnet-cold-start
sudo /usr/local/sbin/kitsnet-vm-startup-control finish-cold-start
The final finish-cold-start operation also publishes the current-boot completion marker and activates kitsnet-vm-startup-complete.target, which starts the configured socat listeners.
Do not start the socat instances manually before the VM startup completion target.
16 Off-mode behavior[edit | edit source]
If the next boot was set to off:
- the generator protects libvirt/QEMU automatic startup;
- the automatic controller records the mode;
- no VMs are released or started;
- the VM-startup-complete target is not activated; and
- socat listeners remain stopped.
The request is consumed for that boot. The next boot returns to the default auto policy unless another mode is explicitly requested.
17 Failure behavior[edit | edit source]
The implementation is fail-safe.
A failure in an earlier dependency gate prevents later VM tiers from starting automatically.
Examples include:
- production host LV unavailable beyond its bounded readiness interval;
- neither NAS VM reaching an accepted coherent HA state;
- ambiguous NAS ownership;
- NAS service path not becoming ready;
- mgr1 failing to become a usable Swarm manager;
- wrk1/wrk2 failing required readiness;
- a required infrastructure VM failing to start; or
- ordinary guest managed-save failure during host shutdown.
The controller does not attempt unsafe storage operations to work around these failures.
18 Operational status commands[edit | edit source]
18.1 Startup controller state[edit | edit source]
sudo /usr/local/sbin/kitsnet-vm-startup-control status
18.2 Automatic startup service[edit | edit source]
systemctl status kitsnet-vm-auto-start.service --no-pager -l
sudo journalctl -b -u kitsnet-vm-auto-start.service --no-pager
18.3 NAS state[edit | edit source]
sudo /usr/local/sbin/kitsnet-nas-ha-status
18.4 VM counts[edit | edit source]
echo "running_guests=$(sudo virsh list --state-running --name | sed '/^[[:space:]]*$/d' | wc -l)"
echo "autostart_links=$(sudo find /etc/libvirt/qemu/autostart -maxdepth 1 -type l 2>/dev/null | wc -l)"
echo "managed_save_files=$(sudo find /var/lib/libvirt/qemu/save -maxdepth 1 -type f 2>/dev/null | wc -l)"
echo "held_autostarts=$(sudo find /var/lib/kitsnet-vm-startup/autostart-hold -maxdepth 1 -type l 2>/dev/null | wc -l)"
Normal complete production state is:
running_guests=19 autostart_links=17 managed_save_files=0 held_autostarts=0
18.5 Completion target and console state[edit | edit source]
systemctl status kitsnet-vm-startup-complete.target --no-pager
cat /proc/sys/kernel/random/boot_id
sudo cat /run/kitsnet-vm-startup/vm-startup-complete
systemctl list-units 'socat-kvm@*.service' --no-pager
The completion marker boot ID must match the current Wort boot ID.
19 Software layout[edit | edit source]
| Installed path | Purpose |
|---|---|
/usr/local/sbin/kitsnet-vm-startup-control
|
Startup-mode, protected-release, ordinary-autostart restoration, and VM-completion control. |
/usr/local/sbin/kitsnet-vm-auto-start
|
Automatic boot-mode dispatcher and host production-LV readiness gate. |
/usr/local/sbin/kitsnet-cold-start
|
Dependency-aware NAS and Docker Swarm infrastructure startup controller. |
/usr/local/sbin/kitsnet-nas-host-shutdown
|
Owner-aware clean shutdown of the NAS pair during Wort shutdown. |
/usr/local/sbin/libvirt-guests-parallel.sh
|
KitsNet-managed libvirt guest suspend/resume wrapper; preserves parallel managed-save for ordinary guests and excludes the NAS pair. |
/etc/systemd/system-generators/kitsnet-vm-killswitch-generator
|
Protects every fresh Wort boot from uncontrolled libvirt/QEMU guest activation. |
/etc/systemd/system/kitsnet-vm-auto-start.service
|
Runs fully automatic controlled startup in normal auto mode.
|
/etc/systemd/system/kitsnet-vm-startup-complete.target
|
Explicit post-VM recovery completion point used to release remote-console listeners. |
/etc/systemd/system/socat-kvm@.service
|
Socat KVM remote-console template, ordered after VM startup completion. |
/etc/systemd/system/libvirt-guests.service.d/override.conf
|
Redirects the vendor libvirt-guests service to the KitsNet-managed parallel wrapper. |
/etc/kitsnet/cold-start.conf
|
Production site-specific startup predicates, VM identities, and bounded timeouts. |
/var/lib/kitsnet-vm-startup/autostart-hold/
|
Persistent temporary holding area for ordinary libvirt autostart links during a released but unfinished controlled startup. |
/run/kitsnet-vm-startup/
|
Current-boot release/mode/completion state. |
20 Persistent autostart hold[edit | edit source]
Ordinary QEMU autostart links are held under:
/var/lib/kitsnet-vm-startup/autostart-hold
The hold is persistent rather than being located only under /run.
This ensures that an interrupted controlled-start operation cannot lose track of the held autostart definitions merely because Wort reboots again before finish-cold-start succeeds.
The hold directory should be empty/absent after a successful complete startup.
21 Git source of record[edit | edit source]
Production source is stored in the KNdkr repository under:
src/wort-vm-coldstart/
Accepted production revision:
751d6af47385794e6c1772eef761f9aaf3370ee8 Automate Wort controlled VM startup and shutdown 2026-09-10 23:19:25 -0400
The committed tree includes the production copies of the relevant scripts, unit files, generator, override, and configuration.
MediaWiki documentation is not stored in the KNdkr Git repository.
22 RPM / DNF ownership and upgrade policy[edit | edit source]
The Wort production audit established that all KitsNet startup-control files listed below are locally managed and are not RPM-owned:
/usr/local/sbin/libvirt-guests-parallel.sh/usr/local/sbin/kitsnet-nas-host-shutdown/usr/local/sbin/kitsnet-vm-auto-start/usr/local/sbin/kitsnet-cold-start/usr/local/sbin/kitsnet-vm-startup-control/etc/systemd/system/kitsnet-vm-auto-start.service/etc/systemd/system/kitsnet-vm-startup-complete.target/etc/systemd/system/socat-kvm@.service/etc/systemd/system/libvirt-guests.service.d/override.conf/etc/systemd/system-generators/kitsnet-vm-killswitch-generator/etc/kitsnet/cold-start.conf
A normal DNF/RPM update therefore should not overwrite these files.
The vendor files:
/usr/libexec/libvirt-guests.sh /usr/lib/systemd/system/libvirt-guests.service
are owned by:
libvirt-daemon-common
and were verified unchanged during production acceptance.
22.1 Important upgrade-maintenance rule[edit | edit source]
/usr/local/sbin/libvirt-guests-parallel.sh is a KitsNet-customized descendant of the vendor libvirt guest-management script.
A future libvirt package update may update:
/usr/libexec/libvirt-guests.sh
without changing the KitsNet copy.
Therefore:
After any update of the libvirt package owning the vendor guest-management script, compare the updated vendor implementation with the KitsNet wrapper and determine whether upstream fixes must be incorporated into the KNdkr-maintained copy.
This is an upstream-drift risk, not an overwrite risk.
23 Socat instance policy[edit | edit source]
At production acceptance, 16 socat instances are enabled under:
/etc/systemd/system/kitsnet-vm-startup-complete.target.wants/
There must be no socat-kvm@*.service enablement links remaining directly under:
/etc/systemd/system/multi-user.target.wants/
The console template retains Restart=on-failure, but correct startup ordering means it should no longer enter a retry storm simply because virtualization is still intentionally protected.
24 Production validation history[edit | edit source]
24.1 Test 2 — protected managed-save reboot[edit | edit source]
Result: PASS
Validated the protected reboot mechanics, managed-save recovery, manual controlled release, dependency-sensitive infrastructure startup, bounded parallel ordinary-guest recovery, and preservation of intentionally-off guests.
24.2 Test 3 — controlled true cold start[edit | edit source]
Result: PASS — 10 September 2026
Git baseline:
8899115d79c228fb6578128bced15e0081722697 Stage Wort VM cold-start controller Test 3 fixes
Validated clean Swarm quiescence, a true guest cold start, protected Wort reboot, Fox-first NAS startup, HEALTHY_FOX, mgr1/NFS readiness, Swarm readiness, bounded parallel ordinary-guest startup, and preservation of intentionally-off guests.
24.3 Initial fully automatic reboot — defects found safely[edit | edit source]
The first fully automatic reboot exposed two lifecycle races:
- Wort attempted Fox before
/dev/T1/uv059-dkrnfs1was active. - Fox had been managed-saved while its live-only production
vdbwas attached.
The HA implementation stopped safely without granting ambiguous storage ownership.
These findings produced the host-LV readiness gate and the permanent NAS clean-shutdown/no-managed-save policy.
During the same investigation, Anchor attempted a MASTER transition while Wort itself was disappearing during shutdown. Remote hypervisor publication timed out; storage ownership was refused; the resulting state was TRANSITION_NO_OWNER with no live or persistent production vdb, no valid lease, and no owner.
24.4 Fully automatic VM recovery — socat ordering defect found[edit | edit source]
After the NAS lifecycle fixes, automatic VM recovery succeeded, but socat units still started directly from multi-user.target. They attempted startup before the controlled VM process and failed against deliberately masked libvirt.
This produced kitsnet-vm-startup-complete.target and the current-boot completion marker.
24.5 Test 4 — fully automatic controlled boot and console recovery[edit | edit source]
Result: PASS — 10 September 2026
Production Git revision subsequently frozen as:
751d6af47385794e6c1772eef761f9aaf3370ee8 Automate Wort controlled VM startup and shutdown
Validated without manual VM or socat intervention:
- protected fresh boot;
- automatic mode;
- 28-second production-LV readiness wait;
- Fox-first startup and preferred safe ownership;
- Anchor joining as BACKUP;
- final
HEALTHY_FOX; - mgr1/NFS readiness;
- Docker Swarm manager and worker readiness;
- bounded parallel ordinary-guest recovery;
- VM-startup-complete target activation only after VM recovery;
- all 16 socat listeners starting afterward and reaching
active; - 19 running guests;
- 17 ordinary autostart links;
- 0 managed-save files;
- 0 held autostarts;
- exactly one live production
vdbon Fox; - 0 persistent production
vdbattachments; and - completion-marker boot ID matching the current Wort boot ID.
25 Journal-retention note[edit | edit source]
At final acceptance Wort did not retain the previous boot's systemd journal persistently. Therefore journalctl -b -1 could not provide post-reboot documentary evidence of the immediately preceding NAS-aware shutdown sequence.
This does not alter the operating design. Shutdown-side historical evidence requires either capture before reboot completes or separate configuration of persistent journald storage.
26 Production acceptance state[edit | edit source]
The accepted normal post-boot state is:
VM startup service active (exited), SUCCESS current boot mode auto running guests 19 ordinary autostart links 17 managed-save files 0 held autostart links 0 NAS coherence HEALTHY_FOX normally NAS production live vdb exactly 1 NAS persistent production vdb 0 VM-startup-complete target active socat listeners 16 active completion marker boot ID matches current boot ID
A legitimate HEALTHY_ANCHOR final NAS state is also accepted and must not trigger automatic failback.
27 Troubleshooting principles[edit | edit source]
When automatic startup stops:
- Do not manually attach/detach the production LV.
- Do not manually move the NFS VIP.
- Do not bypass the NAS authority mechanism.
- Inspect
kitsnet-vm-auto-start.serviceand its current-boot journal. - Inspect
/usr/local/sbin/kitsnet-nas-ha-status. - Determine which dependency gate did not become ready.
- Preserve the protected state until the failure is understood.
- Use manual controlled release/start/finish only when intentionally operating in manual recovery mode.
When socat listeners are absent after an otherwise successful boot:
- Verify
kitsnet-vm-startup-complete.targetis active. - Verify the completion marker exists and contains the current boot ID.
- Verify the relevant VM is actually running.
- Verify the socat instance is enabled under the completion target rather than
multi-user.target. - Inspect the instance journal before manually restarting it.
28 Operational principle[edit | edit source]
Every Wort boot starts protected. Normal recovery is automatic. Shared storage must be ready before NAS startup. NAS ownership must be coherent before Swarm startup. The NAS pair is cleanly shut down rather than managed-saved. Ordinary guests preserve state where possible. Remote consoles start only after VM recovery is complete. Healthy Anchor service is accepted, and automatic failback to Fox is forbidden.