KVM:Cold-Start and Guest Startup Control: Difference between revisions

Peter A. Smode (talk | contribs)
Peter A. Smode (talk | contribs)
 
Line 529: Line 529:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control status
sudo /usr/local/sbin/kitsnet-vm-startup-control status
</syntaxhighlight>
</syntaxhighlight>
Line 536: Line 535:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
systemctl status kitsnet-vm-auto-start.service --no-pager -l
systemctl status kitsnet-vm-auto-start.service --no-pager -l
sudo journalctl -b -u kitsnet-vm-auto-start.service --no-pager
sudo journalctl -b -u kitsnet-vm-auto-start.service --no-pager
Line 544: Line 542:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-nas-ha-status
sudo /usr/local/sbin/kitsnet-nas-ha-status
</syntaxhighlight>
</syntaxhighlight>
Line 551: Line 548:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
echo "running_guests=$(sudo virsh list --state-running --name | sed '/^[[:space:]]*$/d' | wc -l)"
echo "running_guests=$(sudo virsh list --state-running --name | sed '/^[[:space:]]*$/d' | wc -l)"
echo "autostart_links=$(sudo find /etc/libvirt/qemu/autostart -maxdepth 1 -type l 2>/dev/null | wc -l)"
echo "autostart_links=$(sudo find /etc/libvirt/qemu/autostart -maxdepth 1 -type l 2>/dev/null | wc -l)"
Line 570: Line 566:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
systemctl status kitsnet-vm-startup-complete.target --no-pager
systemctl status kitsnet-vm-startup-complete.target --no-pager
cat /proc/sys/kernel/random/boot_id
cat /proc/sys/kernel/random/boot_id

Latest revision as of 13:47, 11 September 2026

Status: Production implementation, revision 4, 10 September 2026

Production Git revision: 751d6af47385794e6c1772eef761f9aaf3370ee8 — Automate Wort controlled VM startup and shutdown

Purpose: Define and operate the production Wort VM shutdown/startup lifecycle so that NAS ownership, Docker Swarm dependencies, ordinary libvirt managed-save recovery, and remote-console listeners recover automatically and safely after an orderly Wort reboot or a true cold start.

1 Scope[edit | edit source]

Wort is the KVM/libvirt hypervisor hosting the KitsNet infrastructure VMs.

Dependency-sensitive infrastructure:

  • uv059 — fox, preferred KitsNet HA NAS VM
  • uv060 — anchor, supported HA NAS fallback VM
  • uv061 — mgr1, Docker Swarm manager
  • uv062 — wrk1, Docker Swarm worker
  • uv063 — wrk2, Docker Swarm worker

Ordinary autostart/managed-save guests:

  • av001
  • av002
  • uv045
  • uv047
  • uv048
  • uv049
  • uv050
  • uv052
  • uv053
  • uv054
  • uv055
  • uv056
  • uv057
  • uv058
  • uv061
  • uv062
  • uv063

The expected ordinary libvirt autostart count is therefore 17.

The NAS pair is deliberately excluded from ordinary libvirt autostart and managed-save:

  • uv059 Fox — Autostart: disable
  • uv060 Anchor — Autostart: disable

The controlled Wort startup implementation owns NAS VM startup exclusively.

2 Governing principles[edit | edit source]

The production policy is:

Preserve state for ordinary guests. Treat the NAS pair specially. Prefer Fox, require a healthy NAS, tolerate a healthy Anchor, never automatically fail back from a healthy Anchor, and never permit dependent infrastructure to outrun its prerequisites.

The Wort controller does not implement a second NAS failover mechanism. It relies on the existing NAS HA authority, lease, fencing, generation, attachment, VRRP, and transition mechanisms.

The controller must never:

  • attach or detach the production LV itself;
  • invoke the NAS transition script directly;
  • manipulate the NFS VIP directly;
  • mount the production XFS filesystem itself;
  • remove nopreempt;
  • force an Anchor-to-Fox failback; or
  • treat VRRP MASTER state alone as proof of safe NAS service.

3 NAS HA behavior used by Wort[edit | edit source]

The production NAS rules are:

  • Fox is preferred.
  • Anchor is a fully supported fallback.
  • Both Keepalived nodes have initial state BACKUP.
  • Fox VRRP priority is 150.
  • Anchor VRRP priority is 100.
  • Both nodes use nopreempt.
  • If both nodes participate in a fresh election, Fox is expected to win.
  • If Fox is unavailable, Anchor may become MASTER and acquire production storage through the guarded Wort authority path.
  • If Fox later returns while Anchor remains healthy MASTER, Fox remains BACKUP and diskless.
  • Returning service from healthy Anchor to Fox is always an explicit planned failback operation.
  • Wort remains the exclusive storage authority.

Accepted coherent production outcomes are:

3.1 Normal NAS state[edit | edit source]

authority_owner = fox
coherence_state = HEALTHY_FOX
fox intent       = MASTER
anchor intent    = BACKUP
live disk owner  = fox

The environment label is NORMAL_NAS_FOX.

3.2 Degraded but operational NAS state[edit | edit source]

authority_owner = anchor
coherence_state = HEALTHY_ANCHOR
anchor intent    = MASTER
fox              = unavailable or BACKUP
live disk owner  = anchor

The environment label is DEGRADED_NAS_ANCHOR.

In HEALTHY_ANCHOR:

  • production NFS/container storage is allowed to serve;
  • Docker Swarm startup may continue;
  • ordinary guests may recover after their normal dependency gates;
  • Veeam remains unavailable because it is intentionally tied to Fox; and
  • Fox must not automatically reclaim service.

4 Production shutdown lifecycle[edit | edit source]

An orderly Wort shutdown or reboot uses a mixed shutdown model.

4.1 Ordinary guests[edit | edit source]

The 17 ordinary autostart guests use the customized libvirt managed-save implementation:

/usr/local/sbin/libvirt-guests-parallel.sh

Managed saves are bounded in parallel according to the configured PARALLEL_SUSPEND value.

4.2 NAS guests[edit | edit source]

Fox and Anchor are excluded from managed-save.

After all ordinary managed saves complete successfully, the wrapper invokes:

/usr/local/sbin/kitsnet-nas-host-shutdown

The helper determines the current production storage owner.

It then shuts down:

  1. the non-owner NAS VM first;
  2. the current storage-owning NAS VM last; and
  3. waits for production storage ownership and lease state to quiesce.

This ordering avoids deliberately provoking an HA failover during a Wort host shutdown.

For the normal HEALTHY_FOX state, the expected order is:

uv060  (Anchor, non-owner)
uv059  (Fox, owner)

For a legitimate HEALTHY_ANCHOR state, the order reverses:

uv059  (Fox, non-owner)
uv060  (Anchor, owner)

The helper does not force-destroy a NAS VM if graceful shutdown fails. A failure stops the shutdown helper with an error rather than deliberately bypassing storage safety.

4.3 Why the NAS pair must not be managed-saved[edit | edit source]

During initial automatic-reboot testing, normal libvirt managed-save captured Fox while its live-only production vdb was attached. The saved QEMU state therefore depended on:

/dev/T1/uv059-dkrnfs1

At the next boot, libvirt attempted to restore that saved state before the T1 volume group was active. The restore failed.

This established the production rule:

Fox and Anchor are special lifecycle guests. They are cleanly shut down and cold-started; they are never ordinary managed-save/autostart guests.

5 Production boot protection[edit | edit source]

Every fresh Wort boot begins with virtualization startup protected by:

/etc/systemd/system-generators/kitsnet-vm-killswitch-generator

The generator masks the relevant libvirt/QEMU automatic startup paths until the controlled startup implementation releases them for the current Wort boot.

This prevents:

  • ordinary QEMU autostart from racing ahead of NAS;
  • libvirt managed-save restore from bypassing dependency sequencing; and
  • console listeners from binding against VMs that do not yet exist.

The generator recognizes a current-boot release marker under /run. Because /run is volatile, a new Wort boot always begins protected.

6 Startup modes[edit | edit source]

The supported next-boot modes are:

Mode Persistence Behavior
auto Default Fully automatic protected startup. No operator command is required after Wort boots.
off One boot Leaves virtualization protected and starts no VMs automatically.
manual One boot Leaves virtualization protected and waits for the operator to execute the controlled release/start/finish sequence.

off and manual are one-shot requests. After they are consumed by a boot, the default returns to auto unless another request is set.

The absence of a next-boot request file represents normal auto mode.

7 Startup-mode commands[edit | edit source]

7.1 Show current and next-boot state[edit | edit source]

sudo /usr/local/sbin/kitsnet-vm-startup-control status

7.2 Normal automatic startup on the next boot[edit | edit source]

sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot auto

Compatibility alias:

sudo /usr/local/sbin/kitsnet-vm-startup-control enable-next-boot

7.3 Keep all VMs off on the next boot[edit | edit source]

sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off

Compatibility alias:

sudo /usr/local/sbin/kitsnet-vm-startup-control disable-next-boot

7.4 Manual controlled startup on the next boot[edit | edit source]

sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot manual

Compatibility alias:

sudo /usr/local/sbin/kitsnet-vm-startup-control manual-next-boot

8 Automatic startup service[edit | edit source]

Normal startup is executed by:

/etc/systemd/system/kitsnet-vm-auto-start.service
/usr/local/sbin/kitsnet-vm-auto-start

The service is:

  • Type=oneshot;
  • enabled under multi-user.target;
  • configured with TimeoutStartSec=0; and
  • governed by the controller's own bounded readiness timeouts rather than a generic systemd service timeout.

A normal auto boot requires no operator command.

9 Host storage readiness gate[edit | edit source]

Automatic startup does not release virtualization until the shared production LV path exists as a block device:

/dev/T1/uv059-dkrnfs1

The controller waits for the device before calling release-cold-start.

This gate exists because Wort networking and multi-user.target can become available before the T1 volume group finishes activation.

During accepted production testing, the controller began at 23:06:23 and waited 28 seconds before the production LV became ready at 23:06:51.

If the block device does not become ready within the configured bound:

  • automatic startup stops;
  • virtualization remains protected; and
  • Anchor is not started merely because the shared host storage is unavailable.

The absence of the shared host LV is an infrastructure-readiness failure, not evidence that Fox itself has failed.

10 Automatic dependency order[edit | edit source]

The normal production path is:

Wort boot
   |
   v
generator protects libvirt/QEMU startup
   |
   v
kitsnet-vm-auto-start.service
   |
   v
wait for /dev/T1/uv059-dkrnfs1
   |
   v
release virtualization management
   |
   v
start Fox
   |
   +-- preferred NAS state established ----------+
   |                                             |
   +-- Fox unavailable / preference expires -----+--> allow Anchor path
                                                 |
                                                 v
                                      HEALTHY_FOX or HEALTHY_ANCHOR
                                                 |
                                                 v
                                      real NAS/NFS service path ready
                                                 |
                                                 v
                                              mgr1
                                                 |
                                                 v
                                      Swarm manager usable
                                                 |
                                                 v
                                           wrk1 + wrk2
                                                 |
                                                 v
                                      Swarm workers ready
                                                 |
                                                 v
                              recover remaining ordinary autostart guests
                                                 |
                                                 v
                                  restore libvirt-guests lifecycle
                                                 |
                                                 v
                               VM-startup-complete target active
                                                 |
                                                 v
                                   socat remote consoles start

There is no separate AD-before-Swarm or AD-before-ordinary-VM startup tier in this design.

KNADA domain controllers are ordinary guests for Wort startup purposes. They are important for interactive authentication and directory-dependent applications, but they are not a prerequisite for Docker Swarm infrastructure recovery.

11 Fox preference window[edit | edit source]

Fox is started first.

The production configuration uses a bounded preference opportunity, for example:

NAS_PRIMARY_PREFERENCE_TIMEOUT=180

This is a preference window, not a requirement that Fox must become healthy.

If Fox cannot be started, the controller enables the Anchor path immediately.

If Fox starts but does not establish the preferred safe state before the preference interval expires, the controller starts Anchor and lets the existing NAS HA implementation determine safe ownership.

The Wort startup controller does not stop Fox merely because the preference window expires.

Startup continues only after the final NAS readiness predicate establishes a coherent accepted state:

  • HEALTHY_FOX; or
  • HEALTHY_ANCHOR.

12 NAS service and Swarm readiness[edit | edit source]

After a safe NAS owner exists, the controller starts mgr1.

It then validates the actual NAS/NFS service path from the manager before accepting NAS service readiness.

Next it validates that mgr1 is a usable Docker Swarm manager.

Only then are wrk1 and wrk2 started.

Startup does not continue to the ordinary-guest completion phase until the required workers reach the accepted Swarm-ready condition.

13 Ordinary guest recovery[edit | edit source]

After infrastructure recovery succeeds, finish-cold-start restores the 17 ordinary autostart definitions and starts/resumes those guests with bounded parallelism.

The production default uses parallelism 6.

The normal virsh start path is used rather than --force-boot. This preserves ordinary libvirt managed-save resume semantics.

Fox and Anchor are absent from the ordinary autostart set and therefore cannot be accidentally restarted by finish-cold-start after a legitimate Anchor-owned recovery.

Expected ordinary autostart count:

17

Expected total normally running guest count after complete recovery:

19

14 Remote console / socat ordering[edit | edit source]

Remote KVM console listeners use the template:

/etc/systemd/system/socat-kvm@.service

They are not enabled directly under multi-user.target.

The enabled instances are instead attached to:

kitsnet-vm-startup-complete.target

The target is activated by finish-cold-start only after:

  • controlled NAS/Swarm startup has succeeded;
  • ordinary guest recovery has completed;
  • normal autostart definitions have been restored; and
  • the libvirt-guests managed-save lifecycle has been re-enabled.

A volatile marker is written at:

/run/kitsnet-vm-startup/vm-startup-complete

It contains the current Wort boot ID.

The socat template also checks that this marker exists. Because /run is cleared at each boot, a previous boot cannot satisfy the condition accidentally.

14.1 Why this ordering exists[edit | edit source]

Before the completion-target integration, socat instances attempted to start under multi-user.target while libvirt was still deliberately masked by the startup protection.

In the observed failure:

  • socat attempted startup at 22:50:22;
  • the controlled VM startup service did not begin until 22:50:36; and
  • VM recovery did not finish until 22:54:08.

The socat units consequently failed and attempted restarts against masked libvirt.

After the completion-target change, accepted testing showed:

  • kitsnet-vm-startup-complete.target reached at 23:10:15;
  • socat startup began only afterward; and
  • all 16 configured socat instances reached active automatically.

No manual socat restart is required during a normal boot.

15 Manual-mode recovery[edit | edit source]

If the next boot was deliberately set to manual, Wort remains protected.

After validating the host, execute:

sudo /usr/local/sbin/kitsnet-vm-startup-control release-cold-start
sudo /usr/local/sbin/kitsnet-cold-start
sudo /usr/local/sbin/kitsnet-vm-startup-control finish-cold-start

The final finish-cold-start operation also publishes the current-boot completion marker and activates kitsnet-vm-startup-complete.target, which starts the configured socat listeners.

Do not start the socat instances manually before the VM startup completion target.

16 Off-mode behavior[edit | edit source]

If the next boot was set to off:

  • the generator protects libvirt/QEMU automatic startup;
  • the automatic controller records the mode;
  • no VMs are released or started;
  • the VM-startup-complete target is not activated; and
  • socat listeners remain stopped.

The request is consumed for that boot. The next boot returns to the default auto policy unless another mode is explicitly requested.

17 Failure behavior[edit | edit source]

The implementation is fail-safe.

A failure in an earlier dependency gate prevents later VM tiers from starting automatically.

Examples include:

  • production host LV unavailable beyond its bounded readiness interval;
  • neither NAS VM reaching an accepted coherent HA state;
  • ambiguous NAS ownership;
  • NAS service path not becoming ready;
  • mgr1 failing to become a usable Swarm manager;
  • wrk1/wrk2 failing required readiness;
  • a required infrastructure VM failing to start; or
  • ordinary guest managed-save failure during host shutdown.

The controller does not attempt unsafe storage operations to work around these failures.

18 Operational status commands[edit | edit source]

18.1 Startup controller state[edit | edit source]

sudo /usr/local/sbin/kitsnet-vm-startup-control status

18.2 Automatic startup service[edit | edit source]

systemctl status kitsnet-vm-auto-start.service --no-pager -l
sudo journalctl -b -u kitsnet-vm-auto-start.service --no-pager

18.3 NAS state[edit | edit source]

sudo /usr/local/sbin/kitsnet-nas-ha-status

18.4 VM counts[edit | edit source]

echo "running_guests=$(sudo virsh list --state-running --name | sed '/^[[:space:]]*$/d' | wc -l)"
echo "autostart_links=$(sudo find /etc/libvirt/qemu/autostart -maxdepth 1 -type l 2>/dev/null | wc -l)"
echo "managed_save_files=$(sudo find /var/lib/libvirt/qemu/save -maxdepth 1 -type f 2>/dev/null | wc -l)"
echo "held_autostarts=$(sudo find /var/lib/kitsnet-vm-startup/autostart-hold -maxdepth 1 -type l 2>/dev/null | wc -l)"

Normal complete production state is:

running_guests=19
autostart_links=17
managed_save_files=0
held_autostarts=0

18.5 Completion target and console state[edit | edit source]

systemctl status kitsnet-vm-startup-complete.target --no-pager
cat /proc/sys/kernel/random/boot_id
sudo cat /run/kitsnet-vm-startup/vm-startup-complete
systemctl list-units 'socat-kvm@*.service' --no-pager

The completion marker boot ID must match the current Wort boot ID.

19 Software layout[edit | edit source]

Installed path Purpose
/usr/local/sbin/kitsnet-vm-startup-control Startup-mode, protected-release, ordinary-autostart restoration, and VM-completion control.
/usr/local/sbin/kitsnet-vm-auto-start Automatic boot-mode dispatcher and host production-LV readiness gate.
/usr/local/sbin/kitsnet-cold-start Dependency-aware NAS and Docker Swarm infrastructure startup controller.
/usr/local/sbin/kitsnet-nas-host-shutdown Owner-aware clean shutdown of the NAS pair during Wort shutdown.
/usr/local/sbin/libvirt-guests-parallel.sh KitsNet-managed libvirt guest suspend/resume wrapper; preserves parallel managed-save for ordinary guests and excludes the NAS pair.
/etc/systemd/system-generators/kitsnet-vm-killswitch-generator Protects every fresh Wort boot from uncontrolled libvirt/QEMU guest activation.
/etc/systemd/system/kitsnet-vm-auto-start.service Runs fully automatic controlled startup in normal auto mode.
/etc/systemd/system/kitsnet-vm-startup-complete.target Explicit post-VM recovery completion point used to release remote-console listeners.
/etc/systemd/system/socat-kvm@.service Socat KVM remote-console template, ordered after VM startup completion.
/etc/systemd/system/libvirt-guests.service.d/override.conf Redirects the vendor libvirt-guests service to the KitsNet-managed parallel wrapper.
/etc/kitsnet/cold-start.conf Production site-specific startup predicates, VM identities, and bounded timeouts.
/var/lib/kitsnet-vm-startup/autostart-hold/ Persistent temporary holding area for ordinary libvirt autostart links during a released but unfinished controlled startup.
/run/kitsnet-vm-startup/ Current-boot release/mode/completion state.

20 Persistent autostart hold[edit | edit source]

Ordinary QEMU autostart links are held under:

/var/lib/kitsnet-vm-startup/autostart-hold

The hold is persistent rather than being located only under /run.

This ensures that an interrupted controlled-start operation cannot lose track of the held autostart definitions merely because Wort reboots again before finish-cold-start succeeds.

The hold directory should be empty/absent after a successful complete startup.

21 Git source of record[edit | edit source]

Production source is stored in the KNdkr repository under:

src/wort-vm-coldstart/

Accepted production revision:

751d6af47385794e6c1772eef761f9aaf3370ee8
Automate Wort controlled VM startup and shutdown
2026-09-10 23:19:25 -0400

The committed tree includes the production copies of the relevant scripts, unit files, generator, override, and configuration.

MediaWiki documentation is not stored in the KNdkr Git repository.

22 RPM / DNF ownership and upgrade policy[edit | edit source]

The Wort production audit established that all KitsNet startup-control files listed below are locally managed and are not RPM-owned:

  • /usr/local/sbin/libvirt-guests-parallel.sh
  • /usr/local/sbin/kitsnet-nas-host-shutdown
  • /usr/local/sbin/kitsnet-vm-auto-start
  • /usr/local/sbin/kitsnet-cold-start
  • /usr/local/sbin/kitsnet-vm-startup-control
  • /etc/systemd/system/kitsnet-vm-auto-start.service
  • /etc/systemd/system/kitsnet-vm-startup-complete.target
  • /etc/systemd/system/socat-kvm@.service
  • /etc/systemd/system/libvirt-guests.service.d/override.conf
  • /etc/systemd/system-generators/kitsnet-vm-killswitch-generator
  • /etc/kitsnet/cold-start.conf

A normal DNF/RPM update therefore should not overwrite these files.

The vendor files:

/usr/libexec/libvirt-guests.sh
/usr/lib/systemd/system/libvirt-guests.service

are owned by:

libvirt-daemon-common

and were verified unchanged during production acceptance.

22.1 Important upgrade-maintenance rule[edit | edit source]

/usr/local/sbin/libvirt-guests-parallel.sh is a KitsNet-customized descendant of the vendor libvirt guest-management script.

A future libvirt package update may update:

/usr/libexec/libvirt-guests.sh

without changing the KitsNet copy.

Therefore:

After any update of the libvirt package owning the vendor guest-management script, compare the updated vendor implementation with the KitsNet wrapper and determine whether upstream fixes must be incorporated into the KNdkr-maintained copy.

This is an upstream-drift risk, not an overwrite risk.

23 Socat instance policy[edit | edit source]

At production acceptance, 16 socat instances are enabled under:

/etc/systemd/system/kitsnet-vm-startup-complete.target.wants/

There must be no socat-kvm@*.service enablement links remaining directly under:

/etc/systemd/system/multi-user.target.wants/

The console template retains Restart=on-failure, but correct startup ordering means it should no longer enter a retry storm simply because virtualization is still intentionally protected.

24 Production validation history[edit | edit source]

24.1 Test 2 — protected managed-save reboot[edit | edit source]

Result: PASS

Validated the protected reboot mechanics, managed-save recovery, manual controlled release, dependency-sensitive infrastructure startup, bounded parallel ordinary-guest recovery, and preservation of intentionally-off guests.

24.2 Test 3 — controlled true cold start[edit | edit source]

Result: PASS — 10 September 2026

Git baseline:

8899115d79c228fb6578128bced15e0081722697
Stage Wort VM cold-start controller Test 3 fixes

Validated clean Swarm quiescence, a true guest cold start, protected Wort reboot, Fox-first NAS startup, HEALTHY_FOX, mgr1/NFS readiness, Swarm readiness, bounded parallel ordinary-guest startup, and preservation of intentionally-off guests.

24.3 Initial fully automatic reboot — defects found safely[edit | edit source]

The first fully automatic reboot exposed two lifecycle races:

  • Wort attempted Fox before /dev/T1/uv059-dkrnfs1 was active.
  • Fox had been managed-saved while its live-only production vdb was attached.

The HA implementation stopped safely without granting ambiguous storage ownership.

These findings produced the host-LV readiness gate and the permanent NAS clean-shutdown/no-managed-save policy.

During the same investigation, Anchor attempted a MASTER transition while Wort itself was disappearing during shutdown. Remote hypervisor publication timed out; storage ownership was refused; the resulting state was TRANSITION_NO_OWNER with no live or persistent production vdb, no valid lease, and no owner.

24.4 Fully automatic VM recovery — socat ordering defect found[edit | edit source]

After the NAS lifecycle fixes, automatic VM recovery succeeded, but socat units still started directly from multi-user.target. They attempted startup before the controlled VM process and failed against deliberately masked libvirt.

This produced kitsnet-vm-startup-complete.target and the current-boot completion marker.

24.5 Test 4 — fully automatic controlled boot and console recovery[edit | edit source]

Result: PASS — 10 September 2026

Production Git revision subsequently frozen as:

751d6af47385794e6c1772eef761f9aaf3370ee8
Automate Wort controlled VM startup and shutdown

Validated without manual VM or socat intervention:

  • protected fresh boot;
  • automatic mode;
  • 28-second production-LV readiness wait;
  • Fox-first startup and preferred safe ownership;
  • Anchor joining as BACKUP;
  • final HEALTHY_FOX;
  • mgr1/NFS readiness;
  • Docker Swarm manager and worker readiness;
  • bounded parallel ordinary-guest recovery;
  • VM-startup-complete target activation only after VM recovery;
  • all 16 socat listeners starting afterward and reaching active;
  • 19 running guests;
  • 17 ordinary autostart links;
  • 0 managed-save files;
  • 0 held autostarts;
  • exactly one live production vdb on Fox;
  • 0 persistent production vdb attachments; and
  • completion-marker boot ID matching the current Wort boot ID.

25 Journal-retention note[edit | edit source]

At final acceptance Wort did not retain the previous boot's systemd journal persistently. Therefore journalctl -b -1 could not provide post-reboot documentary evidence of the immediately preceding NAS-aware shutdown sequence.

This does not alter the operating design. Shutdown-side historical evidence requires either capture before reboot completes or separate configuration of persistent journald storage.

26 Production acceptance state[edit | edit source]

The accepted normal post-boot state is:

VM startup service              active (exited), SUCCESS
current boot mode               auto
running guests                  19
ordinary autostart links        17
managed-save files              0
held autostart links            0
NAS coherence                   HEALTHY_FOX normally
NAS production live vdb         exactly 1
NAS persistent production vdb   0
VM-startup-complete target      active
socat listeners                 16 active
completion marker boot ID       matches current boot ID

A legitimate HEALTHY_ANCHOR final NAS state is also accepted and must not trigger automatic failback.

27 Troubleshooting principles[edit | edit source]

When automatic startup stops:

  1. Do not manually attach/detach the production LV.
  2. Do not manually move the NFS VIP.
  3. Do not bypass the NAS authority mechanism.
  4. Inspect kitsnet-vm-auto-start.service and its current-boot journal.
  5. Inspect /usr/local/sbin/kitsnet-nas-ha-status.
  6. Determine which dependency gate did not become ready.
  7. Preserve the protected state until the failure is understood.
  8. Use manual controlled release/start/finish only when intentionally operating in manual recovery mode.

When socat listeners are absent after an otherwise successful boot:

  1. Verify kitsnet-vm-startup-complete.target is active.
  2. Verify the completion marker exists and contains the current boot ID.
  3. Verify the relevant VM is actually running.
  4. Verify the socat instance is enabled under the completion target rather than multi-user.target.
  5. Inspect the instance journal before manually restarting it.

28 Operational principle[edit | edit source]

Every Wort boot starts protected. Normal recovery is automatic. Shared storage must be ready before NAS startup. NAS ownership must be coherent before Swarm startup. The NAS pair is cleanly shut down rather than managed-saved. Ordinary guests preserve state where possible. Remote consoles start only after VM recovery is complete. Healthy Anchor service is accepted, and automatic failback to Fox is forbidden.