Last edited 3 weeks ago
by Peter A. Smode

KVM:Cold-Start and Guest Startup Control

Revision as of 17:43, 10 September 2026 by Peter A. Smode (talk | contribs) (Created page with "= KitsNet Wort VM Cold-Start and Guest Startup Control = '''Status:''' Design and deployment package, revision 3, 10 September 2026 '''Purpose:''' Preserve the normal UPS-backed KVM managed-save/restore path while providing a controlled, dependency-aware cold-start path for the KitsNet container infrastructure and a host-level next-boot guest-startup kill switch. == Scope == Wort is the KVM/libvirt hypervisor hosting, among other VMs: * <code>fox</code> — preferre...")
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)

1 KitsNet Wort VM Cold-Start and Guest Startup Control[edit | edit source]

Status: Design and deployment package, revision 3, 10 September 2026

Purpose: Preserve the normal UPS-backed KVM managed-save/restore path while providing a controlled, dependency-aware cold-start path for the KitsNet container infrastructure and a host-level next-boot guest-startup kill switch.

1.1 Scope[edit | edit source]

Wort is the KVM/libvirt hypervisor hosting, among other VMs:

  • fox — preferred KitsNet HA NAS VM
  • anchor — secondary KitsNet HA NAS VM
  • mgr1 — Docker Swarm manager
  • wrk1 — Docker Swarm worker
  • wrk2 — Docker Swarm worker

KitsNet normally preserves running guests across an orderly Wort shutdown by using libvirt managed-save/hibernate behavior. A true cold start is exceptional and is expected mainly after a hypervisor bugcheck, disrupted power-event handling, or planned maintenance in which the guests were deliberately shut down and guest activation was withheld until Wort itself had been validated.

The design keeps those paths separate:

  1. Normal recovery: restore managed-save state and resume the previously running environment.
  2. Controlled cold start: sequence the infrastructure VMs according to their real dependencies.
  3. Maintenance interlock: optionally prevent all automatic guest activation on the next and subsequent Wort boots without affecting guests that are running at the time the interlock is armed.

1.2 Governing NAS HA behavior[edit | edit source]

The cold-start design follows the validated KitsNet NAS HA design; it does not create a second storage-failover mechanism.

The production HA rules relevant to cold start are:

  • Fox is the preferred NAS node.
  • Anchor is a fully supported failover NAS node.
  • Both Keepalived instances have initial state BACKUP.
  • Fox has VRRP priority 150 and Anchor priority 100.
  • Both nodes use nopreempt.
  • If both nodes are available for a fresh election, Fox is expected to become MASTER because of its higher priority.
  • If Fox is unavailable, Anchor may become MASTER and acquire the production storage through the normal guarded Wort authority/fencing path.
  • If Fox later returns while Anchor is healthy MASTER, Fox remains BACKUP and diskless. There is no automatic failback.
  • Returning service to Fox is a planned, explicit, orderly failback operation.
  • VRRP MASTER state alone is never permission to mount the shared XFS filesystem. Wort remains the exclusive storage authority.
  • Production fencing and the generation/lease/boot/requester/peer/attachment/authority guards remain in force during cold start.

Therefore the cold-start policy is:

Prefer Fox. Require a healthy NAS, not a healthy Fox. Never automatically fail back from a healthy Anchor to Fox.

1.3 Cold-start dependency order[edit | edit source]

Wort validated stable
        |
        v
Start Fox alone
        |
        +-- HEALTHY_FOX within preference window --> normal path
        |
        +-- unavailable / no healthy Fox ----------> fallback path
                                                        |
                                                        v
                                                   Start Anchor
                                                        |
                     +----------------------------------+----------------------------------+
                     |                                                                     |
                     v                                                                     v
               HEALTHY_FOX                                                        HEALTHY_ANCHOR
               Fox MASTER                                                        Anchor MASTER
               Anchor BACKUP                                                     Fox unavailable/BACKUP
                     |                                                                     |
                     +----------------------------------+----------------------------------+
                                                        |
                                                        v
                                              NAS service readiness
                                                        |
                                                        v
                                                      mgr1
                                                        |
                                               Swarm manager ready
                                                        |
                                                        v
                                                  wrk1 + wrk2
                                                        |
                                               Swarm workers ready
                                                        |
                                                        v
                                          AD / infrastructure tier
                                                        |
                                                        v
                                            dependent applications

A failure of Fox does not fail the environment startup. A failure to establish either a coherent HEALTHY_FOX or coherent HEALTHY_ANCHOR state does fail the cold-start sequence.

1.4 Fox preference window[edit | edit source]

The controller starts Fox first and gives it a bounded opportunity to establish the normal preferred state. The supplied example uses:

NAS_PRIMARY_PREFERENCE_TIMEOUT=180

The value is site-configurable. It is deliberately a preference window rather than a required-Fox timeout.

If Fox cannot be started, the controller immediately moves to the Anchor fallback path. If Fox starts but does not establish HEALTHY_FOX within the preference window, the controller starts Anchor and allows the existing HA implementation to establish whichever safe owner is possible.

The controller does not stop Fox merely because the preference window expires. The existing HA transition, lease, fencing, and storage-authority mechanisms remain responsible for determining safe ownership.

1.5 Accepted NAS outcomes[edit | edit source]

1.5.1 Normal outcome[edit | edit source]

authority_owner = fox
coherence_state = HEALTHY_FOX
fox intent       = MASTER
anchor            = BACKUP after it joins

Startup continues with environment state NORMAL_NAS_FOX.

1.5.2 Degraded but operational outcome[edit | edit source]

authority_owner = anchor
coherence_state = HEALTHY_ANCHOR
anchor intent    = MASTER
fox              = unavailable or BACKUP

Startup continues with environment state DEGRADED_NAS_ANCHOR.

In this state:

  • NFS/container storage is allowed to serve production.
  • Docker Swarm startup may continue.
  • AD and dependent applications may continue after their normal readiness gates.
  • Veeam remains unavailable because Veeam is intentionally tied to Fox.
  • If Fox later returns it must not automatically reclaim service.
  • A later return to Fox uses the documented controlled Anchor-to-Fox failback procedure.

1.6 Conditions that stop cold start[edit | edit source]

The controller stops rather than continuing when:

  • neither Fox nor Anchor can establish a coherent accepted NAS state;
  • the NAS owner/readiness state becomes ambiguous;
  • the required NAS service-path readiness check fails;
  • mgr1 cannot become a usable Swarm manager;
  • wrk1/wrk2 cannot reach the required Swarm-ready state; or
  • a required VM other than preferred Fox cannot be started.

The cold-start controller must never:

  • attach or detach the production LV itself;
  • invoke the NAS transition script directly;
  • manipulate the NFS VIP directly;
  • remove nopreempt;
  • force Anchor-to-Fox failback;
  • mount the production XFS filesystem itself; or
  • treat Keepalived MASTER state alone as proof of safe NAS service.

1.7 Normal managed-save recovery[edit | edit source]

The controlled cold-start mechanism is not intended for an ordinary UPS-backed shutdown/recovery in which valid guest managed-save state exists.

The normal path remains:

running guests
     |
Wort orderly shutdown
     |
libvirt managed-save / guest hibernation
     |
Wort restart
     |
restore saved running instances

Swarm scheduling state should not be deliberately disturbed merely because the guests were suspended and restored.

1.8 Host-level next-boot guest VM kill switch[edit | edit source]

The package provides:

/usr/local/sbin/kitsnet-vm-startup-control

and a boot-ID-aware systemd generator:

/etc/systemd/system-generators/kitsnet-vm-killswitch-generator

The control is specifically designed to meet this requirement:

Disable all automatic guest activation on the next boot without stopping or changing guests that are running now; later permit guest activation on a subsequent boot without affecting currently running guests.

1.8.1 Arm the next-boot interlock[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control disable-next-boot

The marker records the current Wort boot ID. The generator deliberately ignores that marker during the same boot. Therefore arming the switch:

  • does not stop any running guest;
  • does not suspend any running guest;
  • does not alter the current libvirt session; and
  • remains safe across a same-boot systemctl daemon-reload.

On the next boot, the changed boot ID causes generated masks for the relevant modular and monolithic QEMU/libvirt startup paths and libvirt-guests.service.

1.8.2 Check status[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control status

1.8.3 Permit normal activation on a later boot[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control enable-next-boot

Disarming removes the persistent marker. It intentionally does not start virtualization or guests during a boot in which the generated masks are already active.

1.8.4 Controlled release during a guest-blocked boot[edit | edit source]

After Wort is manually validated:

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control release-cold-start
sudo /usr/local/sbin/kitsnet-cold-start

The release procedure temporarily holds QEMU autostart links before releasing libvirt management. This prevents ordinary guest autostarts from racing ahead of the dependency-sensitive infrastructure sequence.

After the infrastructure/AD/application startup sequence has been completed:

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control finish-cold-start

This restores normal autostart definitions and re-establishes the ordinary libvirt-guests managed-save lifecycle.

1.9 Software layout[edit | edit source]

Installed path Purpose
/usr/local/sbin/kitsnet-vm-startup-control Arm/disarm/status/release/finish controls for the host-level guest startup interlock.
/usr/local/sbin/kitsnet-cold-start Dependency-aware infrastructure VM cold-start controller.
/etc/systemd/system-generators/kitsnet-vm-killswitch-generator Boot-ID-aware next-boot libvirt/QEMU startup blocker.
/etc/kitsnet/cold-start.conf Production site-specific readiness predicates and timeouts.
/etc/kitsnet/cold-start.conf.example Supplied staging/example configuration. Must be validated before copying to the production path.
/etc/kitsnet/vm-startup.disabled Persistent marker present only while the next-boot kill switch is armed.
/run/kitsnet-vm-cold-start/ Volatile state used only during a controlled current-boot release.

1.10 Deployment archive[edit | edit source]

The deployment archive is intentionally passive. Its install.sh only installs files. It does not:

  • arm the guest kill switch;
  • stop or start libvirt;
  • start any VM;
  • create the production /etc/kitsnet/cold-start.conf; or
  • alter NAS fencing or failover state.

The archive preserves the intended installed directory structure:

kitsnet-vm-coldstart/
  install.sh
  README.md
  VERSION
  usr/local/sbin/
    kitsnet-vm-startup-control
    kitsnet-cold-start
  etc/systemd/system-generators/
    kitsnet-vm-killswitch-generator
  etc/kitsnet/
    cold-start.conf.example
  tests/
    test-kill-switch.sh

1.11 Required production validation before activation[edit | edit source]

Before the package is trusted for production cold starts:

  1. Confirm Wort's actual libvirt unit model: modular virtqemud or monolithic libvirtd.
  2. Confirm how libvirt-guests.service is presently implementing managed-save shutdown/restore on Wort.
  3. Validate the exact output format of /usr/local/sbin/kitsnet-nas-ha-status and adjust the supplied NAS predicates if necessary.
  4. Ensure the healthy-Fox predicate proves authoritative Fox ownership and HEALTHY_FOX coherence.
  5. Ensure the healthy-Anchor predicate proves authoritative Anchor ownership and HEALTHY_ANCHOR coherence.
  6. Replace or explicitly validate NAS_SERVICE_READY_COMMAND so the production gate verifies the service path and not only HA role metadata.
  7. Validate the approved SSH/readiness path from Wort to mgr1.
  8. Confirm the exact libvirt domain names for Fox, Anchor, mgr1, wrk1, and wrk2.
  9. Run the supplied offline kill-switch test.
  10. During a controlled maintenance window, exercise the next-boot interlock with noncritical guests before depending on it for unattended recovery.
  11. Test both cold-start outcomes: normal Fox ownership and simulated/unavailable-Fox Anchor ownership, without bypassing the normal NAS HA mechanism.

1.12 Validation performed on the packaged code[edit | edit source]

Before packaging revision 3:

  • all shell programs and the example configuration passed bash -n syntax checks;
  • the systemd generator offline test passed the same-boot, next-boot, and disarmed cases;
  • the example configuration was successfully sourced; and
  • the package was built so installation itself performs no activation.

These checks validate package mechanics only. They are not a substitute for the Wort-specific production acceptance steps above.

1.13 Future AD/application sequencing[edit | edit source]

After mgr1, wrk1, and wrk2 are available, the next layer will sequence containerized Samba AD before applications that depend on AD.

Docker Swarm does not provide a complete arbitrary service dependency/start-priority facility. Therefore the final AD-first release should be implemented explicitly at the manager layer once the exact Samba AD service placement and fixed-IP design are finalized.

The intended dependency direction remains:

Wort -> NAS -> Swarm -> AD -> AD-dependent applications

1.14 Operational principle[edit | edit source]

The overall KitsNet recovery policy is:

Preserve state whenever possible. Orchestrate dependencies only when state preservation is not available. Prefer Fox, tolerate a healthy Anchor, and never trade availability for an automatic failback.