KVM:T0 NVMe Replacement Runbook: Difference between revisions

Peter A. Smode (talk | contribs)
Created page with "'''Status:''' Planned maintenance procedure '''Host:''' <code>wort</code> '''Purpose:''' Replace the physical NVMe device containing the <code>T0</code> LVM volume group with a larger NVMe when both devices cannot be installed simultaneously. The migration uses a verified compressed whole-device image stored on a dock-attached disk, followed by restoration to the replacement NVMe, validation of the cloned LVM identity, expansion of the partition and PV, restoration of..."
 
Peter A. Smode (talk | contribs)
mNo edit summary
 
(2 intermediate revisions by the same user not shown)
Line 1: Line 1:
'''Status:''' Planned maintenance procedure
'''Status:''' Completed and production-validated 12 September 2026


'''Host:''' <code>wort</code>
'''Host:''' <code>wort</code>
'''Execution revision:''' Incorporates corrections and lessons learned from the completed Plextor-to-Sabrent T0 migration on 12 September 2026.


'''Purpose:''' Replace the physical NVMe device containing the <code>T0</code> LVM volume group with a larger NVMe when both devices cannot be installed simultaneously. The migration uses a verified compressed whole-device image stored on a dock-attached disk, followed by restoration to the replacement NVMe, validation of the cloned LVM identity, expansion of the partition and PV, restoration of the libvirt managed-save mount, and controlled return of KitsNet virtual machines to service.
'''Purpose:''' Replace the physical NVMe device containing the <code>T0</code> LVM volume group with a larger NVMe when both devices cannot be installed simultaneously. The migration uses a verified compressed whole-device image stored on a dock-attached disk, followed by restoration to the replacement NVMe, validation of the cloned LVM identity, expansion of the partition and PV, restoration of the libvirt managed-save mount, and controlled return of KitsNet virtual machines to service.


'''Depends on:'''
'''Depends on:'''
* [[KitsNet Wort VM Cold-Start and Guest Startup Control]]
* [[KVM:Cold-Start and Guest Startup Control|KitsNet Wort VM Cold-Start and Guest Startup Control]]


== Current storage design ==


At the time this procedure was written:
== Validated execution record and authoritative corrections — 12 September 2026 ==


* <code>T0</code> is a single-PV volume group.
This runbook was executed end-to-end on Wort. The corrections in this section are authoritative where they differ from the original planned steps later in the document. The later procedure has also been patched where practical, but this section records the actual accepted behavior and final state.
* The T0 PV is <code>/dev/nvme0n1p1</code>.
* The underlying whole device is <code>/dev/nvme0n1</code>.
* T0 contains 27 logical volumes used by the Wort virtual-machine environment.
* <code>T0/kvm_save</code> is a 64-GiB LV used for <code>/var/lib/libvirt/qemu/save</code>.
* Wort's host operating-system filesystems are on the separate <code>T0h</code> VG, not on T0.
* T0 can therefore be completely deactivated while Wort itself remains operational.


The migration deliberately clones the '''entire NVMe device''', not merely partition 1.
=== Hardware and preserved identity ===


A whole-device clone preserves the on-disk:
The completed migration replaced:


* partition table;
* source device: <code>PLEXTOR PX-512M8PeG</code>, serial <code>P02709118482</code>, WWN/EUI <code>eui.0023035630011a50</code>;
* disk GUID, if GPT;
* source whole-device size: <code>512110190592</code> bytes;
* partition GUID/PARTUUID;
* replacement device: <code>Sabrent SB-RKT4L-2TB</code>, serial <code>48880885900620</code>, WWN/EUI <code>eui.00000000000000006479a7a11ac00487</code>;
* LVM PV UUID;
* replacement whole-device size: <code>2048408248320</code> bytes; and
* T0 VG UUID;
* logical/physical sector size: <code>512/512</code> bytes on both devices.
* LV UUIDs;
* filesystem UUIDs;
* LVM metadata;
* allocated and unallocated extents; and
* all VM disk and managed-save contents.


The replacement NVMe's physical model, serial number, WWID and other controller-level identifiers are not cloned.
The replacement was installed in Wort's primary NVMe slot. Wort uses an AMD Ryzen 7 5800X processor.


== Governing safety rules ==
The identities preserved through the clone were:


<blockquote>
<pre>
'''Do not improvise past a STOP checkpoint. If a required validation does not produce the documented result, stop the procedure and diagnose the discrepancy before proceeding.'''
GPT disk GUID: 4481283C-CEB0-4E65-BA1D-9997D5708D63
</blockquote>
Partition start: 2048
Partition PARTUUID: 3703AD9C-AB50-486A-B154-E83338BC7D3F
T0 PVID: Xn55sH-OURj-BQdg-GboO-RTia-1HfD-knB8F2
T0 VG UUID: 2Ttv1r-QKPc-hu5W-HDk4-j38t-06wO-WEjFq7
kvm_save filesystem UUID: 28ea4005-788c-44a6-a014-6a10d396dba7
LV count: 27
</pre>


The following rules apply for the entire maintenance:
Final expanded state:
 
<pre>
/dev/nvme0n1p1 start=2048 size=4000795279 sectors
last usable GPT LBA=4000797326
T0 size=1.86 TiB
T0 free=1.50 TiB
</pre>
 
=== Recovery image ===
 
<pre>
/mnt/KitsNet-pbd2/wort-T0-nvme-migration/wort-T0-nvme0n1.full.img.gz
compressed size=139282868502 bytes
SHA256=4329dc57a2ee58c492713108da3c8491e657793c09aa6adfc3c02947bb49b09c
</pre>


# The original NVMe is not erased, modified, reformatted or repurposed after removal until the replacement has passed final production validation.
The image passed <code>pigz -t</code> and a byte-for-byte comparison against the original Plextor before removal. After restoration, the original <code>512110190592</code>-byte source region on the Sabrent compared byte-for-byte identical to the image.
# No <code>pvcreate</code>, <code>vgcreate</code>, <code>vgimportclone</code>, <code>pvchange --uuid</code>, <code>vgchange --uuid</code> or routine <code>vgcfgrestore</code> is used.
# The image is made from the '''whole device''' <code>/dev/nvme0n1</code>.
# The image is restored to the '''whole replacement device'''.
# T0 must be inactive while the source image is created.
# The replacement NVMe must be at least as large in bytes as the source NVMe.
# The replacement NVMe must present the '''same logical sector size''' as the source NVMe.
# The clone is validated before any expansion operation is performed.
# Partition expansion and PV expansion are separate from the cloning operation.
# The <code>/var/lib/libvirt/qemu/save</code> mount remains available until the controlled VM shutdown/managed-save operation is complete.
# The save mount is then unmounted and disabled in <code>/etc/fstab</code> until T0 has been restored and expanded.
# Every maintenance boot remains VM-protected.
# Because <code>next-boot off</code> is one-shot, '''the first administrative action after every protected maintenance boot is to arm <code>next-boot off</code> again for the following boot'''.
# The final production boot is the only boot for which <code>next-boot auto</code> is explicitly selected.


== Migration overview ==
=== Corrected operating rules ===


The controlled sequence is:
# Wort's installed <code>lsblk</code> does not support the <code>START</code> column used in the first draft. Obtain partition start and size from <code>sfdisk --dump</code>.
# The recovery disk's <code>/dev/sdX</code> letter changed across boots. Use the mounted filesystem <code>/mnt/KitsNet-pbd2</code> as the recovery-path authority and identify the physical disk by model/WWN when remounting it. The validated recovery disk is WDC <code>WD8004FRYZ-01VAEB0</code>, WWN <code>0x5000cca0c3c5c772</code>.
# A validation-script STOP is not automatically a storage failure. False STOPs occurred because of bad field parsing, comparing differently formatted LV manifests, attempting to read a root-only file through unprivileged shell redirection, and expecting <code>virsh</code> to work during a deliberately protected boot.
# On a protected <code>off</code> boot, <code>libvirt-guests.service</code> and libvirt management sockets are runtime-masked by design. Do not manually unmask them merely to run <code>virsh</code>. Validate protection from <code>kitsnet-vm-startup-control status</code> plus zero QEMU processes.
# T0 LVs can autoactivate after <code>partprobe</code>/udev activity. During post-clone geometry work, the meaningful safety tests are zero QEMU processes, no T0 mounts, and device-mapper open counts of zero. Autoactivation by itself is not corruption.
# The completed migration did not require a reboot between <code>growpart</code> and <code>pvresize</code>. GPT relocation, partition growth, kernel/udev settle, PARTUUID verification/restoration, and <code>pvresize</code> were completed in one protected boot.
# <code>pvs</code>/<code>vgs</code> may show the enlarged geometry before the explicit <code>pvresize</code>; this happened in the successful run. Still execute <code>pvresize</code> and require rc 0.
# Before every comparison against dock-resident baseline files, verify that <code>/mnt/KitsNet-pbd2</code> is actually mounted. An unmounted recovery filesystem can make a valid baseline look missing.
# Do not use <code>wc -l &lt; /root/file</code> as an unprivileged user and expect a leading <code>sudo</code> elsewhere to help; shell redirection happens before <code>sudo</code>. Read the dock baseline directly or run the complete read under <code>sudo</code>.


<pre>
=== Workload-specific handling ===
Normal production Wort
 
        |
For this maintenance, DONQ (the AlphaServer simulator inside <code>wv902</code>/crownroyal) was manually shut down before <code>wv902</code> was managed-saved because of VMScluster sensitivity. DONQ remains explicitly outside Wort automation and should be restarted manually after <code>wv902</code> returns. MORGAN on <code>uv056</code>/maddog required no special handling.
        v
 
Preflight and record identities
The shutdown produced exactly 18 managed-save files: <code>av001</code>, <code>av002</code>, <code>uv045</code>, <code>uv047</code>, <code>uv048</code>, <code>uv049</code>, <code>uv050</code>, <code>uv052</code>, <code>uv053</code>, <code>uv054</code>, <code>uv055</code>, <code>uv056</code>, <code>uv057</code>, <code>uv058</code>, <code>uv061</code>, <code>uv062</code>, <code>uv063</code>, and <code>wv902</code>. Fox (<code>uv059</code>) and Anchor (<code>uv060</code>) were cleanly shut down rather than managed-saved.
        |
 
        v
=== LVM devices-file lesson ===
Arm next boot OFF
 
        |
After restoration, T0 was initially absent from normal LVM discovery even though the cloned partition contained the correct PVID. The cause was <code>/etc/lvm/devices/system.devices</code>, which still associated the PVID with the old Plextor WWID.
        v
 
Controlled managed-save / NAS shutdown
Before changing the devices file, the clone was proven read-only with explicit <code>--devices /dev/nvme0n1p1</code> overrides. Then <code>lvmdevices --check --refresh</code> reported the new Sabrent <code>sys_wwid</code> and returned rc 5 because an update was required. <code>lvmdevices --update --refresh</code> updated the hardware association, and a final <code>lvmdevices --check</code> returned rc 0. No PVID, VG UUID, or LV UUID changed.
        |
 
        v
=== Partition-growth lesson ===
Unmount and disable kvm_save mount
 
        |
The successful expansion sequence was:
        v
 
Deactivate T0
<pre>
        |
sgdisk -e /dev/nvme0n1
        v
growpart -N /dev/nvme0n1 1
Whole NVMe -> dd -> pigz -p 12 -> dock image
growpart /dev/nvme0n1 1
        |
sync
        v
partprobe /dev/nvme0n1
Verify compressed image against original NVMe
udevadm settle
        |
verify start and PARTUUID
        v
pvresize /dev/nvme0n1p1
Power off Wort
</pre>
        |
 
        v
The start sector remained <code>2048</code>. During the run the partition unique GUID was unexpectedly observed as <code>BA942590-8B70-463D-AA42-5E03D7A280E6</code> rather than the original <code>3703AD9C-AB50-486A-B154-E83338BC7D3F</code>. Because disk GUID, partition contents, PVID, VG UUID, all 27 LV UUIDs, and filesystem identity had already been independently proven, the original PARTUUID was restored with:
Physically replace NVMe
 
        |
<syntaxhighlight lang="bash">
        v
for i in {1..10}; do echo; done
Protected no-VM boot
sudo sgdisk -u 1:3703AD9C-AB50-486A-B154-E83338BC7D3F /dev/nvme0n1
        |
sudo sync
        v
sudo partprobe /dev/nvme0n1
Verify replacement identity / geometry
sudo udevadm settle
        |
sudo sgdisk -v /dev/nvme0n1
        v
</syntaxhighlight>
Dock image -> pigz -d -> dd -> replacement NVMe
 
        |
Do not invent a replacement PARTUUID. Only restore the captured original after clone identity has already been proven.
        v
 
Byte-for-byte verification of restored region
=== Protected validation lesson ===
        |
 
        v
The final protected validation boot intentionally left libvirt management unavailable. The correct acceptance tests were:
Protected reboot
 
        |
* <code>current_boot_mode=off</code>;
        v
* <code>current_boot_protection=PROTECTED</code>;
Prove exact LVM clone before expansion
* <code>next_boot_mode=off</code> re-armed for safety;
        |
* zero QEMU processes;
        v
* GPT structurally clean;
Move GPT backup header if GPT
* partition start/PARTUUID correct;
        |
* T0 PVID/VG UUID and 27-LV inventory correct;
        v
* <code>lvmdevices --check</code> rc 0;
Grow partition 1
* <code>T0/kvm_save</code> automatically mounted with UUID <code>28ea4005-788c-44a6-a014-6a10d396dba7</code>; and
        |
* all 18 managed-save files still present with exact name/size matches.
        v
 
Protected reboot
Do not start a masked <code>virtqemud.socket</code> merely to make <code>virsh</code> available during this boot.
        |
 
        v
=== Production-start timing lesson ===
pvresize existing T0 PV
 
        |
The production startup is intentionally readiness-gated. In the successful boot:
        v
 
Restore and test kvm_save fstab mount
<pre>
        |
14:39:35  automatic startup begins; waits for Wort production storage
        v
14:40:01  production storage READY (26-second wait)
Protected validation reboot
14:40:02  uv059/Fox starts
        |
14:40:59  preferred Fox ownership READY; uv060/Anchor and uv061/mgr1 start
        v
14:41:32  protected NAS service path from mgr1 READY
Final validation with zero VMs
14:41:47  Docker Swarm manager READY
        |
14:41:47  uv062/wrk1 and uv063/wrk2 start
        v
14:42:53  automatic_controlled_vm_startup=COMPLETE
Select AUTO for next boot
</pre>
        |
 
        v
The approximately 48-second gap from mgr1 starting to the workers starting was normal. The controller was waiting first for <code>nas_service</code> and then for <code>swarm_manager</code>; the worker VMs started immediately when those gates passed. Do not manually start wrk1/wrk2 just because mgr1 already appears as running.
Normal controlled KitsNet startup
 
</pre>
An accidental reboot occurred during the first production-start attempt. The subsequent automatic boot completed normally and all expected guests returned. Wort did not retain a persistent previous-boot journal, so <code>journalctl -b -1</code> could not reconstruct the interrupted boot. Persistent journald may be enabled separately if boot-to-boot forensic history is desired.
 
=== Final accepted production state ===
 
<pre>
current_boot_mode=auto
current_boot_protection=RELEASED
next_boot_mode=auto
running_guest_count=20
vm_set_diff_rc=0
remaining_managed_save_count=0
startup_complete_target=active
startup_complete_marker_count=1
startup_process_count=0
T0=1.86 TiB total / 1.50 TiB free
kvm_save UUID=28ea4005-788c-44a6-a014-6a10d396dba7
lvmdevices_check_rc=0
NAS coherence_state=HEALTHY_FOX
NAS authority_owner=fox
nas_status_rc=0
failed systemd units=0
</pre>
 
The exact 20-domain production set is:
 
<pre>
av001
av002
uv045
uv047
uv048
uv049
uv050
uv052
uv053
uv054
uv055
uv056
uv057
uv058
uv059
uv060
uv061
uv062
uv063
wv902
</pre>
 
=== Unrelated systemd warning discovered during this work ===
 
<code>/etc/systemd/system/kitsnet-vm-startup-complete.target</code> contains a free-form <code>Documentation=KitsNet controlled VM startup completion point</code> line. <code>Documentation=</code> expects documentation URIs/identifiers, so systemd emits <code>Invalid URL</code> warnings. This did not prevent the target from becoming active and did not affect the NVMe migration. Remove that line or replace it with a valid documentation URI as separate housekeeping.
 
 
== Current storage design ==
 
At the time this procedure was written:
 
* <code>T0</code> is a single-PV volume group.
* The T0 PV is <code>/dev/nvme0n1p1</code>.
* The underlying whole device is <code>/dev/nvme0n1</code>.
* T0 contains 27 logical volumes used by the Wort virtual-machine environment.
* <code>T0/kvm_save</code> is a 64-GiB LV used for <code>/var/lib/libvirt/qemu/save</code>.
* Wort's host operating-system filesystems are on the separate <code>T0h</code> VG, not on T0.
* T0 can therefore be completely deactivated while Wort itself remains operational.
 
The migration deliberately clones the '''entire NVMe device''', not merely partition 1.
 
A whole-device clone preserves the on-disk:
 
* partition table;
* disk GUID, if GPT;
* partition GUID/PARTUUID;
* LVM PV UUID;
* T0 VG UUID;
* LV UUIDs;
* filesystem UUIDs;
* LVM metadata;
* allocated and unallocated extents; and
* all VM disk and managed-save contents.


== Stage 0 — pre-maintenance validation ==
The replacement NVMe's physical model, serial number, WWID and other controller-level identifiers are not cloned.


This stage is performed while Wort is operating normally and before any VM shutdown.
== Governing safety rules ==


=== Confirm that this is Wort ===
<blockquote>
'''Do not improvise past a STOP checkpoint. If a required validation does not produce the documented result, stop the procedure and diagnose the discrepancy before proceeding.'''
</blockquote>


<syntaxhighlight lang="bash">
The following rules apply for the entire maintenance:
for i in {1..10}; do echo; done
hostname
hostname -f
</syntaxhighlight>


'''Expected:''' the host is Wort.
# The original NVMe is not erased, modified, reformatted or repurposed after removal until the replacement has passed final production validation.
# No <code>pvcreate</code>, <code>vgcreate</code>, <code>vgimportclone</code>, <code>pvchange --uuid</code>, <code>vgchange --uuid</code> or routine <code>vgcfgrestore</code> is used.
# The image is made from the '''whole device''' <code>/dev/nvme0n1</code>.
# The image is restored to the '''whole replacement device'''.
# T0 must be inactive while the source image is created.
# The replacement NVMe must be at least as large in bytes as the source NVMe.
# The replacement NVMe must present the '''same logical sector size''' as the source NVMe.
# The clone is validated before any expansion operation is performed.
# Partition expansion and PV expansion are separate from the cloning operation.
# The <code>/var/lib/libvirt/qemu/save</code> mount remains available until the controlled VM shutdown/managed-save operation is complete.
# The save mount is then unmounted and disabled in <code>/etc/fstab</code> until T0 has been restored and expanded.
# Every maintenance boot remains VM-protected.
# Because <code>next-boot off</code> is one-shot, '''the first administrative action after every protected maintenance boot is to arm <code>next-boot off</code> again for the following boot'''.
# The final production boot is the only boot for which <code>next-boot auto</code> is explicitly selected.


'''STOP''' if these commands are being executed on another host.
== Migration overview ==


=== Define the source-device names ===
The controlled sequence is:


The following names reflect the current T0 design.
<pre>
 
Normal production Wort
<syntaxhighlight lang="bash">
        |
for i in {1..10}; do echo; done
        v
SRC=/dev/nvme0n1
Preflight and record identities
PV=/dev/nvme0n1p1
        |
 
        v
echo "SRC=$SRC"
Arm next boot OFF
echo "PV=$PV"
        |
 
        v
sudo pvs "$PV"
Controlled managed-save / NAS shutdown
sudo vgs T0
        |
sudo lvs T0
        v
</syntaxhighlight>
Unmount and disable kvm_save mount
 
        |
'''Expected:'''
        v
 
Deactivate T0
* <code>/dev/nvme0n1p1</code> belongs to VG <code>T0</code>.
        |
* T0 has exactly one PV.
        v
 
Whole NVMe -> dd -> pigz -p 12 -> dock image
'''STOP''' if the source device or T0 topology differs.
        |
 
        v
=== Confirm the currently mounted libvirt save filesystem ===
Verify compressed image against original NVMe
 
        |
<syntaxhighlight lang="bash">
        v
for i in {1..10}; do echo; done
Power off Wort
sudo findmnt --mountpoint /var/lib/libvirt/qemu/save
        |
echo
        v
sudo lvs -o vg_name,lv_name,lv_size,lv_attr T0/kvm_save
Physically replace NVMe
echo
        |
grep -n '/var/lib/libvirt/qemu/save' /etc/fstab
        v
</syntaxhighlight>
Protected no-VM boot
 
        |
The mount must be understood before continuing.
        v
 
Verify replacement identity / geometry
Do '''not''' unmount it yet. The production VM shutdown mechanism needs this filesystem for ordinary managed-save files.
        |
 
        v
=== Verify the production VM shutdown override ===
Dock image -> pigz -d -> dd -> replacement NVMe
 
        |
<syntaxhighlight lang="bash">
        v
for i in {1..10}; do echo; done
Byte-for-byte verification of restored region
systemctl cat libvirt-guests.service
        |
echo
        v
systemctl cat libvirt-guests.service |
Protected reboot
grep -F '/usr/local/sbin/libvirt-guests-parallel.sh stop'
        |
</syntaxhighlight>
        v
Prove exact LVM clone before expansion
        |
        v
Move GPT backup header if GPT
        |
        v
Grow partition 1
        |
        v
Verify/restore original PARTUUID if required
        |
        v
pvresize existing T0 PV in same protected boot
        |
        v
Restore and test kvm_save fstab mount
        |
        v
Protected validation reboot
        |
        v
Final validation with zero VMs
        |
        v
Select AUTO for next boot
        |
        v
Normal controlled KitsNet startup
</pre>


'''Expected:''' the effective <code>ExecStop</code> path includes:
== Stage 0 — pre-maintenance validation ==


<pre>
This stage is performed while Wort is operating normally and before any VM shutdown.
/usr/local/sbin/libvirt-guests-parallel.sh stop
</pre>


'''STOP''' if the KitsNet override is not in effect.
=== Confirm that this is Wort ===
 
=== Verify required utilities ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
for cmd in \
hostname
  dd pigz cmp blockdev lsblk sfdisk blkid \
hostname -f
  pvs vgs lvs vgchange pvresize pvck vgck vgcfgbackup \
</syntaxhighlight>
  lvmconfig lvmdevices pvscan \
  findmnt mountpoint growpart
do
    if command -v "$cmd" >/dev/null 2>&1; then
        printf 'OK      %s -> %s\n' "$cmd" "$(command -v "$cmd")"
    else
        printf 'MISSING %s\n' "$cmd"
    fi
done


echo
'''Expected:''' the host is Wort.
if command -v nvme >/dev/null 2>&1; then
    echo "OPTIONAL nvme -> $(command -v nvme)"
else
    echo "OPTIONAL nvme command is not installed"
fi


echo
'''STOP''' if these commands are being executed on another host.
if command -v sgdisk >/dev/null 2>&1; then
    echo "sgdisk -> $(command -v sgdisk)"
else
    echo "sgdisk is not installed; it is required if the source disk is GPT"
fi
</syntaxhighlight>


'''STOP''' if any mandatory utility is missing.
=== Define the source-device names ===


If the source is GPT, <code>sgdisk</code> is also mandatory.
The following names reflect the current T0 design.
 
Install/repair required tooling before entering the maintenance outage.
 
=== Record source geometry ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
Line 258: Line 354:
PV=/dev/nvme0n1p1
PV=/dev/nvme0n1p1


echo "===== SOURCE DEVICE ====="
echo "SRC=$SRC"
sudo lsblk -o NAME,PATH,SIZE,TYPE,FSTYPE,PTTYPE,MODEL,SERIAL,WWN,LOG-SEC,PHY-SEC "$SRC"
echo "PV=$PV"


echo
sudo pvs "$PV"
echo "===== SOURCE SIZE IN BYTES ====="
sudo vgs T0
sudo blockdev --getsize64 "$SRC"
sudo lvs T0
</syntaxhighlight>


echo
'''Expected:'''
echo "===== SOURCE LOGICAL SECTOR SIZE ====="
 
sudo blockdev --getss "$SRC"
* <code>/dev/nvme0n1p1</code> belongs to VG <code>T0</code>.
* T0 has exactly one PV.
 
'''STOP''' if the source device or T0 topology differs.


echo
=== Confirm the currently mounted libvirt save filesystem ===
echo "===== PARTITION START ====="
sudo lsblk -no START "$PV"


<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
sudo findmnt --mountpoint /var/lib/libvirt/qemu/save
echo
echo
echo "===== PARTUUID ====="
sudo lvs -o vg_name,lv_name,lv_size,lv_attr T0/kvm_save
sudo blkid -s PARTUUID -o value "$PV"
 
echo
echo
echo "===== PARTITION TABLE TYPE ====="
grep -n '/var/lib/libvirt/qemu/save' /etc/fstab
sudo blkid -p -s PTTYPE -o value "$SRC"
</syntaxhighlight>
</syntaxhighlight>


The partition-table type is expected to be either:
The mount must be understood before continuing.


<pre>
Do '''not''' unmount it yet. The production VM shutdown mechanism needs this filesystem for ordinary managed-save files.
gpt
</pre>


or:
=== Verify the production VM shutdown override ===


<pre>
<syntaxhighlight lang="bash">
dos
for i in {1..10}; do echo; done
</pre>
systemctl cat libvirt-guests.service
echo
systemctl cat libvirt-guests.service |
grep -F '/usr/local/sbin/libvirt-guests-parallel.sh stop'
</syntaxhighlight>


Do not assume GPT until this command has confirmed it.
'''Expected:''' the effective <code>ExecStop</code> path includes:


=== Configure and validate the dock destination ===
<pre>
/usr/local/sbin/libvirt-guests-parallel.sh stop
</pre>


Set <code>DOCK</code> to the actual mounted dock filesystem.
'''STOP''' if the KitsNet override is not in effect.


'''This is the only site-specific path that must be substituted in the command blocks.'''
=== Verify required utilities ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
for cmd in \
SRC=/dev/nvme0n1
  dd pigz cmp blockdev lsblk sfdisk blkid \
 
  pvs vgs lvs vgchange pvresize pvck vgck vgcfgbackup \
echo "DOCK=$DOCK"
  lvmconfig lvmdevices pvscan \
 
  findmnt mountpoint growpart
echo
do
echo "===== DOCK MOUNT ====="
    if command -v "$cmd" >/dev/null 2>&1; then
findmnt --mountpoint "$DOCK"
        printf 'OK      %s -> %s\n' "$cmd" "$(command -v "$cmd")"
    else
        printf 'MISSING %s\n' "$cmd"
    fi
done


echo
echo
echo "===== DOCK FILESYSTEM ====="
if command -v nvme >/dev/null 2>&1; then
df -Th "$DOCK"
    echo "OPTIONAL nvme -> $(command -v nvme)"
else
    echo "OPTIONAL nvme command is not installed"
fi


echo
echo
echo "===== DOCK FREE BYTES ====="
if command -v sgdisk >/dev/null 2>&1; then
DOCK_FREE=$(df -B1 --output=avail "$DOCK" | tail -1 | tr -d ' ')
    echo "sgdisk -> $(command -v sgdisk)"
SRC_BYTES=$(sudo blockdev --getsize64 "$SRC")
else
    echo "sgdisk is not installed; it is required if the source disk is GPT"
fi
</syntaxhighlight>


echo "source_bytes=$SRC_BYTES"
'''STOP''' if any mandatory utility is missing.
echo "dock_free_bytes=$DOCK_FREE"


NEED_BYTES=$((SRC_BYTES + SRC_BYTES / 20))
If the source is GPT, <code>sgdisk</code> is also mandatory.
echo "recommended_minimum_free_bytes=$NEED_BYTES"


if [ "$DOCK_FREE" -ge "$NEED_BYTES" ]; then
Install/repair required tooling before entering the maintenance outage.
    echo "PASS: dock has at least source size plus 5 percent"
else
    echo "STOP: insufficient worst-case dock capacity"
fi
</syntaxhighlight>


The 5-percent allowance ensures that the procedure does not depend on the source being compressible.
=== Record source geometry ===
 
'''STOP''' unless the dock mount and capacity are correct.
 
=== Create the migration directory ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
SRC=/dev/nvme0n1
MIG="$DOCK/wort-T0-nvme-migration"
PV=/dev/nvme0n1p1


sudo mkdir -p "$MIG"
echo "===== SOURCE DEVICE ====="
sudo touch "$MIG/.write-test"
sudo lsblk -o NAME,PATH,SIZE,TYPE,FSTYPE,PTTYPE,MODEL,SERIAL,WWN,LOG-SEC,PHY-SEC "$SRC"
sudo rm -f "$MIG/.write-test"


echo "MIG=$MIG"
echo
sudo ls -ld "$MIG"
echo "===== SOURCE SIZE IN BYTES ====="
</syntaxhighlight>
sudo blockdev --getsize64 "$SRC"


=== Capture the pre-migration identity and recovery metadata ===
echo
 
echo "===== SOURCE LOGICAL SECTOR SIZE ====="
<syntaxhighlight lang="bash">
sudo blockdev --getss "$SRC"
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
SRC=/dev/nvme0n1
PV=/dev/nvme0n1p1


PTTYPE=$(sudo blkid -p -s PTTYPE -o value "$SRC")
echo
echo "===== PARTITION START AND SIZE ====="
sudo sfdisk --dump "$SRC" | grep -F "$PV"
PART_START=$(sudo sfdisk --dump "$SRC" 2>/dev/null | sed -n "\|^$PV |s/.*start=[[:space:]]*\([0-9][0-9]*\),.*/\1/p")
PART_SIZE=$(sudo sfdisk --dump "$SRC" 2>/dev/null | sed -n "\|^$PV |s/.*size=[[:space:]]*\([0-9][0-9]*\),.*/\1/p")
echo "partition_start=$PART_START"
echo "partition_size_sectors=$PART_SIZE"


sudo blockdev --getsize64 "$SRC" |
echo
sudo tee "$MIG/source.bytes"
echo "===== PARTUUID ====="
sudo blkid -s PARTUUID -o value "$PV"


sudo blockdev --getss "$SRC" |
echo
sudo tee "$MIG/source.logical-sector-size"
echo "===== PARTITION TABLE TYPE ====="
sudo blkid -p -s PTTYPE -o value "$SRC"
</syntaxhighlight>


printf '%s\n' "$PTTYPE" |
The partition-table type is expected to be either:
sudo tee "$MIG/source.pttype"


sudo sfdisk --dump "$SRC" |
<pre>
sudo tee "$MIG/source.sfdisk"
gpt
</pre>


sudo lsblk -no START "$PV" |
or:
tr -d ' ' |
sudo tee "$MIG/partition-start.before"


sudo blkid -s PARTUUID -o value "$PV" |
<pre>
sudo tee "$MIG/partuuid.before"
dos
</pre>


sudo pvs --noheadings -o pv_uuid "$PV" |
Do not assume GPT until this command has confirmed it.
tr -d ' ' |
sudo tee "$MIG/pvid.before"


sudo vgs --noheadings -o vg_uuid T0 |
=== Configure and validate the dock destination ===
tr -d ' ' |
sudo tee "$MIG/vguid.before"


sudo lvs --noheadings --separator '|' -o lv_name,lv_uuid T0 |
Set <code>DOCK</code> to the actual mounted dock filesystem.
sed 's/^[[:space:]]*//;s/[[:space:]]*$//' |
sort |
sudo tee "$MIG/lvuuids.before"


sudo blkid -s UUID -o value /dev/T0/kvm_save |
The validated recovery mount is <code>/mnt/KitsNet-pbd2</code>. Verify that this mount is present before every use of dock-resident baselines or the image.
sudo tee "$MIG/kvm-save-fsuuid.before"


sudo findmnt --mountpoint /var/lib/libvirt/qemu/save \
<syntaxhighlight lang="bash">
  -o SOURCE,TARGET,FSTYPE,OPTIONS |
for i in {1..10}; do echo; done
sudo tee "$MIG/kvm-save-mount.before"
DOCK=/mnt/KitsNet-pbd2
SRC=/dev/nvme0n1


sudo cp -a /etc/fstab "$MIG/fstab.before"
echo "DOCK=$DOCK"


sudo vgcfgbackup -f "$MIG/T0.vgcfg" T0
echo
echo "===== DOCK MOUNT ====="
findmnt --mountpoint "$DOCK"


if sudo test -d /etc/lvm/devices; then
echo
    sudo cp -a /etc/lvm/devices "$MIG/lvm-devices.before"
echo "===== DOCK FILESYSTEM ====="
fi
df -Th "$DOCK"


if [ "$PTTYPE" = "gpt" ]; then
echo
    sudo sgdisk --backup="$MIG/source.gpt" "$SRC"
echo "===== DOCK FREE BYTES ====="
DOCK_FREE=$(df -B1 --output=avail "$DOCK" | tail -1 | tr -d ' ')
SRC_BYTES=$(sudo blockdev --getsize64 "$SRC")


    sudo sgdisk -p "$SRC" |
echo "source_bytes=$SRC_BYTES"
    awk '/Disk identifier \(GUID\):/{print $4}' |
echo "dock_free_bytes=$DOCK_FREE"
    sudo tee "$MIG/disk-guid.before"


    sudo sgdisk -v "$SRC"
NEED_BYTES=$((SRC_BYTES + SRC_BYTES / 20))
fi
echo "recommended_minimum_free_bytes=$NEED_BYTES"
</syntaxhighlight>


For GPT, <code>sgdisk -v</code> must not report existing structural corruption.
if [ "$DOCK_FREE" -ge "$NEED_BYTES" ]; then
    echo "PASS: dock has at least source size plus 5 percent"
else
    echo "STOP: insufficient worst-case dock capacity"
fi
</syntaxhighlight>


'''STOP''' if the source partition table or LVM metadata already appears damaged.
The 5-percent allowance ensures that the procedure does not depend on the source being compressible.


== Stage 1 — quiesce all Wort VMs ==
'''STOP''' unless the dock mount and capacity are correct.


=== Arm the next boot for zero VMs ===
=== Create the migration directory ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
DOCK=/mnt/KitsNet-pbd2
echo
MIG="$DOCK/wort-T0-nvme-migration"
sudo /usr/local/sbin/kitsnet-vm-startup-control status
</syntaxhighlight>


Verify that the next boot is configured for <code>off</code> mode.
sudo mkdir -p "$MIG"
sudo touch "$MIG/.write-test"
sudo rm -f "$MIG/.write-test"


=== Run the accepted production shutdown sequence ===
echo "MIG=$MIG"
sudo ls -ld "$MIG"
</syntaxhighlight>


Stopping <code>libvirt-guests.service</code> invokes the KitsNet production wrapper. The wrapper performs managed-save processing for ordinary guests and then invokes the owner-aware NAS shutdown mechanism.
=== Capture the pre-migration identity and recovery metadata ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo systemctl stop libvirt-guests.service
DOCK=/mnt/KitsNet-pbd2
RC=$?
MIG="$DOCK/wort-T0-nvme-migration"
SRC=/dev/nvme0n1
PV=/dev/nvme0n1p1


echo
PTTYPE=$(sudo blkid -p -s PTTYPE -o value "$SRC")
echo "libvirt-guests stop rc=$RC"
echo
systemctl status libvirt-guests.service --no-pager -l
</syntaxhighlight>


'''Required result:'''
sudo blockdev --getsize64 "$SRC" |
sudo tee "$MIG/source.bytes"


<pre>
sudo blockdev --getss "$SRC" |
libvirt-guests stop rc=0
sudo tee "$MIG/source.logical-sector-size"
</pre>


'''STOP''' if the stop operation fails.
printf '%s\n' "$PTTYPE" |
sudo tee "$MIG/source.pttype"


Do not manually destroy a guest to force the procedure onward.
sudo sfdisk --dump "$SRC" |
sudo tee "$MIG/source.sfdisk"


=== Prove that no VM remains running ===
sudo sfdisk --dump "$SRC" 2>/dev/null |
sed -n "\|^$PV |s/.*start=[[:space:]]*\([0-9][0-9]*\),.*/\1/p" |
sudo tee "$MIG/partition-start.before"


<syntaxhighlight lang="bash">
sudo sfdisk --dump "$SRC" 2>/dev/null |
for i in {1..10}; do echo; done
sed -n "\|^$PV |s/.*size=[[:space:]]*\([0-9][0-9]*\),.*/\1/p" |
echo "===== RUNNING LIBVIRT GUESTS ====="
sudo tee "$MIG/partition-size.before"
sudo virsh list --state-running --name


echo
sudo blkid -s PARTUUID -o value "$PV" |
echo "===== RUNNING GUEST COUNT ====="
sudo tee "$MIG/partuuid.before"
RUNNING=$(sudo virsh list --state-running --name |
  sed '/^[[:space:]]*$/d' |
  wc -l)


echo "running_guests=$RUNNING"
sudo pvs --noheadings -o pv_uuid "$PV" |
tr -d ' ' |
sudo tee "$MIG/pvid.before"


echo
sudo vgs --noheadings -o vg_uuid T0 |
echo "===== QEMU PROCESSES ====="
tr -d ' ' |
ps -eo pid,args |
sudo tee "$MIG/vguid.before"
grep -E '[q]emu-system|[q]emu-kvm' || true
</syntaxhighlight>


'''Required result:'''
sudo lvs --noheadings --separator '|' -o lv_name,lv_uuid T0 |
sed 's/^[[:space:]]*//;s/[[:space:]]*$//' |
sort |
sudo tee "$MIG/lvuuids.before"


<pre>
sudo blkid -s UUID -o value /dev/T0/kvm_save |
running_guests=0
sudo tee "$MIG/kvm-save-fsuuid.before"
</pre>


'''STOP''' if any guest or QEMU process remains active.
sudo findmnt --mountpoint /var/lib/libvirt/qemu/save \
  -o SOURCE,TARGET,FSTYPE,OPTIONS |
sudo tee "$MIG/kvm-save-mount.before"


=== Record the managed-save files before unmounting T0/kvm_save ===
sudo cp -a /etc/fstab "$MIG/fstab.before"


<syntaxhighlight lang="bash">
sudo vgcfgbackup -f "$MIG/T0.vgcfg" T0
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"


sudo find /var/lib/libvirt/qemu/save \
if sudo test -d /etc/lvm/devices; then
  -maxdepth 1 \
    sudo cp -a /etc/lvm/devices "$MIG/lvm-devices.before"
  -type f \
fi
  -printf '%f|%s\n' |
sort |
sudo tee "$MIG/managed-save.before"
</syntaxhighlight>


The exact count depends on which ordinary guests were running before shutdown. The file list will later be compared after restoration.
if [ "$PTTYPE" = "gpt" ]; then
    sudo sgdisk --backup="$MIG/source.gpt" "$SRC"


== Stage 2 — remove the libvirt save mount from the maintenance boot path ==
    sudo sgdisk -p "$SRC" |
    awk '/Disk identifier \(GUID\):/{print $4}' |
    sudo tee "$MIG/disk-guid.before"


=== Unmount the save filesystem ===
    sudo sgdisk -v "$SRC"
fi
</syntaxhighlight>
 
For GPT, <code>sgdisk -v</code> must not report existing structural corruption.
 
'''STOP''' if the source partition table or LVM metadata already appears damaged.
 
== Stage 1 — quiesce all Wort VMs ==


<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
sudo umount /var/lib/libvirt/qemu/save
RC=$?


echo "umount rc=$RC"
=== Workload-specific pre-quiesce note ===


if mountpoint -q /var/lib/libvirt/qemu/save; then
For the 12 September 2026 maintenance, DONQ (the AlphaServer simulator running inside <code>wv902</code>/crownroyal) was manually shut down cleanly before <code>wv902</code> was managed-saved. This was an explicit VMScluster sensitivity constraint and remains outside Wort's automatic startup/shutdown model. MORGAN on <code>uv056</code>/maddog required no special handling.
    echo "STOP: /var/lib/libvirt/qemu/save is still mounted"
else
    echo "PASS: save filesystem is unmounted"
fi
</syntaxhighlight>


'''STOP''' unless <code>umount rc=0</code> and the filesystem is no longer a mountpoint.
Do not add DONQ-specific automation to this runbook. If this workload constraint still applies, perform the same manual DONQ shutdown before the host-level quiesce.


=== Verify exactly one active fstab entry exists before commenting it ===
=== Arm the next boot for zero VMs ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
ACTIVE_FSTAB=$(
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
  awk '
echo
    $0 !~ /^[[:space:]]*#/ &&
sudo /usr/local/sbin/kitsnet-vm-startup-control status
    $2 == "/var/lib/libvirt/qemu/save" { n++ }
</syntaxhighlight>
    END { print n+0 }
 
  ' /etc/fstab
Verify that the next boot is configured for <code>off</code> mode.
)
 
=== Run the accepted production shutdown sequence ===
 
Stopping <code>libvirt-guests.service</code> invokes the KitsNet production wrapper. The wrapper performs managed-save processing for ordinary guests and then invokes the owner-aware NAS shutdown mechanism.
 
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
sudo systemctl stop libvirt-guests.service
RC=$?


echo "active_save_fstab_entries=$ACTIVE_FSTAB"
echo
grep -n '/var/lib/libvirt/qemu/save' /etc/fstab
echo "libvirt-guests stop rc=$RC"
echo
systemctl status libvirt-guests.service --no-pager -l
</syntaxhighlight>
</syntaxhighlight>


Line 550: Line 662:


<pre>
<pre>
active_save_fstab_entries=1
libvirt-guests stop rc=0
</pre>
</pre>


'''STOP''' otherwise.
'''STOP''' if the stop operation fails.
 
Do not manually destroy a guest to force the procedure onward.


=== Comment the save mount ===
=== Prove that no VM remains running ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo sed -i \
echo "===== RUNNING LIBVIRT GUESTS ====="
  '\|^[[:space:]]*[^#].*[[:space:]]/var/lib/libvirt/qemu/save[[:space:]]| s|^|# KN-T0-NVME-MIGRATION |' \
sudo virsh list --state-running --name
  /etc/fstab


echo "===== RESULTING FSTAB ENTRY ====="
echo
grep -n '/var/lib/libvirt/qemu/save' /etc/fstab
echo "===== RUNNING GUEST COUNT ====="
RUNNING=$(sudo virsh list --state-running --name |
  sed '/^[[:space:]]*$/d' |
  wc -l)


echo
echo "running_guests=$RUNNING"
echo "===== ACTIVE ENTRIES REMAINING ====="
awk '
  $0 !~ /^[[:space:]]*#/ &&
  $2 == "/var/lib/libvirt/qemu/save" { print }
' /etc/fstab


echo
echo
sudo systemctl daemon-reload
echo "===== QEMU PROCESSES ====="
sudo findmnt --verify --verbose
ps -eo pid,args |
grep -E '[q]emu-system|[q]emu-kvm' || true
</syntaxhighlight>
</syntaxhighlight>


There must be no active fstab entry for the save mount.
'''Required result:'''
 
The original line must remain present with the prefix:


<pre>
<pre>
# KN-T0-NVME-MIGRATION
running_guests=0
</pre>
</pre>


== Stage 3 — deactivate T0 completely ==
'''STOP''' if any guest or QEMU process remains active.
 
=== Record the managed-save files before unmounting T0/kvm_save ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo vgchange -an T0
DOCK=/mnt/KitsNet-pbd2
RC=$?
MIG="$DOCK/wort-T0-nvme-migration"


echo
sudo find /var/lib/libvirt/qemu/save \
echo "vgchange rc=$RC"
  -maxdepth 1 \
  -type f \
  -name '*.save' \
  -printf '%f|%s\n' |
sort |
sudo tee "$MIG/managed-save.before"
</syntaxhighlight>


echo
The exact count depends on which ordinary guests were running before shutdown. The file list will later be compared after restoration.
echo "===== T0 LV STATE ====="
sudo lvs -o vg_name,lv_name,lv_active,lv_attr,lv_size T0


echo
For the validated execution the manifest contained exactly 18 files: <code>av001</code>, <code>av002</code>, <code>uv045</code>, <code>uv047</code>, <code>uv048</code>, <code>uv049</code>, <code>uv050</code>, <code>uv052</code>, <code>uv053</code>, <code>uv054</code>, <code>uv055</code>, <code>uv056</code>, <code>uv057</code>, <code>uv058</code>, <code>uv061</code>, <code>uv062</code>, <code>uv063</code>, and <code>wv902</code>. Fox (<code>uv059</code>) and Anchor (<code>uv060</code>) were cleanly shut down rather than managed-saved.
ACTIVE=$(
  sudo lvs --noheadings -o lv_active T0 |
  awk '$1 == "active" { n++ } END { print n+0 }'
)


echo "active_T0_LVs=$ACTIVE"
== Stage 2 — remove the libvirt save mount from the maintenance boot path ==
</syntaxhighlight>


'''Required results:'''
=== Unmount the save filesystem ===


<pre>
<syntaxhighlight lang="bash">
vgchange rc=0
for i in {1..10}; do echo; done
active_T0_LVs=0
sudo umount /var/lib/libvirt/qemu/save
</pre>
RC=$?


'''STOP''' if T0 cannot be completely deactivated.
echo "umount rc=$RC"


Do not image a partially active T0.
if mountpoint -q /var/lib/libvirt/qemu/save; then
    echo "STOP: /var/lib/libvirt/qemu/save is still mounted"
else
    echo "PASS: save filesystem is unmounted"
fi
</syntaxhighlight>


== Stage 4 — create the compressed whole-NVMe image ==
'''STOP''' unless <code>umount rc=0</code> and the filesystem is no longer a mountpoint.


=== Revalidate source and destination immediately before imaging ===
=== Verify exactly one active fstab entry exists before commenting it ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
ACTIVE_FSTAB=$(
MIG="$DOCK/wort-T0-nvme-migration"
  awk '
SRC=/dev/nvme0n1
    $0 !~ /^[[:space:]]*#/ &&
PV=/dev/nvme0n1p1
    $2 == "/var/lib/libvirt/qemu/save" { n++ }
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"
    END { print n+0 }
  ' /etc/fstab
)
 
echo "active_save_fstab_entries=$ACTIVE_FSTAB"
grep -n '/var/lib/libvirt/qemu/save' /etc/fstab
</syntaxhighlight>


echo "===== SOURCE ====="
'''Required result:'''
sudo lsblk -o NAME,PATH,SIZE,TYPE,FSTYPE,MODEL,SERIAL,LOG-SEC,PHY-SEC "$SRC"


echo
<pre>
echo "===== SOURCE PV ====="
active_save_fstab_entries=1
sudo pvs "$PV"
</pre>


echo
'''STOP''' otherwise.
echo "===== IMAGE DESTINATION ====="
echo "$IMG"


if sudo test -e "$IMG"; then
=== Comment the save mount ===
    echo "STOP: image file already exists; it will NOT be overwritten automatically"
else
    echo "PASS: image pathname is unused"
fi


echo
<syntaxhighlight lang="bash">
df -h "$DOCK"
for i in {1..10}; do echo; done
sudo sed -i \
  '\|^[[:space:]]*[^#].*[[:space:]]/var/lib/libvirt/qemu/save[[:space:]]| s|^|# KN-T0-NVME-MIGRATION |' \
  /etc/fstab
 
echo "===== RESULTING FSTAB ENTRY ====="
grep -n '/var/lib/libvirt/qemu/save' /etc/fstab
 
echo
echo "===== ACTIVE ENTRIES REMAINING ====="
awk '
  $0 !~ /^[[:space:]]*#/ &&
  $2 == "/var/lib/libvirt/qemu/save" { print }
' /etc/fstab
 
echo
sudo systemctl daemon-reload
sudo findmnt --verify --verbose
</syntaxhighlight>
</syntaxhighlight>


'''STOP''' if the source device is not the known T0 NVMe or if the image filename already exists unexpectedly.
There must be no active fstab entry for the save mount.


=== Create the image ===
The original line must remain present with the prefix:


This command:
<pre>
# KN-T0-NVME-MIGRATION
</pre>


* reads the whole source NVMe using <code>dd</code>;
== Stage 3 — deactivate T0 completely ==
* feeds the raw byte stream into <code>pigz</code>;
* allows pigz to use up to 12 compression threads;
* uses normal gzip compression level 6; and
* enables pipeline failure propagation.
 
No <code>conv=noerror</code> option is used. A source read error is a migration failure and must not be silently padded or ignored.


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
sudo vgchange -an T0
MIG="$DOCK/wort-T0-nvme-migration"
SRC=/dev/nvme0n1
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"
 
sudo env SRC="$SRC" IMG="$IMG" \
bash -o pipefail -c '
    if [ -e "$IMG" ]; then
        echo "STOP: refusing to overwrite existing image: $IMG"
        exit 2
    fi
 
    dd if="$SRC" bs=16M iflag=fullblock status=progress |
    pigz -p 12 -6 > "$IMG"
'
 
RC=$?
RC=$?


echo
echo
echo "image_pipeline_rc=$RC"
echo "vgchange rc=$RC"


sudo sync
echo
echo "===== T0 LV STATE ====="
sudo lvs -o vg_name,lv_name,lv_active,lv_attr,lv_size T0


echo
echo
sudo ls -lh "$IMG"
ACTIVE=$(
  sudo lvs --noheadings -o lv_active T0 |
  awk '$1 == "active" { n++ } END { print n+0 }'
)
 
echo "active_T0_LVs=$ACTIVE"
</syntaxhighlight>
</syntaxhighlight>


'''Required result:'''
'''Required results:'''


<pre>
<pre>
image_pipeline_rc=0
vgchange rc=0
active_T0_LVs=0
</pre>
</pre>


'''STOP''' otherwise.
'''STOP''' if T0 cannot be completely deactivated.
 
Do not image a partially active T0.


== Stage 5 — validate the image before removing the original NVMe ==
== Stage 4 — create the compressed whole-NVMe image ==


=== Test the gzip stream ===
=== Revalidate source and destination immediately before imaging ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
MIG="$DOCK/wort-T0-nvme-migration"
SRC=/dev/nvme0n1
PV=/dev/nvme0n1p1
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"


sudo pigz -t "$IMG"
echo "===== SOURCE ====="
RC=$?
sudo lsblk -o NAME,PATH,SIZE,TYPE,FSTYPE,MODEL,SERIAL,LOG-SEC,PHY-SEC "$SRC"


echo "pigz_test_rc=$RC"
echo
</syntaxhighlight>
echo "===== SOURCE PV ====="
sudo pvs "$PV"
 
echo
echo "===== IMAGE DESTINATION ====="
echo "$IMG"


'''Required result:'''
if sudo test -e "$IMG"; then
    echo "STOP: image file already exists; it will NOT be overwritten automatically"
else
    echo "PASS: image pathname is unused"
fi


<pre>
echo
pigz_test_rc=0
df -h "$DOCK"
</pre>
</syntaxhighlight>


=== Compare the decompressed image byte-for-byte with the original NVMe ===
'''STOP''' if the source device is not the known T0 NVMe or if the image filename already exists unexpectedly.


This is the definitive pre-swap verification.
=== Create the image ===


<syntaxhighlight lang="bash">
This command:
for i in {1..10}; do echo; done
 
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
* reads the whole source NVMe using <code>dd</code>;
MIG="$DOCK/wort-T0-nvme-migration"
* feeds the raw byte stream into <code>pigz</code>;
SRC=/dev/nvme0n1
* allows pigz to use up to 12 compression threads;
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"
* uses normal gzip compression level 6; and
* enables pipeline failure propagation.
 
No <code>conv=noerror</code> option is used. A source read error is a migration failure and must not be silently padded or ignored.
 
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
SRC=/dev/nvme0n1
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"


sudo env SRC="$SRC" IMG="$IMG" \
sudo env SRC="$SRC" IMG="$IMG" \
bash -o pipefail -c '
bash -o pipefail -c '
     pigz -dc "$IMG" |
     if [ -e "$IMG" ]; then
     cmp - "$SRC"
        echo "STOP: refusing to overwrite existing image: $IMG"
        exit 2
    fi
 
    dd if="$SRC" bs=16M iflag=fullblock status=progress |
     pigz -p 12 -6 > "$IMG"
'
'


Line 746: Line 895:


echo
echo
echo "source_image_cmp_rc=$RC"
echo "image_pipeline_rc=$RC"
</syntaxhighlight>


<code>cmp</code> normally prints nothing when the inputs are identical.
sudo sync
 
echo
sudo ls -lh "$IMG"
</syntaxhighlight>


'''Required result:'''
'''Required result:'''


<pre>
<pre>
source_image_cmp_rc=0
image_pipeline_rc=0
</pre>
</pre>


'''STOP''' if the comparison is anything other than zero.
'''STOP''' otherwise.
 
== Stage 5 — validate the image before removing the original NVMe ==


=== Record a checksum of the compressed recovery artifact ===
=== Test the gzip stream ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
MIG="$DOCK/wort-T0-nvme-migration"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"


sudo sha256sum "$IMG" |
sudo pigz -t "$IMG"
sudo tee "$IMG.sha256"
RC=$?


sudo ls -lh "$IMG" "$IMG.sha256"
echo "pigz_test_rc=$RC"
</syntaxhighlight>
</syntaxhighlight>


At this point there are two independent recovery assets:
'''Required result:'''


# the untouched original physical NVMe; and
<pre>
# the byte-verified compressed whole-device image.
pigz_test_rc=0
</pre>


== Stage 6 — power off Wort and replace the NVMe ==
=== Compare the decompressed image byte-for-byte with the original NVMe ===


Re-arm protected startup immediately before shutdown, even though it should already be armed.
This is the definitive pre-swap verification.


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
DOCK=/mnt/KitsNet-pbd2
sudo /usr/local/sbin/kitsnet-vm-startup-control status
MIG="$DOCK/wort-T0-nvme-migration"
</syntaxhighlight>
SRC=/dev/nvme0n1
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"


Verify <code>off</code> for the next boot.
sudo env SRC="$SRC" IMG="$IMG" \
bash -o pipefail -c '
    pigz -dc "$IMG" |
    cmp - "$SRC"
'


Then power off:
RC=$?


<syntaxhighlight lang="bash">
echo
for i in {1..10}; do echo; done
echo "source_image_cmp_rc=$RC"
sudo systemctl poweroff
</syntaxhighlight>
</syntaxhighlight>


After Wort is completely powered off:
<code>cmp</code> normally prints nothing when the inputs are identical.


# Remove the original T0 NVMe.
'''Required result:'''
# Label and retain it intact.
# Install the replacement NVMe.
# Do not connect the original NVMe simultaneously with its clone.
# Keep the dock recovery disk connected.


== Stage 7 — first protected boot with the replacement NVMe ==
<pre>
source_image_cmp_rc=0
</pre>


Wort should boot from T0h even though T0 is currently absent from the blank replacement NVMe.
'''STOP''' if the comparison is anything other than zero.
 
=== Record a checksum of the compressed recovery artifact ===
 
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"
 
sudo sha256sum "$IMG" |
sudo tee "$IMG.sha256"
 
sudo ls -lh "$IMG" "$IMG.sha256"
</syntaxhighlight>
 
At this point there are two independent recovery assets:


=== Immediately protect the following boot ===
# the untouched original physical NVMe; and
# the byte-verified compressed whole-device image.


The <code>off</code> request that produced this boot has now been consumed.
== Stage 6 — power off Wort and replace the NVMe ==


The first maintenance action is therefore:
Re-arm protected startup immediately before shutdown, even though it should already be armed.


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control status
sudo /usr/local/sbin/kitsnet-vm-startup-control status
echo
echo "running_guests=$(sudo virsh list --state-running --name 2>/dev/null | sed '/^[[:space:]]*$/d' | wc -l)"
</syntaxhighlight>
</syntaxhighlight>


'''Required result:'''
Verify <code>off</code> for the next boot.


<pre>
Then power off:
running_guests=0
</pre>
 
The following boot must also show <code>off</code> as armed.
 
=== Identify the replacement NVMe ===
 
Do not assume that the Linux device name is correct merely because it is <code>nvme0n1</code>.


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo lsblk -d -o NAME,PATH,SIZE,MODEL,SERIAL,WWN,LOG-SEC,PHY-SEC
sudo systemctl poweroff
 
echo
if command -v nvme >/dev/null 2>&1; then
    sudo nvme list
fi
</syntaxhighlight>
</syntaxhighlight>


Physically correlate model, serial number and capacity with the installed replacement.
After Wort is completely powered off:


Only after positive identification set:
# Remove the original T0 NVMe.
# Label and retain it intact.
# Install the replacement NVMe.
# Do not connect the original NVMe simultaneously with its clone.
# Keep the dock recovery disk connected.


<syntaxhighlight lang="bash">
== Stage 7 — first protected boot with the replacement NVMe ==
for i in {1..10}; do echo; done
NEW=/dev/nvme0n1


sudo lsblk -o NAME,PATH,SIZE,TYPE,FSTYPE,MOUNTPOINTS,MODEL,SERIAL,WWN,LOG-SEC,PHY-SEC "$NEW"
Wort should boot from T0h even though T0 is currently absent from the blank replacement NVMe.
</syntaxhighlight>
 
=== Immediately protect the following boot ===


'''STOP''' if there is any uncertainty regarding the replacement device.
The <code>off</code> request that produced this boot has now been consumed.


=== Verify replacement capacity and logical sector size ===
The first maintenance action is therefore:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
MIG="$DOCK/wort-T0-nvme-migration"
echo
NEW=/dev/nvme0n1
sudo /usr/local/sbin/kitsnet-vm-startup-control status
echo
echo "running_guests=$(sudo virsh list --state-running --name 2>/dev/null | sed '/^[[:space:]]*$/d' | wc -l)"
</syntaxhighlight>
 
'''Required result:'''


OLD_BYTES=$(sudo cat "$MIG/source.bytes")
<pre>
OLD_LOGSEC=$(sudo cat "$MIG/source.logical-sector-size")
running_guests=0
</pre>


NEW_BYTES=$(sudo blockdev --getsize64 "$NEW")
The following boot must also show <code>off</code> as armed.
NEW_LOGSEC=$(sudo blockdev --getss "$NEW")


echo "old_bytes=$OLD_BYTES"
=== Identify the replacement NVMe ===
echo "new_bytes=$NEW_BYTES"
 
echo
Do not assume that the Linux device name is correct merely because it is <code>nvme0n1</code>.
echo "old_logical_sector_size=$OLD_LOGSEC"
 
echo "new_logical_sector_size=$NEW_LOGSEC"
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
sudo lsblk -d -o NAME,PATH,SIZE,MODEL,SERIAL,WWN,LOG-SEC,PHY-SEC


echo
echo
if [ "$NEW_BYTES" -ge "$OLD_BYTES" ]; then
if command -v nvme >/dev/null 2>&1; then
     echo "PASS: replacement is large enough"
     sudo nvme list
else
    echo "STOP: replacement is smaller than source"
fi
 
if [ "$NEW_LOGSEC" -eq "$OLD_LOGSEC" ]; then
    echo "PASS: logical sector sizes match"
else
    echo "STOP: logical sector size mismatch"
fi
fi
</syntaxhighlight>
</syntaxhighlight>


Both tests must pass.
Physically correlate model, serial number and capacity with the installed replacement.


'''A logical-sector-size mismatch is a STOP condition for this whole-device cloning procedure.'''
Only after positive identification set:
 
=== Verify that nothing on the replacement NVMe is mounted ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
Line 903: Line 1,061:
NEW=/dev/nvme0n1
NEW=/dev/nvme0n1


echo "===== REPLACEMENT DEVICE TREE ====="
sudo lsblk -o NAME,PATH,SIZE,TYPE,FSTYPE,MOUNTPOINTS,MODEL,SERIAL,WWN,LOG-SEC,PHY-SEC "$NEW"
sudo lsblk -o NAME,PATH,SIZE,TYPE,FSTYPE,MOUNTPOINTS "$NEW"
 
echo
echo "===== NONEMPTY MOUNTPOINTS ====="
sudo lsblk -nr -o MOUNTPOINTS "$NEW" |
sed '/^[[:space:]]*$/d'
</syntaxhighlight>
</syntaxhighlight>


The second section must be empty.
'''STOP''' if there is any uncertainty regarding the replacement device.


'''STOP''' if any replacement-NVMe filesystem is mounted.
=== Verify replacement capacity and logical sector size ===
 
=== Verify the recovery image again before destructive writing ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
MIG="$DOCK/wort-T0-nvme-migration"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"
NEW=/dev/nvme0n1


sudo sha256sum -c "$IMG.sha256"
OLD_BYTES=$(sudo cat "$MIG/source.bytes")
RC=$?
OLD_LOGSEC=$(sudo cat "$MIG/source.logical-sector-size")


echo "image_sha256_rc=$RC"
NEW_BYTES=$(sudo blockdev --getsize64 "$NEW")
</syntaxhighlight>
NEW_LOGSEC=$(sudo blockdev --getss "$NEW")


'''Required result:'''
echo "old_bytes=$OLD_BYTES"
echo "new_bytes=$NEW_BYTES"
echo
echo "old_logical_sector_size=$OLD_LOGSEC"
echo "new_logical_sector_size=$NEW_LOGSEC"


<pre>
echo
image_sha256_rc=0
if [ "$NEW_BYTES" -ge "$OLD_BYTES" ]; then
</pre>
    echo "PASS: replacement is large enough"
else
    echo "STOP: replacement is smaller than source"
fi


== Stage 8 — restore the whole-device image to the replacement NVMe ==
if [ "$NEW_LOGSEC" -eq "$OLD_LOGSEC" ]; then
    echo "PASS: logical sector sizes match"
else
    echo "STOP: logical sector size mismatch"
fi
</syntaxhighlight>


=== Remove stale partition-table structures from the replacement only ===
Both tests must pass.


This prevents a previously used replacement device from retaining an unrelated backup GPT at its physical end.
'''A logical-sector-size mismatch is a STOP condition for this whole-device cloning procedure.'''


This command is destructive.
=== Verify that nothing on the replacement NVMe is mounted ===
 
Run it '''only''' after positive identification of <code>NEW</code>.


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
Line 950: Line 1,110:
NEW=/dev/nvme0n1
NEW=/dev/nvme0n1


echo "ABOUT TO DESTROY EXISTING PARTITION TABLES ON:"
echo "===== REPLACEMENT DEVICE TREE ====="
sudo lsblk -d -o NAME,PATH,SIZE,MODEL,SERIAL,WWN "$NEW"
sudo lsblk -o NAME,PATH,SIZE,TYPE,FSTYPE,MOUNTPOINTS "$NEW"


echo
echo
read -r -p 'Type WIPE_REPLACEMENT_NVME to continue: ' ANSWER
echo "===== NONEMPTY MOUNTPOINTS ====="
sudo lsblk -nr -o MOUNTPOINTS "$NEW" |
sed '/^[[:space:]]*$/d'
</syntaxhighlight>


if [ "$ANSWER" = "WIPE_REPLACEMENT_NVME" ]; then
The second section must be empty.
    sudo sgdisk --zap-all "$NEW"
    echo "sgdisk zap rc=$?"
else
    echo "ABORTED: replacement NVMe was not modified"
fi
</syntaxhighlight>


Do not continue unless the zap operation completed successfully.
'''STOP''' if any replacement-NVMe filesystem is mounted.


=== Restore the image ===
=== Verify the recovery image again before destructive writing ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
MIG="$DOCK/wort-T0-nvme-migration"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"
NEW=/dev/nvme0n1


echo "IMAGE=$IMG"
sudo sha256sum -c "$IMG.sha256"
echo "DESTINATION=$NEW"
RC=$?


echo
echo "image_sha256_rc=$RC"
sudo lsblk -d -o NAME,PATH,SIZE,MODEL,SERIAL,WWN "$NEW"
</syntaxhighlight>


echo
'''Required result:'''
read -r -p 'Type RESTORE_T0_IMAGE to overwrite the replacement NVMe: ' ANSWER


if [ "$ANSWER" = "RESTORE_T0_IMAGE" ]; then
<pre>
    sudo env IMG="$IMG" NEW="$NEW" \
image_sha256_rc=0
    bash -o pipefail -c '
</pre>
        pigz -dc "$IMG" |
        dd of="$NEW" bs=16M iflag=fullblock status=progress conv=fsync
    '


    RC=$?
== Stage 8 — restore the whole-device image to the replacement NVMe ==
    echo
    echo "restore_pipeline_rc=$RC"
    sudo sync
else
    echo "ABORTED: image was not restored"
fi
</syntaxhighlight>


'''Required result:'''
=== Remove stale partition-table structures from the replacement only ===


<pre>
This prevents a previously used replacement device from retaining an unrelated backup GPT at its physical end.
restore_pipeline_rc=0
</pre>


'''STOP''' otherwise.
This command is destructive.


=== Byte-verify the restored portion of the larger NVMe ===
Run it '''only''' after positive identification of <code>NEW</code>.
 
Because the replacement is larger, compare exactly the number of bytes that existed on the original NVMe.


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"
NEW=/dev/nvme0n1
NEW=/dev/nvme0n1
OLD_BYTES=$(sudo cat "$MIG/source.bytes")


sudo env IMG="$IMG" NEW="$NEW" OLD_BYTES="$OLD_BYTES" \
echo "ABOUT TO DESTROY EXISTING PARTITION TABLES ON:"
bash -o pipefail -c '
sudo lsblk -d -o NAME,PATH,SIZE,MODEL,SERIAL,WWN "$NEW"
    pigz -dc "$IMG" |
    cmp -n "$OLD_BYTES" - "$NEW"
'


RC=$?
echo
read -r -p 'Type WIPE_REPLACEMENT_NVME to continue: ' ANSWER


echo
if [ "$ANSWER" = "WIPE_REPLACEMENT_NVME" ]; then
echo "restored_image_cmp_rc=$RC"
    sudo sgdisk --zap-all "$NEW"
    echo "sgdisk zap rc=$?"
else
    echo "ABORTED: replacement NVMe was not modified"
fi
</syntaxhighlight>
</syntaxhighlight>


'''Required result:'''
Do not continue unless the zap operation completed successfully.


<pre>
=== Restore the image ===
restored_image_cmp_rc=0
</pre>


This proves that every byte belonging to the original disk image was written identically to the replacement.
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"
NEW=/dev/nvme0n1


'''STOP''' otherwise.
echo "IMAGE=$IMG"
echo "DESTINATION=$NEW"


=== Do not expand the disk yet ===
echo
sudo lsblk -d -o NAME,PATH,SIZE,MODEL,SERIAL,WWN "$NEW"


At this point the replacement contains an exact clone of the old NVMe.
echo
read -r -p 'Type RESTORE_T0_IMAGE to overwrite the replacement NVMe: ' ANSWER


If the source uses GPT, a verification command may report that the backup GPT is not at the physical end of the larger device. That condition is expected at this stage.
if [ "$ANSWER" = "RESTORE_T0_IMAGE" ]; then
    sudo env IMG="$IMG" NEW="$NEW" \
    bash -o pipefail -c '
        pigz -dc "$IMG" |
        dd of="$NEW" bs=16M iflag=fullblock status=progress conv=fsync
    '


No correction or expansion is performed until after a clean reboot proves that LVM sees the clone correctly.
    RC=$?
    echo
    echo "restore_pipeline_rc=$RC"
    sudo sync
else
    echo "ABORTED: image was not restored"
fi
</syntaxhighlight>


=== Reboot protected ===
'''Required result:'''
 
<pre>
restore_pipeline_rc=0
</pre>


The following boot should already be armed <code>off</code> from the beginning of this maintenance boot. Verify before rebooting:
'''STOP''' otherwise.


<syntaxhighlight lang="bash">
=== Byte-verify the restored portion of the larger NVMe ===
for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control status
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
sudo /usr/local/sbin/kitsnet-vm-startup-control status
</syntaxhighlight>


Then:
Because the replacement is larger, compare exactly the number of bytes that existed on the original NVMe.


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo systemctl reboot
DOCK=/mnt/KitsNet-pbd2
</syntaxhighlight>
MIG="$DOCK/wort-T0-nvme-migration"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"
NEW=/dev/nvme0n1
OLD_BYTES=$(sudo cat "$MIG/source.bytes")
 
sudo env IMG="$IMG" NEW="$NEW" OLD_BYTES="$OLD_BYTES" \
bash -o pipefail -c '
    pigz -dc "$IMG" |
    cmp -n "$OLD_BYTES" - "$NEW"
'
 
RC=$?


== Stage 9 — prove the exact clone before expansion ==
=== Immediately protect the following boot ===
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
echo
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control status
echo "restored_image_cmp_rc=$RC"
echo
echo "running_guests=$(sudo virsh list --state-running --name 2>/dev/null | sed '/^[[:space:]]*$/d' | wc -l)"
</syntaxhighlight>
</syntaxhighlight>


'''Required:'''
'''Required result:'''


<pre>
<pre>
running_guests=0
restored_image_cmp_rc=0
</pre>
</pre>


=== Let normal boot discovery run first ===
This proves that every byte belonging to the original disk image was written identically to the replacement.
 
'''STOP''' otherwise.


Do not issue LVM repair commands preemptively.
=== Do not expand the disk yet ===


Begin with read-only inspection:
At this point the replacement contains an exact clone of the old NVMe.


<syntaxhighlight lang="bash">
If the source uses GPT, a verification command may report that the backup GPT is not at the physical end of the larger device. That condition is expected at this stage.
for i in {1..10}; do echo; done
sudo pvs
echo
sudo vgs
echo
sudo lvs T0 2>&1 || true
</syntaxhighlight>


The expected outcome is that <code>/dev/nvme0n1p1</code> appears as the existing T0 PV.
No correction or expansion is performed until after a clean reboot proves that LVM sees the clone correctly.


If T0 appears normally, '''do not run any additional scan/refresh command'''.
=== Reboot protected ===


=== Only if T0 is missing: refresh the LVM device identity ===
The following boot should already be armed <code>off</code> from the beginning of this maintenance boot. Verify before rebooting:
 
The replacement NVMe has a different physical WWID/serial number. If Wort uses <code>/etc/lvm/devices/system.devices</code>, the old hardware ID may prevent immediate discovery even though the cloned PVID is correct.
 
First inspect:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo lvmconfig --type current devices/use_devicesfile
sudo /usr/local/sbin/kitsnet-vm-startup-control status
echo
echo
sudo lvmdevices 2>&1 || true
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
sudo /usr/local/sbin/kitsnet-vm-startup-control status
</syntaxhighlight>
</syntaxhighlight>


If <code>use_devicesfile=1</code>, perform the supported device-ID refresh:
Then:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo lvmdevices --check --refresh
sudo systemctl reboot
echo
</syntaxhighlight>
sudo lvmdevices --update --refresh
RC=$?


echo
== Stage 9 — prove the exact clone before expansion ==
echo "lvmdevices_refresh_rc=$RC"


echo
=== Immediately protect the following boot ===
sudo pvscan --cache /dev/nvme0n1p1


<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control status
echo
echo
sudo pvs
echo "running_guests=$(sudo virsh list --state-running --name 2>/dev/null | sed '/^[[:space:]]*$/d' | wc -l)"
</syntaxhighlight>
</syntaxhighlight>


If the devices file is not in use, a normal cache refresh may be used:
'''Required:'''
 
<pre>
running_guests=0
</pre>
 
=== Let normal boot discovery run first ===
 
Do not issue LVM repair commands preemptively.
 
Begin with read-only inspection:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo pvscan --cache /dev/nvme0n1p1
sudo pvs
echo
echo
sudo pvs
sudo vgs
echo
sudo lvs T0 2>&1 || true
</syntaxhighlight>
</syntaxhighlight>


'''STOP''' if T0 still does not appear.
The expected outcome is that <code>/dev/nvme0n1p1</code> appears as the existing T0 PV.
 
If T0 appears normally, '''do not run any additional scan/refresh command'''.


Do not initialize or repair the PV merely because discovery failed.
=== If T0 is missing: prove the clone first, then update the LVM devices file ===


=== Check LVM metadata without repairing it ===
The completed migration encountered this exact condition. The cloned partition contained the correct PVID, but <code>/etc/lvm/devices/system.devices</code> still associated that PVID with the old Plextor WWID. Normal discovery therefore omitted T0 even though the clone itself was correct.
 
First inspect whether the devices file is in use:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
PV=/dev/nvme0n1p1
sudo lvmconfig --type current devices/use_devicesfile
echo
sudo lvmdevices 2>&1 || true
</syntaxhighlight>


sudo pvck "$PV"
Before updating the devices file, use an explicit device override to prove the cloned identities:
PVCK_RC=$?


<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
PV=/dev/nvme0n1p1
sudo pvs --devices "$PV" -o pv_name,pv_uuid,vg_name,pv_size,pv_free "$PV"
echo
echo
sudo vgck T0
sudo vgs --devices "$PV" -o vg_name,vg_uuid,pv_count,lv_count,vg_size,vg_free T0
VGCK_RC=$?
 
echo
echo
echo "pvck_rc=$PVCK_RC"
sudo lvs --devices "$PV" -o vg_name,lv_name,lv_uuid,lv_size T0
echo "vgck_rc=$VGCK_RC"
</syntaxhighlight>
</syntaxhighlight>


Do not use <code>pvck --repair</code> or <code>vgck --updatemetadata</code> during normal migration validation.
The PVID, VG UUID and 27-LV inventory must match the captured baseline. '''STOP''' if they do not.


'''STOP''' on unexplained metadata errors.
If <code>use_devicesfile=1</code> and the clone identity is correct:
 
=== Compare the cloned PV UUID ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
sudo lvmdevices --check --refresh
MIG="$DOCK/wort-T0-nvme-migration"
CHECK_RC=$?
PV=/dev/nvme0n1p1
echo "lvmdevices_check_refresh_rc=$CHECK_RC"


sudo pvs --noheadings -o pv_uuid "$PV" |
echo
tr -d ' ' > /tmp/T0-pvid.after
sudo lvmdevices --update --refresh
UPDATE_RC=$?
echo "lvmdevices_update_refresh_rc=$UPDATE_RC"


echo "===== BEFORE ====="
echo
sudo cat "$MIG/pvid.before"
sudo lvmdevices --check
FINAL_RC=$?
echo "lvmdevices_final_check_rc=$FINAL_RC"


echo "===== AFTER ====="
echo
cat /tmp/T0-pvid.after
sudo pvs
sudo vgs T0
</syntaxhighlight>


echo "===== DIFF ====="
In the validated migration, <code>lvmdevices --check --refresh</code> returned rc 5 and reported that the PVID had a new <code>sys_wwid</code>. This was a diagnostic result indicating that <code>system.devices</code> needed an update. <code>lvmdevices --update --refresh</code> then replaced the Plextor WWID with <code>eui.00000000000000006479a7a11ac00487</code>, and the final <code>lvmdevices --check</code> returned zero.
sudo diff -u "$MIG/pvid.before" /tmp/T0-pvid.after
</syntaxhighlight>


'''Required:''' no difference.
Do not initialize or repair the PV merely because normal discovery initially fails.


=== Compare the cloned VG UUID ===
=== Check LVM metadata without repairing it ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
PV=/dev/nvme0n1p1
MIG="$DOCK/wort-T0-nvme-migration"
 
sudo pvck "$PV"
PVCK_RC=$?


sudo vgs --noheadings -o vg_uuid T0 |
echo
tr -d ' ' > /tmp/T0-vguid.after
sudo vgck T0
VGCK_RC=$?


sudo diff -u "$MIG/vguid.before" /tmp/T0-vguid.after
echo
echo "pvck_rc=$PVCK_RC"
echo "vgck_rc=$VGCK_RC"
</syntaxhighlight>
</syntaxhighlight>


'''Required:''' no difference.
Do not use <code>pvck --repair</code> or <code>vgck --updatemetadata</code> during normal migration validation.


=== Compare all LV UUIDs ===
'''STOP''' on unexplained metadata errors.
 
=== Compare the cloned PV UUID ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
MIG="$DOCK/wort-T0-nvme-migration"
PV=/dev/nvme0n1p1
sudo pvs --noheadings -o pv_uuid "$PV" |
tr -d ' ' > /tmp/T0-pvid.after


sudo lvs --noheadings --separator '|' -o lv_name,lv_uuid T0 |
echo "===== BEFORE ====="
sed 's/^[[:space:]]*//;s/[[:space:]]*$//' |
sudo cat "$MIG/pvid.before"
sort > /tmp/T0-lvuuids.after
 
echo "===== AFTER ====="
cat /tmp/T0-pvid.after


sudo diff -u "$MIG/lvuuids.before" /tmp/T0-lvuuids.after
echo "===== DIFF ====="
sudo diff -u "$MIG/pvid.before" /tmp/T0-pvid.after
</syntaxhighlight>
</syntaxhighlight>


'''Required:''' no difference.
'''Required:''' no difference.


=== Compare the partition UUID and start sector ===
=== Compare the cloned VG UUID ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
MIG="$DOCK/wort-T0-nvme-migration"
PV=/dev/nvme0n1p1


sudo blkid -s PARTUUID -o value "$PV" > /tmp/partuuid.after
sudo vgs --noheadings -o vg_uuid T0 |
tr -d ' ' > /tmp/T0-vguid.after


sudo lsblk -no START "$PV" |
sudo diff -u "$MIG/vguid.before" /tmp/T0-vguid.after
tr -d ' ' > /tmp/partition-start.after
</syntaxhighlight>


echo "===== PARTUUID DIFF ====="
'''Required:''' no difference.
sudo diff -u "$MIG/partuuid.before" /tmp/partuuid.after
 
=== Compare all LV UUIDs ===
 
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
 
sudo lvs --noheadings --separator '|' -o lv_name,lv_uuid T0 |
sed 's/^[[:space:]]*//;s/[[:space:]]*$//' |
sort > /tmp/T0-lvuuids.after
 
sudo diff -u "$MIG/lvuuids.before" /tmp/T0-lvuuids.after
</syntaxhighlight>
 
'''Required:''' no difference.
 
=== Compare the partition UUID and start sector ===
 
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
PV=/dev/nvme0n1p1
 
sudo blkid -s PARTUUID -o value "$PV" > /tmp/partuuid.after
 
sudo sfdisk --dump /dev/nvme0n1 2>/dev/null |
sed -n "\|^$PV |s/.*start=[[:space:]]*\([0-9][0-9]*\),.*/\1/p" > /tmp/partition-start.after
 
echo "===== PARTUUID DIFF ====="
sudo diff -u "$MIG/partuuid.before" /tmp/partuuid.after


echo
echo
Line 1,255: Line 1,471:
<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
MIG="$DOCK/wort-T0-nvme-migration"
NEW=/dev/nvme0n1
NEW=/dev/nvme0n1
Line 1,278: Line 1,494:
<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
MIG="$DOCK/wort-T0-nvme-migration"


Line 1,301: Line 1,517:
At this point the clone has been proven.
At this point the clone has been proven.


== Stage 10 — expand the replacement partition ==
== Stage 10 — expand partition 1 and the existing T0 PV ==


Before modifying the partition table, deactivate T0 again:
The executed maintenance proved that partition growth and <code>pvresize</code> can be completed in the same protected boot. The reboot between the original draft's Stage 10 and Stage 11 is removed.
 
=== Establish the actual safety conditions ===
 
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
echo "===== QEMU ====="
QEMU=$(pgrep -fc 'qemu-system|qemu-kvm' 2>/dev/null)
echo "qemu_processes=$QEMU"
 
echo
echo "===== T0 MOUNTS ====="
findmnt -rn -S '/dev/mapper/T0-*' || true
 
echo
echo "===== T0 DEVICE-MAPPER OPEN COUNTS ====="
sudo dmsetup info -c --noheadings -o name,open | awk '$1 ~ /^T0-/ {print}'
</syntaxhighlight>
 
'''Required:''' zero QEMU processes, no T0 filesystem mounts, and open count zero for every T0 device-mapper node.
 
T0 may autoactivate after partition/udev activity. If the safety conditions above remain true, autoactivation alone is not a STOP condition. Deactivate T0 before the partition write when practical:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
Line 1,309: Line 1,546:
sudo vgchange -an T0
sudo vgchange -an T0
RC=$?
RC=$?
echo "vgchange_deactivate_rc=$RC"
echo "vgchange_deactivate_rc=$RC"
</syntaxhighlight>


ACTIVE=$(
=== Relocate the GPT backup header ===
  sudo lvs --noheadings -o lv_active T0 |
  awk '$1 == "active" { n++ } END { print n+0 }'
)


echo "active_T0_LVs=$ACTIVE"
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
NEW=/dev/nvme0n1
sudo sgdisk -e "$NEW"
RC=$?
echo "sgdisk_move_header_rc=$RC"
echo
sudo sgdisk -v "$NEW"
</syntaxhighlight>
</syntaxhighlight>


'''Required:'''
'''Required:''' rc 0 and no GPT structural errors.
 
<pre>
vgchange_deactivate_rc=0
active_T0_LVs=0
</pre>


=== If the source is GPT, move the backup GPT to the real end of the replacement ===
=== Dry-run and perform partition growth ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
NEW=/dev/nvme0n1
NEW=/dev/nvme0n1
PTTYPE=$(sudo cat "$MIG/source.pttype")
PV=/dev/nvme0n1p1


if [ "$PTTYPE" = "gpt" ]; then
sudo growpart -N "$NEW" 1
    echo "Moving backup GPT to the actual end of $NEW"
DRY_RC=$?
    sudo sgdisk -e "$NEW"
echo "growpart_dry_run_rc=$DRY_RC"
    RC=$?
 
    echo "sgdisk_move_header_rc=$RC"
echo
sudo growpart "$NEW" 1
GROW_RC=$?
echo "growpart_rc=$GROW_RC"
 
sudo sync
sudo partprobe "$NEW"
sudo udevadm settle


    echo
echo
    sudo sgdisk -v "$NEW"
sudo sfdisk --dump "$NEW"
elif [ "$PTTYPE" = "dos" ]; then
    echo "Source uses DOS/MBR; no GPT backup-header relocation is required"
else
    echo "STOP: unexpected partition-table type: $PTTYPE"
fi
</syntaxhighlight>
</syntaxhighlight>


For GPT, <code>sgdisk_move_header_rc</code> must be zero.
The dry run must propose only extension of partition 1 while preserving start sector <code>2048</code>. On the validated Sabrent the final partition size was <code>4000795279</code> sectors and the last usable GPT LBA was <code>4000797326</code>.


=== Dry-run partition growth ===
=== Verify partition start and PARTUUID ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
MIG=/mnt/KitsNet-pbd2/wort-T0-nvme-migration
NEW=/dev/nvme0n1
NEW=/dev/nvme0n1
PV=/dev/nvme0n1p1


sudo growpart -N "$NEW" 1
EXPECTED_START=$(cat "$MIG/partition-start.before")
RC=$?
EXPECTED_PARTUUID=$(tr '[:lower:]' '[:upper:]' < "$MIG/partuuid.before")
CURRENT_START=$(sudo sfdisk --dump "$NEW" 2>/dev/null | sed -n "\|^$PV |s/.*start=[[:space:]]*\([0-9][0-9]*\),.*/\1/p")
CURRENT_PARTUUID=$(sudo blkid -s PARTUUID -o value "$PV" | tr '[:lower:]' '[:upper:]')


echo
echo "expected_start=$EXPECTED_START"
echo "growpart_dry_run_rc=$RC"
echo "current_start=$CURRENT_START"
echo "expected_partuuid=$EXPECTED_PARTUUID"
echo "current_partuuid=$CURRENT_PARTUUID"
</syntaxhighlight>
</syntaxhighlight>


The dry run should show that partition 1 can grow into the additional capacity.
The start sector must remain identical.
 
'''STOP''' if the proposed change is not exactly an extension of partition 1 with the original starting sector preserved.


=== Grow partition 1 ===
During the completed migration, the partition unique GUID was unexpectedly observed as <code>BA942590-8B70-463D-AA42-5E03D7A280E6</code> instead of the captured original <code>3703AD9C-AB50-486A-B154-E83338BC7D3F</code>. Because the disk GUID, partition contents, PVID, VG UUID, all 27 LV UUIDs, and kvm_save filesystem UUID were already proven, the original partition GUID was restored:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
NEW=/dev/nvme0n1
sudo sgdisk -u 1:3703AD9C-AB50-486A-B154-E83338BC7D3F /dev/nvme0n1
 
sudo growpart "$NEW" 1
RC=$?
RC=$?
echo "partition_guid_restore_rc=$RC"
sudo sync
sudo partprobe /dev/nvme0n1
sudo udevadm settle
sudo sgdisk -v /dev/nvme0n1
sudo blkid -s PARTUUID -o value /dev/nvme0n1p1
</syntaxhighlight>


echo
Only restore a PARTUUID when the exact original value was captured before migration and all other clone identities have already been proven. Never generate a new identity as part of this procedure.
echo "growpart_rc=$RC"


echo
=== Resize the existing PV without an intervening reboot ===
sudo sfdisk --dump "$NEW"
</syntaxhighlight>


'''Required:'''
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
PV=/dev/nvme0n1p1
sudo pvresize "$PV"
RC=$?
echo "pvresize_rc=$RC"
echo
sudo pvs -o pv_name,pv_uuid,pv_size,pv_free "$PV"
sudo vgs -o vg_name,vg_uuid,pv_count,lv_count,vg_size,vg_free T0
</syntaxhighlight>


<pre>
'''Required:''' <code>pvresize_rc=0</code>. The successful run ended with T0 at approximately <code>1.86t</code> total and <code>1.50t</code> free. The LVM reports had already begun showing the larger geometry immediately before the explicit <code>pvresize</code>; that is not a failure.
growpart_rc=0
</pre>


=== Verify that the partition start and PARTUUID did not change ===
=== Revalidate GPT and all LVM identities ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG=/mnt/KitsNet-pbd2/wort-T0-nvme-migration
MIG="$DOCK/wort-T0-nvme-migration"
PV=/dev/nvme0n1p1
PV=/dev/nvme0n1p1


sudo lsblk -no START "$PV" |
sudo sgdisk -v /dev/nvme0n1
tr -d ' ' > /tmp/partition-start.grown
sudo pvck "$PV"
sudo vgck T0


sudo blkid -s PARTUUID -o value "$PV" > /tmp/partuuid.grown
sudo pvs --noheadings -o pv_uuid "$PV" | tr -d ' ' > /tmp/T0-pvid.final
sudo vgs --noheadings -o vg_uuid T0 | tr -d ' ' > /tmp/T0-vguid.final
sudo lvs --noheadings --separator '|' -o lv_name,lv_uuid T0 |
  sed 's/^[[:space:]]*//;s/[[:space:]]*$//' |
  sort > /tmp/T0-lvuuids.final


echo "===== START-SECTOR DIFF ====="
sudo diff -u "$MIG/pvid.before" /tmp/T0-pvid.final
sudo diff -u "$MIG/partition-start.before" /tmp/partition-start.grown
sudo diff -u "$MIG/vguid.before" /tmp/T0-vguid.final
sudo diff -u "$MIG/lvuuids.before" /tmp/T0-lvuuids.final
</syntaxhighlight>


echo
All comparisons must be identical.
echo "===== PARTUUID DIFF ====="
sudo diff -u "$MIG/partuuid.before" /tmp/partuuid.grown
</syntaxhighlight>


Both must show no difference.
== Stage 11 — restore the libvirt save filesystem ==


For GPT, verify structure again:
=== Activate T0 ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
sudo vgchange -ay T0
MIG="$DOCK/wort-T0-nvme-migration"
RC=$?
NEW=/dev/nvme0n1
PTTYPE=$(sudo cat "$MIG/source.pttype")


if [ "$PTTYPE" = "gpt" ]; then
echo "vgchange_activate_rc=$RC"
    sudo sgdisk -v "$NEW"
 
fi
echo
sudo lvs -o vg_name,lv_name,lv_active,lv_attr,lv_size T0
</syntaxhighlight>
</syntaxhighlight>


'''Do not run <code>pvresize</code> yet.'''
'''Required:'''


The next reboot establishes the enlarged partition geometry from a clean kernel/device discovery cycle.
<pre>
vgchange_activate_rc=0
</pre>


=== Reboot protected ===
=== Restore only the fstab line that this procedure commented ===


<syntaxhighlight lang="bash">
First verify that the marker exists exactly as expected:
for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
sudo /usr/local/sbin/kitsnet-vm-startup-control status
</syntaxhighlight>
 
Then:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo systemctl reboot
grep -n '^# KN-T0-NVME-MIGRATION ' /etc/fstab
</syntaxhighlight>
</syntaxhighlight>


== Stage 11 — resize the existing T0 PV ==
There should be exactly one line, and it must be the original <code>/var/lib/libvirt/qemu/save</code> entry.


=== Immediately protect the following boot ===
Restore it:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
sudo sed -i \
  's|^# KN-T0-NVME-MIGRATION ||' \
  /etc/fstab


echo
echo "===== RESTORED ENTRY ====="
sudo /usr/local/sbin/kitsnet-vm-startup-control status
grep -n '/var/lib/libvirt/qemu/save' /etc/fstab


echo
echo
echo "running_guests=$(sudo virsh list --state-running --name 2>/dev/null | sed '/^[[:space:]]*$/d' | wc -l)"
sudo systemctl daemon-reload
sudo findmnt --verify --verbose
</syntaxhighlight>
</syntaxhighlight>


'''Required:'''
=== Mount the save filesystem manually ===
 
<pre>
running_guests=0
</pre>
 
=== Confirm the kernel sees the enlarged partition ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
NEW=/dev/nvme0n1
sudo mount /var/lib/libvirt/qemu/save
PV=/dev/nvme0n1p1
RC=$?


sudo lsblk -o NAME,PATH,SIZE,TYPE,FSTYPE,PTTYPE,START,LOG-SEC,PHY-SEC "$NEW"
echo "mount rc=$RC"


echo
echo
sudo pvs -o pv_name,pv_uuid,vg_name,pv_size,pv_free "$PV"
sudo findmnt --mountpoint /var/lib/libvirt/qemu/save \
  -o SOURCE,TARGET,FSTYPE,SIZE,USED,AVAIL,OPTIONS
</syntaxhighlight>


echo
'''Required:'''
sudo vgs -o vg_name,vg_uuid,vg_size,vg_free T0
</syntaxhighlight>


The partition should now show the new size.
<pre>
mount rc=0
</pre>


The LVM PV may still show the original size. That is expected until <code>pvresize</code> is executed.
The source must resolve to the restored <code>T0/kvm_save</code> LV.


=== Deactivate T0 before changing PV geometry ===
=== Verify the filesystem UUID ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo vgchange -an T0
DOCK=/mnt/KitsNet-pbd2
RC=$?
MIG="$DOCK/wort-T0-nvme-migration"
 
sudo blkid -s UUID -o value /dev/T0/kvm_save > /tmp/kvm-save-fsuuid.final


echo "vgchange_deactivate_rc=$RC"
sudo diff -u \
  "$MIG/kvm-save-fsuuid.before" \
  /tmp/kvm-save-fsuuid.final
</syntaxhighlight>
</syntaxhighlight>


'''Required:'''
'''Required:''' no difference.


<pre>
=== Verify the managed-save inventory survived ===
vgchange_deactivate_rc=0
</pre>


=== Resize the existing PV ===
Verify that the recovery filesystem is mounted before reading the baseline.
 
This changes the usable size of the existing PVID. It does not create or replace the PV.


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
PV=/dev/nvme0n1p1
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
SAVE=/var/lib/libvirt/qemu/save
 
findmnt --mountpoint "$DOCK"
RECOVERY_RC=$?
echo "recovery_mount_rc=$RECOVERY_RC"


sudo pvresize "$PV"
sudo find "$SAVE" \
RC=$?
  -maxdepth 1 \
  -type f \
  -name '*.save' \
  -printf '%f|%s\n' |
sort > /tmp/managed-save.current


echo
EXPECTED_COUNT=$(awk 'END {print NR}' "$MIG/managed-save.before")
echo "pvresize_rc=$RC"
CURRENT_COUNT=$(awk 'END {print NR}' /tmp/managed-save.current)


echo
echo "expected_save_count=$EXPECTED_COUNT"
sudo pvs -o pv_name,pv_uuid,vg_name,pv_size,pv_free "$PV"
echo "current_save_count=$CURRENT_COUNT"


echo
diff -u "$MIG/managed-save.before" /tmp/managed-save.current
sudo vgs -o vg_name,vg_uuid,vg_size,vg_free T0
DIFF_RC=$?
echo "managed_save_manifest_diff_rc=$DIFF_RC"
</syntaxhighlight>
</syntaxhighlight>


'''Required:'''
For the validated execution the expected and current counts were both <code>18</code> and the diff returned zero.


<pre>
Do not copy the baseline to <code>/root</code> and then use unprivileged shell redirection such as <code>wc -l &lt; /root/file</code>; that produced a false STOP during this maintenance because the shell opens the file before <code>sudo</code> can affect a child command.
pvresize_rc=0
 
</pre>
== Stage 12 — final no-VM validation reboot ==
 
This reboot tests the final intended storage configuration:
 
* replacement NVMe installed;
* clone proven;
* partition expanded;
* T0 PV expanded;
* T0 identities unchanged;
* fstab restored; and
* <code>/var/lib/libvirt/qemu/save</code> restored.


T0 should now show substantial new free space corresponding to the larger NVMe.
No VM is yet permitted to start.


=== Revalidate LVM metadata and identity ===
=== Arm the protected boot ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
MIG="$DOCK/wort-T0-nvme-migration"
sudo /usr/local/sbin/kitsnet-vm-startup-control status
PV=/dev/nvme0n1p1
</syntaxhighlight>


sudo pvck "$PV"
Then:
echo
sudo vgck T0


echo
<syntaxhighlight lang="bash">
sudo pvs --noheadings -o pv_uuid "$PV" |
for i in {1..10}; do echo; done
tr -d ' ' > /tmp/T0-pvid.final
sudo systemctl reboot
 
</syntaxhighlight>
sudo vgs --noheadings -o vg_uuid T0 |
tr -d ' ' > /tmp/T0-vguid.final


sudo lvs --noheadings --separator '|' -o lv_name,lv_uuid T0 |
== Stage 13 — final validation with all VMs still off ==
sed 's/^[[:space:]]*//;s/[[:space:]]*$//' |
sort > /tmp/T0-lvuuids.final


echo "===== PVID ====="
=== Immediately preserve protection against an unexpected additional reboot ===
sudo diff -u "$MIG/pvid.before" /tmp/T0-pvid.final


<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
echo
echo
echo "===== VG UUID ====="
sudo /usr/local/sbin/kitsnet-vm-startup-control status
sudo diff -u "$MIG/vguid.before" /tmp/T0-vguid.final
 
echo
echo "===== LV UUIDS ====="
sudo diff -u "$MIG/lvuuids.before" /tmp/T0-lvuuids.final
</syntaxhighlight>
</syntaxhighlight>


All UUID comparisons must remain identical.
=== Verify the protected no-VM state ===


== Stage 12 — restore the libvirt save filesystem ==
A protected <code>off</code> boot intentionally runtime-masks <code>libvirt-guests.service</code> and the <code>virtqemud</code>/<code>libvirtd</code> sockets. <code>virsh</code> can therefore fail with a missing or masked socket even though the boot is correctly protected.


=== Activate T0 ===
Do not unmask or start those services for validation. Use the controller state and QEMU process count:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo vgchange -ay T0
sudo /usr/local/sbin/kitsnet-vm-startup-control status
RC=$?


echo "vgchange_activate_rc=$RC"
echo
QEMU=$(pgrep -fc 'qemu-system|qemu-kvm' 2>/dev/null)
echo "qemu_processes=$QEMU"


echo
echo
sudo lvs -o vg_name,lv_name,lv_active,lv_attr,lv_size T0
systemctl status libvirt-guests.service --no-pager 2>&1 || true
systemctl list-unit-files --no-pager |
  grep -E '^(virtqemud|libvirtd|libvirt-guests)\.(service|socket)' || true
</syntaxhighlight>
</syntaxhighlight>


'''Required:'''
'''Required:''' <code>current_boot_mode=off</code>, <code>current_boot_protection=PROTECTED</code>, and <code>qemu_processes=0</code>. Runtime-masked libvirt units are expected here.


<pre>
=== Verify final T0 configuration ===
vgchange_activate_rc=0
</pre>


=== Restore only the fstab line that this procedure commented ===
<syntaxhighlight lang="bash">
 
First verify that the marker exists exactly as expected:
 
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
grep -n '^# KN-T0-NVME-MIGRATION ' /etc/fstab
sudo pvs -o pv_name,pv_uuid,vg_name,pv_size,pv_free
echo
sudo vgs -o vg_name,vg_uuid,pv_count,lv_count,vg_size,vg_free T0
echo
sudo lvs -o vg_name,lv_name,lv_uuid,lv_size,lv_attr T0
</syntaxhighlight>
</syntaxhighlight>


There should be exactly one line, and it must be the original <code>/var/lib/libvirt/qemu/save</code> entry.
Confirm:
 
* T0 has exactly one PV;
* the PVID is unchanged;
* the VG UUID is unchanged;
* all expected LVs exist;
* their UUIDs are unchanged; and
* the additional replacement-NVMe space now appears as T0 free space.


Restore it:
=== Verify the save mount survived a clean boot ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo sed -i \
sudo findmnt --mountpoint /var/lib/libvirt/qemu/save \
  's|^# KN-T0-NVME-MIGRATION ||' \
  -o SOURCE,TARGET,FSTYPE,SIZE,USED,AVAIL,OPTIONS
  /etc/fstab
 
echo "===== RESTORED ENTRY ====="
grep -n '/var/lib/libvirt/qemu/save' /etc/fstab


echo
echo
sudo systemctl daemon-reload
sudo find /var/lib/libvirt/qemu/save \
sudo findmnt --verify --verbose
  -maxdepth 1 \
  -type f \
  -name '*.save' \
  -printf '%f|%s\n' |
sort
</syntaxhighlight>
</syntaxhighlight>


=== Mount the save filesystem manually ===
=== Reconfirm the managed-save manifest after the protected reboot ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo mount /var/lib/libvirt/qemu/save
MIG=/mnt/KitsNet-pbd2/wort-T0-nvme-migration
RC=$?
SAVE=/var/lib/libvirt/qemu/save


echo "mount rc=$RC"
sudo find "$SAVE" -maxdepth 1 -type f -name '*.save' -printf '%f|%s\n' |
  sort > /tmp/wort-T0-managed-save.after-reboot


echo
EXPECTED=$(awk 'END {print NR}' "$MIG/managed-save.before")
sudo findmnt --mountpoint /var/lib/libvirt/qemu/save \
CURRENT=$(awk 'END {print NR}' /tmp/wort-T0-managed-save.after-reboot)
  -o SOURCE,TARGET,FSTYPE,SIZE,USED,AVAIL,OPTIONS
</syntaxhighlight>


'''Required:'''
diff -u "$MIG/managed-save.before" /tmp/wort-T0-managed-save.after-reboot
DIFF_RC=$?


<pre>
echo "expected_save_count=$EXPECTED"
mount rc=0
echo "current_save_count=$CURRENT"
</pre>
echo "managed_save_diff_rc=$DIFF_RC"
</syntaxhighlight>


The source must resolve to the restored <code>T0/kvm_save</code> LV.
For the validated maintenance the required result was <code>18</code>, <code>18</code>, and diff rc <code>0</code>.


=== Verify the filesystem UUID ===
=== Final metadata checks ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
sudo pvck /dev/nvme0n1p1
MIG="$DOCK/wort-T0-nvme-migration"
PVCK_RC=$?


sudo blkid -s UUID -o value /dev/T0/kvm_save > /tmp/kvm-save-fsuuid.final
echo
sudo vgck T0
VGCK_RC=$?


sudo diff -u \
echo
  "$MIG/kvm-save-fsuuid.before" \
echo "pvck_rc=$PVCK_RC"
  /tmp/kvm-save-fsuuid.final
echo "vgck_rc=$VGCK_RC"
</syntaxhighlight>
</syntaxhighlight>


'''Required:''' no difference.
=== Check host health ===
 
=== Verify the managed-save inventory survived ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
echo "===== FAILED SYSTEMD UNITS ====="
MIG="$DOCK/wort-T0-nvme-migration"
systemctl --failed --no-pager


sudo find /var/lib/libvirt/qemu/save \
echo
  -maxdepth 1 \
echo "===== T0 ====="
  -type f \
sudo pvs /dev/nvme0n1p1
  -printf '%f|%s\n' |
sudo vgs T0
sort > /tmp/managed-save.final


sudo diff -u \
echo
  "$MIG/managed-save.before" \
echo "===== SAVE MOUNT ====="
  /tmp/managed-save.final
sudo findmnt --mountpoint /var/lib/libvirt/qemu/save
</syntaxhighlight>


'''Required:''' no difference.
echo
echo "===== STARTUP CONTROL ====="
sudo /usr/local/sbin/kitsnet-vm-startup-control status
</syntaxhighlight>


== Stage 13 — final no-VM validation reboot ==
Do not return to production if there is an unexplained storage, LVM, mount or host failure.


This reboot tests the final intended storage configuration:
== Stage 14 — return KitsNet to normal automatic startup ==


* replacement NVMe installed;
Only after every preceding validation has passed should the next boot be changed from <code>off</code> to <code>auto</code>.
* clone proven;
* partition expanded;
* T0 PV expanded;
* T0 identities unchanged;
* fstab restored; and
* <code>/var/lib/libvirt/qemu/save</code> restored.


No VM is yet permitted to start.
=== Explicitly select normal automatic startup ===
 
=== Arm the protected boot ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot auto
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control status
sudo /usr/local/sbin/kitsnet-vm-startup-control status
</syntaxhighlight>
</syntaxhighlight>


Then:
Verify that the next boot is configured for normal automatic controlled startup.
 
=== Perform the production boot ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
Line 1,706: Line 1,969:
</syntaxhighlight>
</syntaxhighlight>


== Stage 14 — final validation with all VMs still off ==
The normal Wort controller now owns startup sequencing.


=== Immediately preserve protection against an unexpected additional reboot ===
Do not manually start Fox, Anchor, mgr1, workers, ordinary VMs or socat listeners unless the established startup-control procedure reports a failure requiring diagnosis.
 
== Stage 15 — post-return production validation ==


<syntaxhighlight lang="bash">
=== Startup-controller status ===
 
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
sudo /usr/local/sbin/kitsnet-vm-startup-control status
 
echo
systemctl status kitsnet-vm-auto-start.service --no-pager -l
 
echo
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control status
sudo journalctl -b -u kitsnet-vm-auto-start.service --no-pager
</syntaxhighlight>
</syntaxhighlight>


=== Verify zero VMs ===
=== NAS status ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
echo "running_guests=$(sudo virsh list --state-running --name 2>/dev/null | sed '/^[[:space:]]*$/d' | wc -l)"
sudo /usr/local/sbin/kitsnet-nas-ha-status
</syntaxhighlight>
</syntaxhighlight>


'''Required:'''
The expected normal state is <code>HEALTHY_FOX</code> unless a legitimate accepted <code>HEALTHY_ANCHOR</code> state exists.


<pre>
=== Exact production VM inventory and managed-save consumption ===
running_guests=0
</pre>


=== Verify final T0 configuration ===
The accepted 12 September 2026 production baseline is '''20''' running libvirt domains.


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo pvs -o pv_name,pv_uuid,vg_name,pv_size,pv_free
cat >/tmp/wort-expected-running <<'EOF'
echo
av001
sudo vgs -o vg_name,vg_uuid,pv_count,lv_count,vg_size,vg_free T0
av002
echo
uv045
sudo lvs -o vg_name,lv_name,lv_uuid,lv_size,lv_attr T0
uv047
</syntaxhighlight>
uv048
uv049
uv050
uv052
uv053
uv054
uv055
uv056
uv057
uv058
uv059
uv060
uv061
uv062
uv063
wv902
EOF


Confirm:
sort -o /tmp/wort-expected-running /tmp/wort-expected-running


* T0 has exactly one PV;
sudo virsh -c qemu:///system list --state-running --name |
* the PVID is unchanged;
  sed '/^[[:space:]]*$/d' |
* the VG UUID is unchanged;
  sort > /tmp/wort-running-now
* all expected LVs exist;
 
* their UUIDs are unchanged; and
cat /tmp/wort-running-now
* the additional replacement-NVMe space now appears as T0 free space.
RUNNING_COUNT=$(wc -l < /tmp/wort-running-now)


=== Verify the save mount survived a clean boot ===
diff -u /tmp/wort-expected-running /tmp/wort-running-now
VM_DIFF_RC=$?


<syntaxhighlight lang="bash">
SAVE_COUNT=$(sudo find /var/lib/libvirt/qemu/save -maxdepth 1 -type f -name '*.save' | wc -l)
for i in {1..10}; do echo; done
sudo findmnt --mountpoint /var/lib/libvirt/qemu/save \
  -o SOURCE,TARGET,FSTYPE,SIZE,USED,AVAIL,OPTIONS


echo
echo "running_guest_count=$RUNNING_COUNT"
sudo find /var/lib/libvirt/qemu/save \
echo "vm_set_diff_rc=$VM_DIFF_RC"
  -maxdepth 1 \
echo "remaining_managed_save_count=$SAVE_COUNT"
  -type f \
  -printf '%f|%s\n' |
sort
</syntaxhighlight>
</syntaxhighlight>


=== Final metadata checks ===
Validated result:
 
<pre>
running_guest_count=20
vm_set_diff_rc=0
remaining_managed_save_count=0
</pre>
 
DONQ inside <code>wv902</code> is intentionally not a separate Wort libvirt domain and remains outside this count. Its simulator startup remains manual by design.
 
=== Completion target and remote consoles ===


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
sudo pvck /dev/nvme0n1p1
systemctl status kitsnet-vm-startup-complete.target --no-pager
PVCK_RC=$?


echo
echo
sudo vgck T0
cat /proc/sys/kernel/random/boot_id
VGCK_RC=$?
 
echo
sudo cat /run/kitsnet-vm-startup/vm-startup-complete


echo
echo
echo "pvck_rc=$PVCK_RC"
systemctl list-units 'socat-kvm@*.service' --no-pager
echo "vgck_rc=$VGCK_RC"
</syntaxhighlight>
</syntaxhighlight>


=== Check host health ===
The completion-marker boot ID must match the current Wort boot ID.


<syntaxhighlight lang="bash">
== Production startup timing and execution lessons ==
for i in {1..10}; do echo; done
echo "===== FAILED SYSTEMD UNITS ====="
systemctl --failed --no-pager


echo
The controlled startup is deliberately readiness-gated. A VM appearing as <code>running</code> does not mean the service it hosts is ready for its dependents.
echo "===== T0 ====="
sudo pvs /dev/nvme0n1p1
sudo vgs T0


echo
The successful production boot progressed as follows:
echo "===== SAVE MOUNT ====="
sudo findmnt --mountpoint /var/lib/libvirt/qemu/save


echo
<pre>
echo "===== STARTUP CONTROL ====="
14:39:35  Wort auto-start begins; waits for production storage
sudo /usr/local/sbin/kitsnet-vm-startup-control status
14:40:01  production storage READY (26-second wait)
</syntaxhighlight>
14:40:02  uv059/Fox starts
14:40:59  Fox preferred-owner condition READY; uv060/Anchor and uv061/mgr1 start
14:41:32  protected NAS service path from mgr1 READY
14:41:47  Docker Swarm manager READY
14:41:47  uv062/wrk1 and uv063/wrk2 start
14:42:53  automatic_controlled_vm_startup=COMPLETE
</pre>


Do not return to production if there is an unexplained storage, LVM, mount or host failure.
The approximately 48-second interval from mgr1 starting to wrk1/wrk2 starting was normal. The cold-start logic waits for <code>nas_service</code>, then <code>swarm_manager</code>, then starts the workers and waits for <code>swarm_workers</code>. The default manager and worker readiness timeouts are 300 seconds. Do not manually start workers merely because mgr1 is already visible in <code>virsh list</code>.


== Stage 15 — return KitsNet to normal automatic startup ==
During the first production-start attempt, Wort was accidentally rebooted while startup was still in progress. The following automatic boot recovered normally and reached the full 20-domain accepted state. Wort had no persistent previous-boot journal available, so <code>journalctl -b -1</code> could not reconstruct the interrupted boot. Persistent journald may be configured separately if future boot-to-boot forensics are desired.


Only after every preceding validation has passed should the next boot be changed from <code>off</code> to <code>auto</code>.
=== Known unrelated systemd warning ===


=== Explicitly select normal automatic startup ===
The target file <code>/etc/systemd/system/kitsnet-vm-startup-complete.target</code> contains:


<syntaxhighlight lang="bash">
<pre>
for i in {1..10}; do echo; done
Documentation=KitsNet controlled VM startup completion point
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot auto
</pre>
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control status
</syntaxhighlight>
 
Verify that the next boot is configured for normal automatic controlled startup.
 
=== Perform the production boot ===
 
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
sudo systemctl reboot
</syntaxhighlight>
 
The normal Wort controller now owns startup sequencing.
 
Do not manually start Fox, Anchor, mgr1, workers, ordinary VMs or socat listeners unless the established startup-control procedure reports a failure requiring diagnosis.
 
== Stage 16 — post-return production validation ==
 
=== Startup-controller status ===
 
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control status
 
echo
systemctl status kitsnet-vm-auto-start.service --no-pager -l
 
echo
sudo journalctl -b -u kitsnet-vm-auto-start.service --no-pager
</syntaxhighlight>
 
=== NAS status ===
 
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-nas-ha-status
</syntaxhighlight>


The expected normal state is <code>HEALTHY_FOX</code> unless a legitimate accepted <code>HEALTHY_ANCHOR</code> state exists.
<code>Documentation=</code> expects documentation URIs/identifiers, not free-form prose. Systemd therefore logs <code>Invalid URL</code> warnings. The warning did not prevent the target from becoming active and did not affect this migration. Remove the line or replace it with a valid documentation URI as separate housekeeping.


=== VM counts ===
== Final production acceptance record ==
 
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
echo "running_guests=$(sudo virsh list --state-running --name | sed '/^[[:space:]]*$/d' | wc -l)"
echo "autostart_links=$(sudo find /etc/libvirt/qemu/autostart -maxdepth 1 -type l 2>/dev/null | wc -l)"
echo "managed_save_files=$(sudo find /var/lib/libvirt/qemu/save -maxdepth 1 -type f 2>/dev/null | wc -l)"
echo "held_autostarts=$(sudo find /var/lib/kitsnet-vm-startup/autostart-hold -maxdepth 1 -type l 2>/dev/null | wc -l)"
</syntaxhighlight>
 
At the accepted Wort production baseline the normal complete state is:


<pre>
<pre>
running_guests=19
PASS: WORT PRODUCTION STARTUP COMPLETE
autostart_links=17
PASS: ALL 20 EXPECTED GUESTS ARE RUNNING
managed_save_files=0
PASS: ALL MANAGED-SAVE STATES HAVE BEEN CONSUMED
held_autostarts=0
PASS: T0 NVME REPLACEMENT AND EXPANSION VALIDATION COMPLETE
</pre>
</pre>


If the intentionally-running VM inventory has changed since this runbook was written, compare with the current documented production inventory rather than forcing those historical counts.
Final NAS state was <code>HEALTHY_FOX</code>, authority owner <code>fox</code>, Anchor BACKUP, one live production VDB, valid lease, and zero failed systemd units.
 
=== Completion target and remote consoles ===
 
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
systemctl status kitsnet-vm-startup-complete.target --no-pager
 
echo
cat /proc/sys/kernel/random/boot_id
 
echo
sudo cat /run/kitsnet-vm-startup/vm-startup-complete
 
echo
systemctl list-units 'socat-kvm@*.service' --no-pager
</syntaxhighlight>
 
The completion-marker boot ID must match the current Wort boot ID.


== Rollback ==
== Rollback ==
Line 1,995: Line 2,217:
* the managed-save inventory survived the migration;
* the managed-save inventory survived the migration;
* a complete no-VM boot succeeded using the final storage configuration;
* a complete no-VM boot succeeded using the final storage configuration;
* the subsequent automatic KitsNet startup completed successfully;
* the subsequent automatic KitsNet startup completed successfully with the exact 20-domain production set;
* NAS ownership is coherent;
* all managed-save files have been consumed after successful restoration;
* the Wort startup-completion target is active; and
* NAS ownership is coherent (<code>HEALTHY_FOX</code>, owner <code>fox</code> in the validated execution);
* the Wort startup-completion target is active and <code>automatic_controlled_vm_startup=COMPLETE</code> was logged;
* <code>lvmdevices --check</code> returns zero;
* <code>systemctl --failed</code> reports zero failed units; and
* expected VMs and console listeners are operational.
* expected VMs and console listeners are operational.


Only after a suitable observation period should the original NVMe or dock recovery image be considered available for reuse.
Only after a suitable observation period should the original NVMe or dock recovery image be considered available for reuse.

Latest revision as of 14:59, 12 September 2026

Status: Completed and production-validated 12 September 2026

Host: wort

Execution revision: Incorporates corrections and lessons learned from the completed Plextor-to-Sabrent T0 migration on 12 September 2026.

Purpose: Replace the physical NVMe device containing the T0 LVM volume group with a larger NVMe when both devices cannot be installed simultaneously. The migration uses a verified compressed whole-device image stored on a dock-attached disk, followed by restoration to the replacement NVMe, validation of the cloned LVM identity, expansion of the partition and PV, restoration of the libvirt managed-save mount, and controlled return of KitsNet virtual machines to service.

Depends on:


1 Validated execution record and authoritative corrections — 12 September 2026[edit | edit source]

This runbook was executed end-to-end on Wort. The corrections in this section are authoritative where they differ from the original planned steps later in the document. The later procedure has also been patched where practical, but this section records the actual accepted behavior and final state.

1.1 Hardware and preserved identity[edit | edit source]

The completed migration replaced:

  • source device: PLEXTOR PX-512M8PeG, serial P02709118482, WWN/EUI eui.0023035630011a50;
  • source whole-device size: 512110190592 bytes;
  • replacement device: Sabrent SB-RKT4L-2TB, serial 48880885900620, WWN/EUI eui.00000000000000006479a7a11ac00487;
  • replacement whole-device size: 2048408248320 bytes; and
  • logical/physical sector size: 512/512 bytes on both devices.

The replacement was installed in Wort's primary NVMe slot. Wort uses an AMD Ryzen 7 5800X processor.

The identities preserved through the clone were:

GPT disk GUID: 4481283C-CEB0-4E65-BA1D-9997D5708D63
Partition start: 2048
Partition PARTUUID: 3703AD9C-AB50-486A-B154-E83338BC7D3F
T0 PVID: Xn55sH-OURj-BQdg-GboO-RTia-1HfD-knB8F2
T0 VG UUID: 2Ttv1r-QKPc-hu5W-HDk4-j38t-06wO-WEjFq7
kvm_save filesystem UUID: 28ea4005-788c-44a6-a014-6a10d396dba7
LV count: 27

Final expanded state:

/dev/nvme0n1p1 start=2048 size=4000795279 sectors
last usable GPT LBA=4000797326
T0 size=1.86 TiB
T0 free=1.50 TiB

1.2 Recovery image[edit | edit source]

/mnt/KitsNet-pbd2/wort-T0-nvme-migration/wort-T0-nvme0n1.full.img.gz
compressed size=139282868502 bytes
SHA256=4329dc57a2ee58c492713108da3c8491e657793c09aa6adfc3c02947bb49b09c

The image passed pigz -t and a byte-for-byte comparison against the original Plextor before removal. After restoration, the original 512110190592-byte source region on the Sabrent compared byte-for-byte identical to the image.

1.3 Corrected operating rules[edit | edit source]

  1. Wort's installed lsblk does not support the START column used in the first draft. Obtain partition start and size from sfdisk --dump.
  2. The recovery disk's /dev/sdX letter changed across boots. Use the mounted filesystem /mnt/KitsNet-pbd2 as the recovery-path authority and identify the physical disk by model/WWN when remounting it. The validated recovery disk is WDC WD8004FRYZ-01VAEB0, WWN 0x5000cca0c3c5c772.
  3. A validation-script STOP is not automatically a storage failure. False STOPs occurred because of bad field parsing, comparing differently formatted LV manifests, attempting to read a root-only file through unprivileged shell redirection, and expecting virsh to work during a deliberately protected boot.
  4. On a protected off boot, libvirt-guests.service and libvirt management sockets are runtime-masked by design. Do not manually unmask them merely to run virsh. Validate protection from kitsnet-vm-startup-control status plus zero QEMU processes.
  5. T0 LVs can autoactivate after partprobe/udev activity. During post-clone geometry work, the meaningful safety tests are zero QEMU processes, no T0 mounts, and device-mapper open counts of zero. Autoactivation by itself is not corruption.
  6. The completed migration did not require a reboot between growpart and pvresize. GPT relocation, partition growth, kernel/udev settle, PARTUUID verification/restoration, and pvresize were completed in one protected boot.
  7. pvs/vgs may show the enlarged geometry before the explicit pvresize; this happened in the successful run. Still execute pvresize and require rc 0.
  8. Before every comparison against dock-resident baseline files, verify that /mnt/KitsNet-pbd2 is actually mounted. An unmounted recovery filesystem can make a valid baseline look missing.
  9. Do not use wc -l < /root/file as an unprivileged user and expect a leading sudo elsewhere to help; shell redirection happens before sudo. Read the dock baseline directly or run the complete read under sudo.

1.4 Workload-specific handling[edit | edit source]

For this maintenance, DONQ (the AlphaServer simulator inside wv902/crownroyal) was manually shut down before wv902 was managed-saved because of VMScluster sensitivity. DONQ remains explicitly outside Wort automation and should be restarted manually after wv902 returns. MORGAN on uv056/maddog required no special handling.

The shutdown produced exactly 18 managed-save files: av001, av002, uv045, uv047, uv048, uv049, uv050, uv052, uv053, uv054, uv055, uv056, uv057, uv058, uv061, uv062, uv063, and wv902. Fox (uv059) and Anchor (uv060) were cleanly shut down rather than managed-saved.

1.5 LVM devices-file lesson[edit | edit source]

After restoration, T0 was initially absent from normal LVM discovery even though the cloned partition contained the correct PVID. The cause was /etc/lvm/devices/system.devices, which still associated the PVID with the old Plextor WWID.

Before changing the devices file, the clone was proven read-only with explicit --devices /dev/nvme0n1p1 overrides. Then lvmdevices --check --refresh reported the new Sabrent sys_wwid and returned rc 5 because an update was required. lvmdevices --update --refresh updated the hardware association, and a final lvmdevices --check returned rc 0. No PVID, VG UUID, or LV UUID changed.

1.6 Partition-growth lesson[edit | edit source]

The successful expansion sequence was:

sgdisk -e /dev/nvme0n1
growpart -N /dev/nvme0n1 1
growpart /dev/nvme0n1 1
sync
partprobe /dev/nvme0n1
udevadm settle
verify start and PARTUUID
pvresize /dev/nvme0n1p1

The start sector remained 2048. During the run the partition unique GUID was unexpectedly observed as BA942590-8B70-463D-AA42-5E03D7A280E6 rather than the original 3703AD9C-AB50-486A-B154-E83338BC7D3F. Because disk GUID, partition contents, PVID, VG UUID, all 27 LV UUIDs, and filesystem identity had already been independently proven, the original PARTUUID was restored with:

for i in {1..10}; do echo; done
sudo sgdisk -u 1:3703AD9C-AB50-486A-B154-E83338BC7D3F /dev/nvme0n1
sudo sync
sudo partprobe /dev/nvme0n1
sudo udevadm settle
sudo sgdisk -v /dev/nvme0n1

Do not invent a replacement PARTUUID. Only restore the captured original after clone identity has already been proven.

1.7 Protected validation lesson[edit | edit source]

The final protected validation boot intentionally left libvirt management unavailable. The correct acceptance tests were:

  • current_boot_mode=off;
  • current_boot_protection=PROTECTED;
  • next_boot_mode=off re-armed for safety;
  • zero QEMU processes;
  • GPT structurally clean;
  • partition start/PARTUUID correct;
  • T0 PVID/VG UUID and 27-LV inventory correct;
  • lvmdevices --check rc 0;
  • T0/kvm_save automatically mounted with UUID 28ea4005-788c-44a6-a014-6a10d396dba7; and
  • all 18 managed-save files still present with exact name/size matches.

Do not start a masked virtqemud.socket merely to make virsh available during this boot.

1.8 Production-start timing lesson[edit | edit source]

The production startup is intentionally readiness-gated. In the successful boot:

14:39:35  automatic startup begins; waits for Wort production storage
14:40:01  production storage READY (26-second wait)
14:40:02  uv059/Fox starts
14:40:59  preferred Fox ownership READY; uv060/Anchor and uv061/mgr1 start
14:41:32  protected NAS service path from mgr1 READY
14:41:47  Docker Swarm manager READY
14:41:47  uv062/wrk1 and uv063/wrk2 start
14:42:53  automatic_controlled_vm_startup=COMPLETE

The approximately 48-second gap from mgr1 starting to the workers starting was normal. The controller was waiting first for nas_service and then for swarm_manager; the worker VMs started immediately when those gates passed. Do not manually start wrk1/wrk2 just because mgr1 already appears as running.

An accidental reboot occurred during the first production-start attempt. The subsequent automatic boot completed normally and all expected guests returned. Wort did not retain a persistent previous-boot journal, so journalctl -b -1 could not reconstruct the interrupted boot. Persistent journald may be enabled separately if boot-to-boot forensic history is desired.

1.9 Final accepted production state[edit | edit source]

current_boot_mode=auto
current_boot_protection=RELEASED
next_boot_mode=auto
running_guest_count=20
vm_set_diff_rc=0
remaining_managed_save_count=0
startup_complete_target=active
startup_complete_marker_count=1
startup_process_count=0
T0=1.86 TiB total / 1.50 TiB free
kvm_save UUID=28ea4005-788c-44a6-a014-6a10d396dba7
lvmdevices_check_rc=0
NAS coherence_state=HEALTHY_FOX
NAS authority_owner=fox
nas_status_rc=0
failed systemd units=0

The exact 20-domain production set is:

av001
av002
uv045
uv047
uv048
uv049
uv050
uv052
uv053
uv054
uv055
uv056
uv057
uv058
uv059
uv060
uv061
uv062
uv063
wv902

1.10 Unrelated systemd warning discovered during this work[edit | edit source]

/etc/systemd/system/kitsnet-vm-startup-complete.target contains a free-form Documentation=KitsNet controlled VM startup completion point line. Documentation= expects documentation URIs/identifiers, so systemd emits Invalid URL warnings. This did not prevent the target from becoming active and did not affect the NVMe migration. Remove that line or replace it with a valid documentation URI as separate housekeeping.


2 Current storage design[edit | edit source]

At the time this procedure was written:

  • T0 is a single-PV volume group.
  • The T0 PV is /dev/nvme0n1p1.
  • The underlying whole device is /dev/nvme0n1.
  • T0 contains 27 logical volumes used by the Wort virtual-machine environment.
  • T0/kvm_save is a 64-GiB LV used for /var/lib/libvirt/qemu/save.
  • Wort's host operating-system filesystems are on the separate T0h VG, not on T0.
  • T0 can therefore be completely deactivated while Wort itself remains operational.

The migration deliberately clones the entire NVMe device, not merely partition 1.

A whole-device clone preserves the on-disk:

  • partition table;
  • disk GUID, if GPT;
  • partition GUID/PARTUUID;
  • LVM PV UUID;
  • T0 VG UUID;
  • LV UUIDs;
  • filesystem UUIDs;
  • LVM metadata;
  • allocated and unallocated extents; and
  • all VM disk and managed-save contents.

The replacement NVMe's physical model, serial number, WWID and other controller-level identifiers are not cloned.

3 Governing safety rules[edit | edit source]

Do not improvise past a STOP checkpoint. If a required validation does not produce the documented result, stop the procedure and diagnose the discrepancy before proceeding.

The following rules apply for the entire maintenance:

  1. The original NVMe is not erased, modified, reformatted or repurposed after removal until the replacement has passed final production validation.
  2. No pvcreate, vgcreate, vgimportclone, pvchange --uuid, vgchange --uuid or routine vgcfgrestore is used.
  3. The image is made from the whole device /dev/nvme0n1.
  4. The image is restored to the whole replacement device.
  5. T0 must be inactive while the source image is created.
  6. The replacement NVMe must be at least as large in bytes as the source NVMe.
  7. The replacement NVMe must present the same logical sector size as the source NVMe.
  8. The clone is validated before any expansion operation is performed.
  9. Partition expansion and PV expansion are separate from the cloning operation.
  10. The /var/lib/libvirt/qemu/save mount remains available until the controlled VM shutdown/managed-save operation is complete.
  11. The save mount is then unmounted and disabled in /etc/fstab until T0 has been restored and expanded.
  12. Every maintenance boot remains VM-protected.
  13. Because next-boot off is one-shot, the first administrative action after every protected maintenance boot is to arm next-boot off again for the following boot.
  14. The final production boot is the only boot for which next-boot auto is explicitly selected.

4 Migration overview[edit | edit source]

The controlled sequence is:

Normal production Wort
        |
        v
Preflight and record identities
        |
        v
Arm next boot OFF
        |
        v
Controlled managed-save / NAS shutdown
        |
        v
Unmount and disable kvm_save mount
        |
        v
Deactivate T0
        |
        v
Whole NVMe -> dd -> pigz -p 12 -> dock image
        |
        v
Verify compressed image against original NVMe
        |
        v
Power off Wort
        |
        v
Physically replace NVMe
        |
        v
Protected no-VM boot
        |
        v
Verify replacement identity / geometry
        |
        v
Dock image -> pigz -d -> dd -> replacement NVMe
        |
        v
Byte-for-byte verification of restored region
        |
        v
Protected reboot
        |
        v
Prove exact LVM clone before expansion
        |
        v
Move GPT backup header if GPT
        |
        v
Grow partition 1
        |
        v
Verify/restore original PARTUUID if required
        |
        v
pvresize existing T0 PV in same protected boot
        |
        v
Restore and test kvm_save fstab mount
        |
        v
Protected validation reboot
        |
        v
Final validation with zero VMs
        |
        v
Select AUTO for next boot
        |
        v
Normal controlled KitsNet startup

5 Stage 0 — pre-maintenance validation[edit | edit source]

This stage is performed while Wort is operating normally and before any VM shutdown.

5.1 Confirm that this is Wort[edit | edit source]

for i in {1..10}; do echo; done
hostname
hostname -f

Expected: the host is Wort.

STOP if these commands are being executed on another host.

5.2 Define the source-device names[edit | edit source]

The following names reflect the current T0 design.

for i in {1..10}; do echo; done
SRC=/dev/nvme0n1
PV=/dev/nvme0n1p1

echo "SRC=$SRC"
echo "PV=$PV"

sudo pvs "$PV"
sudo vgs T0
sudo lvs T0

Expected:

  • /dev/nvme0n1p1 belongs to VG T0.
  • T0 has exactly one PV.

STOP if the source device or T0 topology differs.

5.3 Confirm the currently mounted libvirt save filesystem[edit | edit source]

for i in {1..10}; do echo; done
sudo findmnt --mountpoint /var/lib/libvirt/qemu/save
echo
sudo lvs -o vg_name,lv_name,lv_size,lv_attr T0/kvm_save
echo
grep -n '/var/lib/libvirt/qemu/save' /etc/fstab

The mount must be understood before continuing.

Do not unmount it yet. The production VM shutdown mechanism needs this filesystem for ordinary managed-save files.

5.4 Verify the production VM shutdown override[edit | edit source]

for i in {1..10}; do echo; done
systemctl cat libvirt-guests.service
echo
systemctl cat libvirt-guests.service |
grep -F '/usr/local/sbin/libvirt-guests-parallel.sh stop'

Expected: the effective ExecStop path includes:

/usr/local/sbin/libvirt-guests-parallel.sh stop

STOP if the KitsNet override is not in effect.

5.5 Verify required utilities[edit | edit source]

for i in {1..10}; do echo; done
for cmd in \
  dd pigz cmp blockdev lsblk sfdisk blkid \
  pvs vgs lvs vgchange pvresize pvck vgck vgcfgbackup \
  lvmconfig lvmdevices pvscan \
  findmnt mountpoint growpart
do
    if command -v "$cmd" >/dev/null 2>&1; then
        printf 'OK      %s -> %s\n' "$cmd" "$(command -v "$cmd")"
    else
        printf 'MISSING %s\n' "$cmd"
    fi
done

echo
if command -v nvme >/dev/null 2>&1; then
    echo "OPTIONAL nvme -> $(command -v nvme)"
else
    echo "OPTIONAL nvme command is not installed"
fi

echo
if command -v sgdisk >/dev/null 2>&1; then
    echo "sgdisk -> $(command -v sgdisk)"
else
    echo "sgdisk is not installed; it is required if the source disk is GPT"
fi

STOP if any mandatory utility is missing.

If the source is GPT, sgdisk is also mandatory.

Install/repair required tooling before entering the maintenance outage.

5.6 Record source geometry[edit | edit source]

for i in {1..10}; do echo; done
SRC=/dev/nvme0n1
PV=/dev/nvme0n1p1

echo "===== SOURCE DEVICE ====="
sudo lsblk -o NAME,PATH,SIZE,TYPE,FSTYPE,PTTYPE,MODEL,SERIAL,WWN,LOG-SEC,PHY-SEC "$SRC"

echo
echo "===== SOURCE SIZE IN BYTES ====="
sudo blockdev --getsize64 "$SRC"

echo
echo "===== SOURCE LOGICAL SECTOR SIZE ====="
sudo blockdev --getss "$SRC"

echo
echo "===== PARTITION START AND SIZE ====="
sudo sfdisk --dump "$SRC" | grep -F "$PV"
PART_START=$(sudo sfdisk --dump "$SRC" 2>/dev/null | sed -n "\|^$PV |s/.*start=[[:space:]]*\([0-9][0-9]*\),.*/\1/p")
PART_SIZE=$(sudo sfdisk --dump "$SRC" 2>/dev/null | sed -n "\|^$PV |s/.*size=[[:space:]]*\([0-9][0-9]*\),.*/\1/p")
echo "partition_start=$PART_START"
echo "partition_size_sectors=$PART_SIZE"

echo
echo "===== PARTUUID ====="
sudo blkid -s PARTUUID -o value "$PV"

echo
echo "===== PARTITION TABLE TYPE ====="
sudo blkid -p -s PTTYPE -o value "$SRC"

The partition-table type is expected to be either:

gpt

or:

dos

Do not assume GPT until this command has confirmed it.

5.7 Configure and validate the dock destination[edit | edit source]

Set DOCK to the actual mounted dock filesystem.

The validated recovery mount is /mnt/KitsNet-pbd2. Verify that this mount is present before every use of dock-resident baselines or the image.

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
SRC=/dev/nvme0n1

echo "DOCK=$DOCK"

echo
echo "===== DOCK MOUNT ====="
findmnt --mountpoint "$DOCK"

echo
echo "===== DOCK FILESYSTEM ====="
df -Th "$DOCK"

echo
echo "===== DOCK FREE BYTES ====="
DOCK_FREE=$(df -B1 --output=avail "$DOCK" | tail -1 | tr -d ' ')
SRC_BYTES=$(sudo blockdev --getsize64 "$SRC")

echo "source_bytes=$SRC_BYTES"
echo "dock_free_bytes=$DOCK_FREE"

NEED_BYTES=$((SRC_BYTES + SRC_BYTES / 20))
echo "recommended_minimum_free_bytes=$NEED_BYTES"

if [ "$DOCK_FREE" -ge "$NEED_BYTES" ]; then
    echo "PASS: dock has at least source size plus 5 percent"
else
    echo "STOP: insufficient worst-case dock capacity"
fi

The 5-percent allowance ensures that the procedure does not depend on the source being compressible.

STOP unless the dock mount and capacity are correct.

5.8 Create the migration directory[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"

sudo mkdir -p "$MIG"
sudo touch "$MIG/.write-test"
sudo rm -f "$MIG/.write-test"

echo "MIG=$MIG"
sudo ls -ld "$MIG"

5.9 Capture the pre-migration identity and recovery metadata[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
SRC=/dev/nvme0n1
PV=/dev/nvme0n1p1

PTTYPE=$(sudo blkid -p -s PTTYPE -o value "$SRC")

sudo blockdev --getsize64 "$SRC" |
sudo tee "$MIG/source.bytes"

sudo blockdev --getss "$SRC" |
sudo tee "$MIG/source.logical-sector-size"

printf '%s\n' "$PTTYPE" |
sudo tee "$MIG/source.pttype"

sudo sfdisk --dump "$SRC" |
sudo tee "$MIG/source.sfdisk"

sudo sfdisk --dump "$SRC" 2>/dev/null |
sed -n "\|^$PV |s/.*start=[[:space:]]*\([0-9][0-9]*\),.*/\1/p" |
sudo tee "$MIG/partition-start.before"

sudo sfdisk --dump "$SRC" 2>/dev/null |
sed -n "\|^$PV |s/.*size=[[:space:]]*\([0-9][0-9]*\),.*/\1/p" |
sudo tee "$MIG/partition-size.before"

sudo blkid -s PARTUUID -o value "$PV" |
sudo tee "$MIG/partuuid.before"

sudo pvs --noheadings -o pv_uuid "$PV" |
tr -d ' ' |
sudo tee "$MIG/pvid.before"

sudo vgs --noheadings -o vg_uuid T0 |
tr -d ' ' |
sudo tee "$MIG/vguid.before"

sudo lvs --noheadings --separator '|' -o lv_name,lv_uuid T0 |
sed 's/^[[:space:]]*//;s/[[:space:]]*$//' |
sort |
sudo tee "$MIG/lvuuids.before"

sudo blkid -s UUID -o value /dev/T0/kvm_save |
sudo tee "$MIG/kvm-save-fsuuid.before"

sudo findmnt --mountpoint /var/lib/libvirt/qemu/save \
  -o SOURCE,TARGET,FSTYPE,OPTIONS |
sudo tee "$MIG/kvm-save-mount.before"

sudo cp -a /etc/fstab "$MIG/fstab.before"

sudo vgcfgbackup -f "$MIG/T0.vgcfg" T0

if sudo test -d /etc/lvm/devices; then
    sudo cp -a /etc/lvm/devices "$MIG/lvm-devices.before"
fi

if [ "$PTTYPE" = "gpt" ]; then
    sudo sgdisk --backup="$MIG/source.gpt" "$SRC"

    sudo sgdisk -p "$SRC" |
    awk '/Disk identifier \(GUID\):/{print $4}' |
    sudo tee "$MIG/disk-guid.before"

    sudo sgdisk -v "$SRC"
fi

For GPT, sgdisk -v must not report existing structural corruption.

STOP if the source partition table or LVM metadata already appears damaged.

6 Stage 1 — quiesce all Wort VMs[edit | edit source]

6.1 Workload-specific pre-quiesce note[edit | edit source]

For the 12 September 2026 maintenance, DONQ (the AlphaServer simulator running inside wv902/crownroyal) was manually shut down cleanly before wv902 was managed-saved. This was an explicit VMScluster sensitivity constraint and remains outside Wort's automatic startup/shutdown model. MORGAN on uv056/maddog required no special handling.

Do not add DONQ-specific automation to this runbook. If this workload constraint still applies, perform the same manual DONQ shutdown before the host-level quiesce.

6.2 Arm the next boot for zero VMs[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control status

Verify that the next boot is configured for off mode.

6.3 Run the accepted production shutdown sequence[edit | edit source]

Stopping libvirt-guests.service invokes the KitsNet production wrapper. The wrapper performs managed-save processing for ordinary guests and then invokes the owner-aware NAS shutdown mechanism.

for i in {1..10}; do echo; done
sudo systemctl stop libvirt-guests.service
RC=$?

echo
echo "libvirt-guests stop rc=$RC"
echo
systemctl status libvirt-guests.service --no-pager -l

Required result:

libvirt-guests stop rc=0

STOP if the stop operation fails.

Do not manually destroy a guest to force the procedure onward.

6.4 Prove that no VM remains running[edit | edit source]

for i in {1..10}; do echo; done
echo "===== RUNNING LIBVIRT GUESTS ====="
sudo virsh list --state-running --name

echo
echo "===== RUNNING GUEST COUNT ====="
RUNNING=$(sudo virsh list --state-running --name |
  sed '/^[[:space:]]*$/d' |
  wc -l)

echo "running_guests=$RUNNING"

echo
echo "===== QEMU PROCESSES ====="
ps -eo pid,args |
grep -E '[q]emu-system|[q]emu-kvm' || true

Required result:

running_guests=0

STOP if any guest or QEMU process remains active.

6.5 Record the managed-save files before unmounting T0/kvm_save[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"

sudo find /var/lib/libvirt/qemu/save \
  -maxdepth 1 \
  -type f \
  -name '*.save' \
  -printf '%f|%s\n' |
sort |
sudo tee "$MIG/managed-save.before"

The exact count depends on which ordinary guests were running before shutdown. The file list will later be compared after restoration.

For the validated execution the manifest contained exactly 18 files: av001, av002, uv045, uv047, uv048, uv049, uv050, uv052, uv053, uv054, uv055, uv056, uv057, uv058, uv061, uv062, uv063, and wv902. Fox (uv059) and Anchor (uv060) were cleanly shut down rather than managed-saved.

7 Stage 2 — remove the libvirt save mount from the maintenance boot path[edit | edit source]

7.1 Unmount the save filesystem[edit | edit source]

for i in {1..10}; do echo; done
sudo umount /var/lib/libvirt/qemu/save
RC=$?

echo "umount rc=$RC"

if mountpoint -q /var/lib/libvirt/qemu/save; then
    echo "STOP: /var/lib/libvirt/qemu/save is still mounted"
else
    echo "PASS: save filesystem is unmounted"
fi

STOP unless umount rc=0 and the filesystem is no longer a mountpoint.

7.2 Verify exactly one active fstab entry exists before commenting it[edit | edit source]

for i in {1..10}; do echo; done
ACTIVE_FSTAB=$(
  awk '
    $0 !~ /^[[:space:]]*#/ &&
    $2 == "/var/lib/libvirt/qemu/save" { n++ }
    END { print n+0 }
  ' /etc/fstab
)

echo "active_save_fstab_entries=$ACTIVE_FSTAB"
grep -n '/var/lib/libvirt/qemu/save' /etc/fstab

Required result:

active_save_fstab_entries=1

STOP otherwise.

7.3 Comment the save mount[edit | edit source]

for i in {1..10}; do echo; done
sudo sed -i \
  '\|^[[:space:]]*[^#].*[[:space:]]/var/lib/libvirt/qemu/save[[:space:]]| s|^|# KN-T0-NVME-MIGRATION |' \
  /etc/fstab

echo "===== RESULTING FSTAB ENTRY ====="
grep -n '/var/lib/libvirt/qemu/save' /etc/fstab

echo
echo "===== ACTIVE ENTRIES REMAINING ====="
awk '
  $0 !~ /^[[:space:]]*#/ &&
  $2 == "/var/lib/libvirt/qemu/save" { print }
' /etc/fstab

echo
sudo systemctl daemon-reload
sudo findmnt --verify --verbose

There must be no active fstab entry for the save mount.

The original line must remain present with the prefix:

# KN-T0-NVME-MIGRATION

8 Stage 3 — deactivate T0 completely[edit | edit source]

for i in {1..10}; do echo; done
sudo vgchange -an T0
RC=$?

echo
echo "vgchange rc=$RC"

echo
echo "===== T0 LV STATE ====="
sudo lvs -o vg_name,lv_name,lv_active,lv_attr,lv_size T0

echo
ACTIVE=$(
  sudo lvs --noheadings -o lv_active T0 |
  awk '$1 == "active" { n++ } END { print n+0 }'
)

echo "active_T0_LVs=$ACTIVE"

Required results:

vgchange rc=0
active_T0_LVs=0

STOP if T0 cannot be completely deactivated.

Do not image a partially active T0.

9 Stage 4 — create the compressed whole-NVMe image[edit | edit source]

9.1 Revalidate source and destination immediately before imaging[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
SRC=/dev/nvme0n1
PV=/dev/nvme0n1p1
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"

echo "===== SOURCE ====="
sudo lsblk -o NAME,PATH,SIZE,TYPE,FSTYPE,MODEL,SERIAL,LOG-SEC,PHY-SEC "$SRC"

echo
echo "===== SOURCE PV ====="
sudo pvs "$PV"

echo
echo "===== IMAGE DESTINATION ====="
echo "$IMG"

if sudo test -e "$IMG"; then
    echo "STOP: image file already exists; it will NOT be overwritten automatically"
else
    echo "PASS: image pathname is unused"
fi

echo
df -h "$DOCK"

STOP if the source device is not the known T0 NVMe or if the image filename already exists unexpectedly.

9.2 Create the image[edit | edit source]

This command:

  • reads the whole source NVMe using dd;
  • feeds the raw byte stream into pigz;
  • allows pigz to use up to 12 compression threads;
  • uses normal gzip compression level 6; and
  • enables pipeline failure propagation.

No conv=noerror option is used. A source read error is a migration failure and must not be silently padded or ignored.

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
SRC=/dev/nvme0n1
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"

sudo env SRC="$SRC" IMG="$IMG" \
bash -o pipefail -c '
    if [ -e "$IMG" ]; then
        echo "STOP: refusing to overwrite existing image: $IMG"
        exit 2
    fi

    dd if="$SRC" bs=16M iflag=fullblock status=progress |
    pigz -p 12 -6 > "$IMG"
'

RC=$?

echo
echo "image_pipeline_rc=$RC"

sudo sync

echo
sudo ls -lh "$IMG"

Required result:

image_pipeline_rc=0

STOP otherwise.

10 Stage 5 — validate the image before removing the original NVMe[edit | edit source]

10.1 Test the gzip stream[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"

sudo pigz -t "$IMG"
RC=$?

echo "pigz_test_rc=$RC"

Required result:

pigz_test_rc=0

10.2 Compare the decompressed image byte-for-byte with the original NVMe[edit | edit source]

This is the definitive pre-swap verification.

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
SRC=/dev/nvme0n1
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"

sudo env SRC="$SRC" IMG="$IMG" \
bash -o pipefail -c '
    pigz -dc "$IMG" |
    cmp - "$SRC"
'

RC=$?

echo
echo "source_image_cmp_rc=$RC"

cmp normally prints nothing when the inputs are identical.

Required result:

source_image_cmp_rc=0

STOP if the comparison is anything other than zero.

10.3 Record a checksum of the compressed recovery artifact[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"

sudo sha256sum "$IMG" |
sudo tee "$IMG.sha256"

sudo ls -lh "$IMG" "$IMG.sha256"

At this point there are two independent recovery assets:

  1. the untouched original physical NVMe; and
  2. the byte-verified compressed whole-device image.

11 Stage 6 — power off Wort and replace the NVMe[edit | edit source]

Re-arm protected startup immediately before shutdown, even though it should already be armed.

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
sudo /usr/local/sbin/kitsnet-vm-startup-control status

Verify off for the next boot.

Then power off:

for i in {1..10}; do echo; done
sudo systemctl poweroff

After Wort is completely powered off:

  1. Remove the original T0 NVMe.
  2. Label and retain it intact.
  3. Install the replacement NVMe.
  4. Do not connect the original NVMe simultaneously with its clone.
  5. Keep the dock recovery disk connected.

12 Stage 7 — first protected boot with the replacement NVMe[edit | edit source]

Wort should boot from T0h even though T0 is currently absent from the blank replacement NVMe.

12.1 Immediately protect the following boot[edit | edit source]

The off request that produced this boot has now been consumed.

The first maintenance action is therefore:

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control status
echo
echo "running_guests=$(sudo virsh list --state-running --name 2>/dev/null | sed '/^[[:space:]]*$/d' | wc -l)"

Required result:

running_guests=0

The following boot must also show off as armed.

12.2 Identify the replacement NVMe[edit | edit source]

Do not assume that the Linux device name is correct merely because it is nvme0n1.

for i in {1..10}; do echo; done
sudo lsblk -d -o NAME,PATH,SIZE,MODEL,SERIAL,WWN,LOG-SEC,PHY-SEC

echo
if command -v nvme >/dev/null 2>&1; then
    sudo nvme list
fi

Physically correlate model, serial number and capacity with the installed replacement.

Only after positive identification set:

for i in {1..10}; do echo; done
NEW=/dev/nvme0n1

sudo lsblk -o NAME,PATH,SIZE,TYPE,FSTYPE,MOUNTPOINTS,MODEL,SERIAL,WWN,LOG-SEC,PHY-SEC "$NEW"

STOP if there is any uncertainty regarding the replacement device.

12.3 Verify replacement capacity and logical sector size[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
NEW=/dev/nvme0n1

OLD_BYTES=$(sudo cat "$MIG/source.bytes")
OLD_LOGSEC=$(sudo cat "$MIG/source.logical-sector-size")

NEW_BYTES=$(sudo blockdev --getsize64 "$NEW")
NEW_LOGSEC=$(sudo blockdev --getss "$NEW")

echo "old_bytes=$OLD_BYTES"
echo "new_bytes=$NEW_BYTES"
echo
echo "old_logical_sector_size=$OLD_LOGSEC"
echo "new_logical_sector_size=$NEW_LOGSEC"

echo
if [ "$NEW_BYTES" -ge "$OLD_BYTES" ]; then
    echo "PASS: replacement is large enough"
else
    echo "STOP: replacement is smaller than source"
fi

if [ "$NEW_LOGSEC" -eq "$OLD_LOGSEC" ]; then
    echo "PASS: logical sector sizes match"
else
    echo "STOP: logical sector size mismatch"
fi

Both tests must pass.

A logical-sector-size mismatch is a STOP condition for this whole-device cloning procedure.

12.4 Verify that nothing on the replacement NVMe is mounted[edit | edit source]

for i in {1..10}; do echo; done
NEW=/dev/nvme0n1

echo "===== REPLACEMENT DEVICE TREE ====="
sudo lsblk -o NAME,PATH,SIZE,TYPE,FSTYPE,MOUNTPOINTS "$NEW"

echo
echo "===== NONEMPTY MOUNTPOINTS ====="
sudo lsblk -nr -o MOUNTPOINTS "$NEW" |
sed '/^[[:space:]]*$/d'

The second section must be empty.

STOP if any replacement-NVMe filesystem is mounted.

12.5 Verify the recovery image again before destructive writing[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"

sudo sha256sum -c "$IMG.sha256"
RC=$?

echo "image_sha256_rc=$RC"

Required result:

image_sha256_rc=0

13 Stage 8 — restore the whole-device image to the replacement NVMe[edit | edit source]

13.1 Remove stale partition-table structures from the replacement only[edit | edit source]

This prevents a previously used replacement device from retaining an unrelated backup GPT at its physical end.

This command is destructive.

Run it only after positive identification of NEW.

for i in {1..10}; do echo; done
NEW=/dev/nvme0n1

echo "ABOUT TO DESTROY EXISTING PARTITION TABLES ON:"
sudo lsblk -d -o NAME,PATH,SIZE,MODEL,SERIAL,WWN "$NEW"

echo
read -r -p 'Type WIPE_REPLACEMENT_NVME to continue: ' ANSWER

if [ "$ANSWER" = "WIPE_REPLACEMENT_NVME" ]; then
    sudo sgdisk --zap-all "$NEW"
    echo "sgdisk zap rc=$?"
else
    echo "ABORTED: replacement NVMe was not modified"
fi

Do not continue unless the zap operation completed successfully.

13.2 Restore the image[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"
NEW=/dev/nvme0n1

echo "IMAGE=$IMG"
echo "DESTINATION=$NEW"

echo
sudo lsblk -d -o NAME,PATH,SIZE,MODEL,SERIAL,WWN "$NEW"

echo
read -r -p 'Type RESTORE_T0_IMAGE to overwrite the replacement NVMe: ' ANSWER

if [ "$ANSWER" = "RESTORE_T0_IMAGE" ]; then
    sudo env IMG="$IMG" NEW="$NEW" \
    bash -o pipefail -c '
        pigz -dc "$IMG" |
        dd of="$NEW" bs=16M iflag=fullblock status=progress conv=fsync
    '

    RC=$?
    echo
    echo "restore_pipeline_rc=$RC"
    sudo sync
else
    echo "ABORTED: image was not restored"
fi

Required result:

restore_pipeline_rc=0

STOP otherwise.

13.3 Byte-verify the restored portion of the larger NVMe[edit | edit source]

Because the replacement is larger, compare exactly the number of bytes that existed on the original NVMe.

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"
NEW=/dev/nvme0n1
OLD_BYTES=$(sudo cat "$MIG/source.bytes")

sudo env IMG="$IMG" NEW="$NEW" OLD_BYTES="$OLD_BYTES" \
bash -o pipefail -c '
    pigz -dc "$IMG" |
    cmp -n "$OLD_BYTES" - "$NEW"
'

RC=$?

echo
echo "restored_image_cmp_rc=$RC"

Required result:

restored_image_cmp_rc=0

This proves that every byte belonging to the original disk image was written identically to the replacement.

STOP otherwise.

13.4 Do not expand the disk yet[edit | edit source]

At this point the replacement contains an exact clone of the old NVMe.

If the source uses GPT, a verification command may report that the backup GPT is not at the physical end of the larger device. That condition is expected at this stage.

No correction or expansion is performed until after a clean reboot proves that LVM sees the clone correctly.

13.5 Reboot protected[edit | edit source]

The following boot should already be armed off from the beginning of this maintenance boot. Verify before rebooting:

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control status
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
sudo /usr/local/sbin/kitsnet-vm-startup-control status

Then:

for i in {1..10}; do echo; done
sudo systemctl reboot

14 Stage 9 — prove the exact clone before expansion[edit | edit source]

14.1 Immediately protect the following boot[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control status
echo
echo "running_guests=$(sudo virsh list --state-running --name 2>/dev/null | sed '/^[[:space:]]*$/d' | wc -l)"

Required:

running_guests=0

14.2 Let normal boot discovery run first[edit | edit source]

Do not issue LVM repair commands preemptively.

Begin with read-only inspection:

for i in {1..10}; do echo; done
sudo pvs
echo
sudo vgs
echo
sudo lvs T0 2>&1 || true

The expected outcome is that /dev/nvme0n1p1 appears as the existing T0 PV.

If T0 appears normally, do not run any additional scan/refresh command.

14.3 If T0 is missing: prove the clone first, then update the LVM devices file[edit | edit source]

The completed migration encountered this exact condition. The cloned partition contained the correct PVID, but /etc/lvm/devices/system.devices still associated that PVID with the old Plextor WWID. Normal discovery therefore omitted T0 even though the clone itself was correct.

First inspect whether the devices file is in use:

for i in {1..10}; do echo; done
sudo lvmconfig --type current devices/use_devicesfile
echo
sudo lvmdevices 2>&1 || true

Before updating the devices file, use an explicit device override to prove the cloned identities:

for i in {1..10}; do echo; done
PV=/dev/nvme0n1p1
sudo pvs --devices "$PV" -o pv_name,pv_uuid,vg_name,pv_size,pv_free "$PV"
echo
sudo vgs --devices "$PV" -o vg_name,vg_uuid,pv_count,lv_count,vg_size,vg_free T0
echo
sudo lvs --devices "$PV" -o vg_name,lv_name,lv_uuid,lv_size T0

The PVID, VG UUID and 27-LV inventory must match the captured baseline. STOP if they do not.

If use_devicesfile=1 and the clone identity is correct:

for i in {1..10}; do echo; done
sudo lvmdevices --check --refresh
CHECK_RC=$?
echo "lvmdevices_check_refresh_rc=$CHECK_RC"

echo
sudo lvmdevices --update --refresh
UPDATE_RC=$?
echo "lvmdevices_update_refresh_rc=$UPDATE_RC"

echo
sudo lvmdevices --check
FINAL_RC=$?
echo "lvmdevices_final_check_rc=$FINAL_RC"

echo
sudo pvs
sudo vgs T0

In the validated migration, lvmdevices --check --refresh returned rc 5 and reported that the PVID had a new sys_wwid. This was a diagnostic result indicating that system.devices needed an update. lvmdevices --update --refresh then replaced the Plextor WWID with eui.00000000000000006479a7a11ac00487, and the final lvmdevices --check returned zero.

Do not initialize or repair the PV merely because normal discovery initially fails.

14.4 Check LVM metadata without repairing it[edit | edit source]

for i in {1..10}; do echo; done
PV=/dev/nvme0n1p1

sudo pvck "$PV"
PVCK_RC=$?

echo
sudo vgck T0
VGCK_RC=$?

echo
echo "pvck_rc=$PVCK_RC"
echo "vgck_rc=$VGCK_RC"

Do not use pvck --repair or vgck --updatemetadata during normal migration validation.

STOP on unexplained metadata errors.

14.5 Compare the cloned PV UUID[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
PV=/dev/nvme0n1p1

sudo pvs --noheadings -o pv_uuid "$PV" |
tr -d ' ' > /tmp/T0-pvid.after

echo "===== BEFORE ====="
sudo cat "$MIG/pvid.before"

echo "===== AFTER ====="
cat /tmp/T0-pvid.after

echo "===== DIFF ====="
sudo diff -u "$MIG/pvid.before" /tmp/T0-pvid.after

Required: no difference.

14.6 Compare the cloned VG UUID[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"

sudo vgs --noheadings -o vg_uuid T0 |
tr -d ' ' > /tmp/T0-vguid.after

sudo diff -u "$MIG/vguid.before" /tmp/T0-vguid.after

Required: no difference.

14.7 Compare all LV UUIDs[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"

sudo lvs --noheadings --separator '|' -o lv_name,lv_uuid T0 |
sed 's/^[[:space:]]*//;s/[[:space:]]*$//' |
sort > /tmp/T0-lvuuids.after

sudo diff -u "$MIG/lvuuids.before" /tmp/T0-lvuuids.after

Required: no difference.

14.8 Compare the partition UUID and start sector[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
PV=/dev/nvme0n1p1

sudo blkid -s PARTUUID -o value "$PV" > /tmp/partuuid.after

sudo sfdisk --dump /dev/nvme0n1 2>/dev/null |
sed -n "\|^$PV |s/.*start=[[:space:]]*\([0-9][0-9]*\),.*/\1/p" > /tmp/partition-start.after

echo "===== PARTUUID DIFF ====="
sudo diff -u "$MIG/partuuid.before" /tmp/partuuid.after

echo
echo "===== START-SECTOR DIFF ====="
sudo diff -u "$MIG/partition-start.before" /tmp/partition-start.after

Both comparisons must show no difference.

14.9 For GPT: compare the cloned disk GUID[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
NEW=/dev/nvme0n1
PTTYPE=$(sudo cat "$MIG/source.pttype")

if [ "$PTTYPE" = "gpt" ]; then
    sudo sgdisk -p "$NEW" |
    awk '/Disk identifier \(GUID\):/{print $4}' > /tmp/disk-guid.after

    sudo diff -u "$MIG/disk-guid.before" /tmp/disk-guid.after
else
    echo "Source partition table is $PTTYPE; GPT disk-GUID test does not apply"
fi

For GPT, the GUID must match.

14.10 Activate T0 temporarily and verify kvm_save filesystem identity[edit | edit source]

Do not mount it yet.

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"

sudo vgchange -ay T0
RC=$?

echo "vgchange_activate_rc=$RC"

sudo blkid -s UUID -o value /dev/T0/kvm_save > /tmp/kvm-save-fsuuid.after

echo
sudo diff -u "$MIG/kvm-save-fsuuid.before" /tmp/kvm-save-fsuuid.after

echo
mountpoint -q /var/lib/libvirt/qemu/save &&
    echo "STOP: save filesystem unexpectedly mounted" ||
    echo "PASS: save filesystem remains unmounted"

The filesystem UUID must be identical and the filesystem must remain unmounted.

At this point the clone has been proven.

15 Stage 10 — expand partition 1 and the existing T0 PV[edit | edit source]

The executed maintenance proved that partition growth and pvresize can be completed in the same protected boot. The reboot between the original draft's Stage 10 and Stage 11 is removed.

15.1 Establish the actual safety conditions[edit | edit source]

for i in {1..10}; do echo; done
echo "===== QEMU ====="
QEMU=$(pgrep -fc 'qemu-system|qemu-kvm' 2>/dev/null)
echo "qemu_processes=$QEMU"

echo
echo "===== T0 MOUNTS ====="
findmnt -rn -S '/dev/mapper/T0-*' || true

echo
echo "===== T0 DEVICE-MAPPER OPEN COUNTS ====="
sudo dmsetup info -c --noheadings -o name,open | awk '$1 ~ /^T0-/ {print}'

Required: zero QEMU processes, no T0 filesystem mounts, and open count zero for every T0 device-mapper node.

T0 may autoactivate after partition/udev activity. If the safety conditions above remain true, autoactivation alone is not a STOP condition. Deactivate T0 before the partition write when practical:

for i in {1..10}; do echo; done
sudo vgchange -an T0
RC=$?
echo "vgchange_deactivate_rc=$RC"

15.2 Relocate the GPT backup header[edit | edit source]

for i in {1..10}; do echo; done
NEW=/dev/nvme0n1
sudo sgdisk -e "$NEW"
RC=$?
echo "sgdisk_move_header_rc=$RC"
echo
sudo sgdisk -v "$NEW"

Required: rc 0 and no GPT structural errors.

15.3 Dry-run and perform partition growth[edit | edit source]

for i in {1..10}; do echo; done
NEW=/dev/nvme0n1
PV=/dev/nvme0n1p1

sudo growpart -N "$NEW" 1
DRY_RC=$?
echo "growpart_dry_run_rc=$DRY_RC"

echo
sudo growpart "$NEW" 1
GROW_RC=$?
echo "growpart_rc=$GROW_RC"

sudo sync
sudo partprobe "$NEW"
sudo udevadm settle

echo
sudo sfdisk --dump "$NEW"

The dry run must propose only extension of partition 1 while preserving start sector 2048. On the validated Sabrent the final partition size was 4000795279 sectors and the last usable GPT LBA was 4000797326.

15.4 Verify partition start and PARTUUID[edit | edit source]

for i in {1..10}; do echo; done
MIG=/mnt/KitsNet-pbd2/wort-T0-nvme-migration
NEW=/dev/nvme0n1
PV=/dev/nvme0n1p1

EXPECTED_START=$(cat "$MIG/partition-start.before")
EXPECTED_PARTUUID=$(tr '[:lower:]' '[:upper:]' < "$MIG/partuuid.before")
CURRENT_START=$(sudo sfdisk --dump "$NEW" 2>/dev/null | sed -n "\|^$PV |s/.*start=[[:space:]]*\([0-9][0-9]*\),.*/\1/p")
CURRENT_PARTUUID=$(sudo blkid -s PARTUUID -o value "$PV" | tr '[:lower:]' '[:upper:]')

echo "expected_start=$EXPECTED_START"
echo "current_start=$CURRENT_START"
echo "expected_partuuid=$EXPECTED_PARTUUID"
echo "current_partuuid=$CURRENT_PARTUUID"

The start sector must remain identical.

During the completed migration, the partition unique GUID was unexpectedly observed as BA942590-8B70-463D-AA42-5E03D7A280E6 instead of the captured original 3703AD9C-AB50-486A-B154-E83338BC7D3F. Because the disk GUID, partition contents, PVID, VG UUID, all 27 LV UUIDs, and kvm_save filesystem UUID were already proven, the original partition GUID was restored:

for i in {1..10}; do echo; done
sudo sgdisk -u 1:3703AD9C-AB50-486A-B154-E83338BC7D3F /dev/nvme0n1
RC=$?
echo "partition_guid_restore_rc=$RC"
sudo sync
sudo partprobe /dev/nvme0n1
sudo udevadm settle
sudo sgdisk -v /dev/nvme0n1
sudo blkid -s PARTUUID -o value /dev/nvme0n1p1

Only restore a PARTUUID when the exact original value was captured before migration and all other clone identities have already been proven. Never generate a new identity as part of this procedure.

15.5 Resize the existing PV without an intervening reboot[edit | edit source]

for i in {1..10}; do echo; done
PV=/dev/nvme0n1p1
sudo pvresize "$PV"
RC=$?
echo "pvresize_rc=$RC"
echo
sudo pvs -o pv_name,pv_uuid,pv_size,pv_free "$PV"
sudo vgs -o vg_name,vg_uuid,pv_count,lv_count,vg_size,vg_free T0

Required: pvresize_rc=0. The successful run ended with T0 at approximately 1.86t total and 1.50t free. The LVM reports had already begun showing the larger geometry immediately before the explicit pvresize; that is not a failure.

15.6 Revalidate GPT and all LVM identities[edit | edit source]

for i in {1..10}; do echo; done
MIG=/mnt/KitsNet-pbd2/wort-T0-nvme-migration
PV=/dev/nvme0n1p1

sudo sgdisk -v /dev/nvme0n1
sudo pvck "$PV"
sudo vgck T0

sudo pvs --noheadings -o pv_uuid "$PV" | tr -d ' ' > /tmp/T0-pvid.final
sudo vgs --noheadings -o vg_uuid T0 | tr -d ' ' > /tmp/T0-vguid.final
sudo lvs --noheadings --separator '|' -o lv_name,lv_uuid T0 |
  sed 's/^[[:space:]]*//;s/[[:space:]]*$//' |
  sort > /tmp/T0-lvuuids.final

sudo diff -u "$MIG/pvid.before" /tmp/T0-pvid.final
sudo diff -u "$MIG/vguid.before" /tmp/T0-vguid.final
sudo diff -u "$MIG/lvuuids.before" /tmp/T0-lvuuids.final

All comparisons must be identical.

16 Stage 11 — restore the libvirt save filesystem[edit | edit source]

16.1 Activate T0[edit | edit source]

for i in {1..10}; do echo; done
sudo vgchange -ay T0
RC=$?

echo "vgchange_activate_rc=$RC"

echo
sudo lvs -o vg_name,lv_name,lv_active,lv_attr,lv_size T0

Required:

vgchange_activate_rc=0

16.2 Restore only the fstab line that this procedure commented[edit | edit source]

First verify that the marker exists exactly as expected:

for i in {1..10}; do echo; done
grep -n '^# KN-T0-NVME-MIGRATION ' /etc/fstab

There should be exactly one line, and it must be the original /var/lib/libvirt/qemu/save entry.

Restore it:

for i in {1..10}; do echo; done
sudo sed -i \
  's|^# KN-T0-NVME-MIGRATION ||' \
  /etc/fstab

echo "===== RESTORED ENTRY ====="
grep -n '/var/lib/libvirt/qemu/save' /etc/fstab

echo
sudo systemctl daemon-reload
sudo findmnt --verify --verbose

16.3 Mount the save filesystem manually[edit | edit source]

for i in {1..10}; do echo; done
sudo mount /var/lib/libvirt/qemu/save
RC=$?

echo "mount rc=$RC"

echo
sudo findmnt --mountpoint /var/lib/libvirt/qemu/save \
  -o SOURCE,TARGET,FSTYPE,SIZE,USED,AVAIL,OPTIONS

Required:

mount rc=0

The source must resolve to the restored T0/kvm_save LV.

16.4 Verify the filesystem UUID[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"

sudo blkid -s UUID -o value /dev/T0/kvm_save > /tmp/kvm-save-fsuuid.final

sudo diff -u \
  "$MIG/kvm-save-fsuuid.before" \
  /tmp/kvm-save-fsuuid.final

Required: no difference.

16.5 Verify the managed-save inventory survived[edit | edit source]

Verify that the recovery filesystem is mounted before reading the baseline.

for i in {1..10}; do echo; done
DOCK=/mnt/KitsNet-pbd2
MIG="$DOCK/wort-T0-nvme-migration"
SAVE=/var/lib/libvirt/qemu/save

findmnt --mountpoint "$DOCK"
RECOVERY_RC=$?
echo "recovery_mount_rc=$RECOVERY_RC"

sudo find "$SAVE" \
  -maxdepth 1 \
  -type f \
  -name '*.save' \
  -printf '%f|%s\n' |
sort > /tmp/managed-save.current

EXPECTED_COUNT=$(awk 'END {print NR}' "$MIG/managed-save.before")
CURRENT_COUNT=$(awk 'END {print NR}' /tmp/managed-save.current)

echo "expected_save_count=$EXPECTED_COUNT"
echo "current_save_count=$CURRENT_COUNT"

diff -u "$MIG/managed-save.before" /tmp/managed-save.current
DIFF_RC=$?
echo "managed_save_manifest_diff_rc=$DIFF_RC"

For the validated execution the expected and current counts were both 18 and the diff returned zero.

Do not copy the baseline to /root and then use unprivileged shell redirection such as wc -l < /root/file; that produced a false STOP during this maintenance because the shell opens the file before sudo can affect a child command.

17 Stage 12 — final no-VM validation reboot[edit | edit source]

This reboot tests the final intended storage configuration:

  • replacement NVMe installed;
  • clone proven;
  • partition expanded;
  • T0 PV expanded;
  • T0 identities unchanged;
  • fstab restored; and
  • /var/lib/libvirt/qemu/save restored.

No VM is yet permitted to start.

17.1 Arm the protected boot[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
sudo /usr/local/sbin/kitsnet-vm-startup-control status

Then:

for i in {1..10}; do echo; done
sudo systemctl reboot

18 Stage 13 — final validation with all VMs still off[edit | edit source]

18.1 Immediately preserve protection against an unexpected additional reboot[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control status

18.2 Verify the protected no-VM state[edit | edit source]

A protected off boot intentionally runtime-masks libvirt-guests.service and the virtqemud/libvirtd sockets. virsh can therefore fail with a missing or masked socket even though the boot is correctly protected.

Do not unmask or start those services for validation. Use the controller state and QEMU process count:

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control status

echo
QEMU=$(pgrep -fc 'qemu-system|qemu-kvm' 2>/dev/null)
echo "qemu_processes=$QEMU"

echo
systemctl status libvirt-guests.service --no-pager 2>&1 || true
systemctl list-unit-files --no-pager |
  grep -E '^(virtqemud|libvirtd|libvirt-guests)\.(service|socket)' || true

Required: current_boot_mode=off, current_boot_protection=PROTECTED, and qemu_processes=0. Runtime-masked libvirt units are expected here.

18.3 Verify final T0 configuration[edit | edit source]

for i in {1..10}; do echo; done
sudo pvs -o pv_name,pv_uuid,vg_name,pv_size,pv_free
echo
sudo vgs -o vg_name,vg_uuid,pv_count,lv_count,vg_size,vg_free T0
echo
sudo lvs -o vg_name,lv_name,lv_uuid,lv_size,lv_attr T0

Confirm:

  • T0 has exactly one PV;
  • the PVID is unchanged;
  • the VG UUID is unchanged;
  • all expected LVs exist;
  • their UUIDs are unchanged; and
  • the additional replacement-NVMe space now appears as T0 free space.

18.4 Verify the save mount survived a clean boot[edit | edit source]

for i in {1..10}; do echo; done
sudo findmnt --mountpoint /var/lib/libvirt/qemu/save \
  -o SOURCE,TARGET,FSTYPE,SIZE,USED,AVAIL,OPTIONS

echo
sudo find /var/lib/libvirt/qemu/save \
  -maxdepth 1 \
  -type f \
  -name '*.save' \
  -printf '%f|%s\n' |
sort

18.5 Reconfirm the managed-save manifest after the protected reboot[edit | edit source]

for i in {1..10}; do echo; done
MIG=/mnt/KitsNet-pbd2/wort-T0-nvme-migration
SAVE=/var/lib/libvirt/qemu/save

sudo find "$SAVE" -maxdepth 1 -type f -name '*.save' -printf '%f|%s\n' |
  sort > /tmp/wort-T0-managed-save.after-reboot

EXPECTED=$(awk 'END {print NR}' "$MIG/managed-save.before")
CURRENT=$(awk 'END {print NR}' /tmp/wort-T0-managed-save.after-reboot)

diff -u "$MIG/managed-save.before" /tmp/wort-T0-managed-save.after-reboot
DIFF_RC=$?

echo "expected_save_count=$EXPECTED"
echo "current_save_count=$CURRENT"
echo "managed_save_diff_rc=$DIFF_RC"

For the validated maintenance the required result was 18, 18, and diff rc 0.

18.6 Final metadata checks[edit | edit source]

for i in {1..10}; do echo; done
sudo pvck /dev/nvme0n1p1
PVCK_RC=$?

echo
sudo vgck T0
VGCK_RC=$?

echo
echo "pvck_rc=$PVCK_RC"
echo "vgck_rc=$VGCK_RC"

18.7 Check host health[edit | edit source]

for i in {1..10}; do echo; done
echo "===== FAILED SYSTEMD UNITS ====="
systemctl --failed --no-pager

echo
echo "===== T0 ====="
sudo pvs /dev/nvme0n1p1
sudo vgs T0

echo
echo "===== SAVE MOUNT ====="
sudo findmnt --mountpoint /var/lib/libvirt/qemu/save

echo
echo "===== STARTUP CONTROL ====="
sudo /usr/local/sbin/kitsnet-vm-startup-control status

Do not return to production if there is an unexplained storage, LVM, mount or host failure.

19 Stage 14 — return KitsNet to normal automatic startup[edit | edit source]

Only after every preceding validation has passed should the next boot be changed from off to auto.

19.1 Explicitly select normal automatic startup[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot auto
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control status

Verify that the next boot is configured for normal automatic controlled startup.

19.2 Perform the production boot[edit | edit source]

for i in {1..10}; do echo; done
sudo systemctl reboot

The normal Wort controller now owns startup sequencing.

Do not manually start Fox, Anchor, mgr1, workers, ordinary VMs or socat listeners unless the established startup-control procedure reports a failure requiring diagnosis.

20 Stage 15 — post-return production validation[edit | edit source]

20.1 Startup-controller status[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control status

echo
systemctl status kitsnet-vm-auto-start.service --no-pager -l

echo
sudo journalctl -b -u kitsnet-vm-auto-start.service --no-pager

20.2 NAS status[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-nas-ha-status

The expected normal state is HEALTHY_FOX unless a legitimate accepted HEALTHY_ANCHOR state exists.

20.3 Exact production VM inventory and managed-save consumption[edit | edit source]

The accepted 12 September 2026 production baseline is 20 running libvirt domains.

for i in {1..10}; do echo; done
cat >/tmp/wort-expected-running <<'EOF'
av001
av002
uv045
uv047
uv048
uv049
uv050
uv052
uv053
uv054
uv055
uv056
uv057
uv058
uv059
uv060
uv061
uv062
uv063
wv902
EOF

sort -o /tmp/wort-expected-running /tmp/wort-expected-running

sudo virsh -c qemu:///system list --state-running --name |
  sed '/^[[:space:]]*$/d' |
  sort > /tmp/wort-running-now

cat /tmp/wort-running-now
RUNNING_COUNT=$(wc -l < /tmp/wort-running-now)

diff -u /tmp/wort-expected-running /tmp/wort-running-now
VM_DIFF_RC=$?

SAVE_COUNT=$(sudo find /var/lib/libvirt/qemu/save -maxdepth 1 -type f -name '*.save' | wc -l)

echo "running_guest_count=$RUNNING_COUNT"
echo "vm_set_diff_rc=$VM_DIFF_RC"
echo "remaining_managed_save_count=$SAVE_COUNT"

Validated result:

running_guest_count=20
vm_set_diff_rc=0
remaining_managed_save_count=0

DONQ inside wv902 is intentionally not a separate Wort libvirt domain and remains outside this count. Its simulator startup remains manual by design.

20.4 Completion target and remote consoles[edit | edit source]

for i in {1..10}; do echo; done
systemctl status kitsnet-vm-startup-complete.target --no-pager

echo
cat /proc/sys/kernel/random/boot_id

echo
sudo cat /run/kitsnet-vm-startup/vm-startup-complete

echo
systemctl list-units 'socat-kvm@*.service' --no-pager

The completion-marker boot ID must match the current Wort boot ID.

21 Production startup timing and execution lessons[edit | edit source]

The controlled startup is deliberately readiness-gated. A VM appearing as running does not mean the service it hosts is ready for its dependents.

The successful production boot progressed as follows:

14:39:35  Wort auto-start begins; waits for production storage
14:40:01  production storage READY (26-second wait)
14:40:02  uv059/Fox starts
14:40:59  Fox preferred-owner condition READY; uv060/Anchor and uv061/mgr1 start
14:41:32  protected NAS service path from mgr1 READY
14:41:47  Docker Swarm manager READY
14:41:47  uv062/wrk1 and uv063/wrk2 start
14:42:53  automatic_controlled_vm_startup=COMPLETE

The approximately 48-second interval from mgr1 starting to wrk1/wrk2 starting was normal. The cold-start logic waits for nas_service, then swarm_manager, then starts the workers and waits for swarm_workers. The default manager and worker readiness timeouts are 300 seconds. Do not manually start workers merely because mgr1 is already visible in virsh list.

During the first production-start attempt, Wort was accidentally rebooted while startup was still in progress. The following automatic boot recovered normally and reached the full 20-domain accepted state. Wort had no persistent previous-boot journal available, so journalctl -b -1 could not reconstruct the interrupted boot. Persistent journald may be configured separately if future boot-to-boot forensics are desired.

21.1 Known unrelated systemd warning[edit | edit source]

The target file /etc/systemd/system/kitsnet-vm-startup-complete.target contains:

Documentation=KitsNet controlled VM startup completion point

Documentation= expects documentation URIs/identifiers, not free-form prose. Systemd therefore logs Invalid URL warnings. The warning did not prevent the target from becoming active and did not affect this migration. Remove the line or replace it with a valid documentation URI as separate housekeeping.

22 Final production acceptance record[edit | edit source]

PASS: WORT PRODUCTION STARTUP COMPLETE
PASS: ALL 20 EXPECTED GUESTS ARE RUNNING
PASS: ALL MANAGED-SAVE STATES HAVE BEEN CONSUMED
PASS: T0 NVME REPLACEMENT AND EXPANSION VALIDATION COMPLETE

Final NAS state was HEALTHY_FOX, authority owner fox, Anchor BACKUP, one live production VDB, valid lease, and zero failed systemd units.

23 Rollback[edit | edit source]

23.1 Before physical NVMe replacement[edit | edit source]

If any source-image creation or verification step fails:

  1. Do not remove the original NVMe.
  2. Do not proceed to hardware replacement.
  3. Correct the image/dock problem or abandon the maintenance.

To abandon after T0 has already been deactivated:

for i in {1..10}; do echo; done
sudo vgchange -ay T0

sudo sed -i \
  's|^# KN-T0-NVME-MIGRATION ||' \
  /etc/fstab

sudo systemctl daemon-reload
sudo findmnt --verify --verbose

sudo mount /var/lib/libvirt/qemu/save

sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot auto
sudo /usr/local/sbin/kitsnet-vm-startup-control status

A controlled normal reboot may then be used to restore the accepted VM environment.

23.2 After physical replacement but before successful production validation[edit | edit source]

The original NVMe remains the primary rollback object.

If the replacement cannot pass a STOP checkpoint:

  1. Do not repair or reinitialize T0 speculatively.
  2. If Wort is operational, arm next-boot off.
  3. Power Wort off.
  4. Remove the replacement NVMe.
  5. Reinstall the untouched original NVMe.
  6. Boot protected.
  7. Validate T0 using its original device.
  8. Restore the save-mount fstab entry if it is still commented.
  9. Validate /var/lib/libvirt/qemu/save.
  10. Only then return to normal auto startup.

23.3 Duplicate-PV warning[edit | edit source]

The original NVMe and the successfully cloned replacement NVMe contain the same:

  • partition identities;
  • LVM PVID;
  • T0 VG UUID; and
  • LV UUIDs.

They must not later be attached to Wort simultaneously as normally scanned LVM devices.

Doing so would create a duplicate-PV/duplicate-VG condition.

24 Commands that are intentionally absent[edit | edit source]

A successful migration does not require any of the following:

pvcreate
vgcreate
vgextend
vgreduce
vgimport
vgimportclone
pvchange --uuid
vgchange --uuid
vgcfgrestore
pvck --repair
vgck --updatemetadata

Their unexpected apparent necessity is a reason to stop and diagnose the clone rather than continue.

25 Completion criteria[edit | edit source]

The T0 NVMe replacement is complete only when all of the following are true:

  • the replacement NVMe is installed;
  • the original NVMe remains preserved as rollback media;
  • the compressed source image remains available on the dock;
  • the image was byte-verified against the original source before removal;
  • the restored portion of the replacement was byte-verified against the image;
  • source and replacement logical sector sizes matched;
  • T0 retains the original PVID;
  • T0 retains the original VG UUID;
  • all T0 LVs retain their original LV UUIDs;
  • the partition start sector is unchanged;
  • the partition identity is unchanged;
  • the partition is expanded to the replacement capacity;
  • pvresize has exposed the added capacity as free space in T0;
  • pvck and vgck show no unexplained errors;
  • /var/lib/libvirt/qemu/save mounts normally from T0/kvm_save;
  • the managed-save inventory survived the migration;
  • a complete no-VM boot succeeded using the final storage configuration;
  • the subsequent automatic KitsNet startup completed successfully with the exact 20-domain production set;
  • all managed-save files have been consumed after successful restoration;
  • NAS ownership is coherent (HEALTHY_FOX, owner fox in the validated execution);
  • the Wort startup-completion target is active and automatic_controlled_vm_startup=COMPLETE was logged;
  • lvmdevices --check returns zero;
  • systemctl --failed reports zero failed units; and
  • expected VMs and console listeners are operational.

Only after a suitable observation period should the original NVMe or dock recovery image be considered available for reuse.