Last edited 3 weeks ago
by Peter A. Smode

KVM:T0 NVMe Replacement Runbook

Revision as of 10:14, 12 September 2026 by Peter A. Smode (talk | contribs) (Created page with "'''Status:''' Planned maintenance procedure '''Host:''' <code>wort</code> '''Purpose:''' Replace the physical NVMe device containing the <code>T0</code> LVM volume group with a larger NVMe when both devices cannot be installed simultaneously. The migration uses a verified compressed whole-device image stored on a dock-attached disk, followed by restoration to the replacement NVMe, validation of the cloned LVM identity, expansion of the partition and PV, restoration of...")
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)

Status: Planned maintenance procedure

Host: wort

Purpose: Replace the physical NVMe device containing the T0 LVM volume group with a larger NVMe when both devices cannot be installed simultaneously. The migration uses a verified compressed whole-device image stored on a dock-attached disk, followed by restoration to the replacement NVMe, validation of the cloned LVM identity, expansion of the partition and PV, restoration of the libvirt managed-save mount, and controlled return of KitsNet virtual machines to service.

Depends on:

1 Current storage design[edit | edit source]

At the time this procedure was written:

  • T0 is a single-PV volume group.
  • The T0 PV is /dev/nvme0n1p1.
  • The underlying whole device is /dev/nvme0n1.
  • T0 contains 27 logical volumes used by the Wort virtual-machine environment.
  • T0/kvm_save is a 64-GiB LV used for /var/lib/libvirt/qemu/save.
  • Wort's host operating-system filesystems are on the separate T0h VG, not on T0.
  • T0 can therefore be completely deactivated while Wort itself remains operational.

The migration deliberately clones the entire NVMe device, not merely partition 1.

A whole-device clone preserves the on-disk:

  • partition table;
  • disk GUID, if GPT;
  • partition GUID/PARTUUID;
  • LVM PV UUID;
  • T0 VG UUID;
  • LV UUIDs;
  • filesystem UUIDs;
  • LVM metadata;
  • allocated and unallocated extents; and
  • all VM disk and managed-save contents.

The replacement NVMe's physical model, serial number, WWID and other controller-level identifiers are not cloned.

2 Governing safety rules[edit | edit source]

Do not improvise past a STOP checkpoint. If a required validation does not produce the documented result, stop the procedure and diagnose the discrepancy before proceeding.

The following rules apply for the entire maintenance:

  1. The original NVMe is not erased, modified, reformatted or repurposed after removal until the replacement has passed final production validation.
  2. No pvcreate, vgcreate, vgimportclone, pvchange --uuid, vgchange --uuid or routine vgcfgrestore is used.
  3. The image is made from the whole device /dev/nvme0n1.
  4. The image is restored to the whole replacement device.
  5. T0 must be inactive while the source image is created.
  6. The replacement NVMe must be at least as large in bytes as the source NVMe.
  7. The replacement NVMe must present the same logical sector size as the source NVMe.
  8. The clone is validated before any expansion operation is performed.
  9. Partition expansion and PV expansion are separate from the cloning operation.
  10. The /var/lib/libvirt/qemu/save mount remains available until the controlled VM shutdown/managed-save operation is complete.
  11. The save mount is then unmounted and disabled in /etc/fstab until T0 has been restored and expanded.
  12. Every maintenance boot remains VM-protected.
  13. Because next-boot off is one-shot, the first administrative action after every protected maintenance boot is to arm next-boot off again for the following boot.
  14. The final production boot is the only boot for which next-boot auto is explicitly selected.

3 Migration overview[edit | edit source]

The controlled sequence is:

Normal production Wort
        |
        v
Preflight and record identities
        |
        v
Arm next boot OFF
        |
        v
Controlled managed-save / NAS shutdown
        |
        v
Unmount and disable kvm_save mount
        |
        v
Deactivate T0
        |
        v
Whole NVMe -> dd -> pigz -p 12 -> dock image
        |
        v
Verify compressed image against original NVMe
        |
        v
Power off Wort
        |
        v
Physically replace NVMe
        |
        v
Protected no-VM boot
        |
        v
Verify replacement identity / geometry
        |
        v
Dock image -> pigz -d -> dd -> replacement NVMe
        |
        v
Byte-for-byte verification of restored region
        |
        v
Protected reboot
        |
        v
Prove exact LVM clone before expansion
        |
        v
Move GPT backup header if GPT
        |
        v
Grow partition 1
        |
        v
Protected reboot
        |
        v
pvresize existing T0 PV
        |
        v
Restore and test kvm_save fstab mount
        |
        v
Protected validation reboot
        |
        v
Final validation with zero VMs
        |
        v
Select AUTO for next boot
        |
        v
Normal controlled KitsNet startup

4 Stage 0 — pre-maintenance validation[edit | edit source]

This stage is performed while Wort is operating normally and before any VM shutdown.

4.1 Confirm that this is Wort[edit | edit source]

for i in {1..10}; do echo; done
hostname
hostname -f

Expected: the host is Wort.

STOP if these commands are being executed on another host.

4.2 Define the source-device names[edit | edit source]

The following names reflect the current T0 design.

for i in {1..10}; do echo; done
SRC=/dev/nvme0n1
PV=/dev/nvme0n1p1

echo "SRC=$SRC"
echo "PV=$PV"

sudo pvs "$PV"
sudo vgs T0
sudo lvs T0

Expected:

  • /dev/nvme0n1p1 belongs to VG T0.
  • T0 has exactly one PV.

STOP if the source device or T0 topology differs.

4.3 Confirm the currently mounted libvirt save filesystem[edit | edit source]

for i in {1..10}; do echo; done
sudo findmnt --mountpoint /var/lib/libvirt/qemu/save
echo
sudo lvs -o vg_name,lv_name,lv_size,lv_attr T0/kvm_save
echo
grep -n '/var/lib/libvirt/qemu/save' /etc/fstab

The mount must be understood before continuing.

Do not unmount it yet. The production VM shutdown mechanism needs this filesystem for ordinary managed-save files.

4.4 Verify the production VM shutdown override[edit | edit source]

for i in {1..10}; do echo; done
systemctl cat libvirt-guests.service
echo
systemctl cat libvirt-guests.service |
grep -F '/usr/local/sbin/libvirt-guests-parallel.sh stop'

Expected: the effective ExecStop path includes:

/usr/local/sbin/libvirt-guests-parallel.sh stop

STOP if the KitsNet override is not in effect.

4.5 Verify required utilities[edit | edit source]

for i in {1..10}; do echo; done
for cmd in \
  dd pigz cmp blockdev lsblk sfdisk blkid \
  pvs vgs lvs vgchange pvresize pvck vgck vgcfgbackup \
  lvmconfig lvmdevices pvscan \
  findmnt mountpoint growpart
do
    if command -v "$cmd" >/dev/null 2>&1; then
        printf 'OK      %s -> %s\n' "$cmd" "$(command -v "$cmd")"
    else
        printf 'MISSING %s\n' "$cmd"
    fi
done

echo
if command -v nvme >/dev/null 2>&1; then
    echo "OPTIONAL nvme -> $(command -v nvme)"
else
    echo "OPTIONAL nvme command is not installed"
fi

echo
if command -v sgdisk >/dev/null 2>&1; then
    echo "sgdisk -> $(command -v sgdisk)"
else
    echo "sgdisk is not installed; it is required if the source disk is GPT"
fi

STOP if any mandatory utility is missing.

If the source is GPT, sgdisk is also mandatory.

Install/repair required tooling before entering the maintenance outage.

4.6 Record source geometry[edit | edit source]

for i in {1..10}; do echo; done
SRC=/dev/nvme0n1
PV=/dev/nvme0n1p1

echo "===== SOURCE DEVICE ====="
sudo lsblk -o NAME,PATH,SIZE,TYPE,FSTYPE,PTTYPE,MODEL,SERIAL,WWN,LOG-SEC,PHY-SEC "$SRC"

echo
echo "===== SOURCE SIZE IN BYTES ====="
sudo blockdev --getsize64 "$SRC"

echo
echo "===== SOURCE LOGICAL SECTOR SIZE ====="
sudo blockdev --getss "$SRC"

echo
echo "===== PARTITION START ====="
sudo lsblk -no START "$PV"

echo
echo "===== PARTUUID ====="
sudo blkid -s PARTUUID -o value "$PV"

echo
echo "===== PARTITION TABLE TYPE ====="
sudo blkid -p -s PTTYPE -o value "$SRC"

The partition-table type is expected to be either:

gpt

or:

dos

Do not assume GPT until this command has confirmed it.

4.7 Configure and validate the dock destination[edit | edit source]

Set DOCK to the actual mounted dock filesystem.

This is the only site-specific path that must be substituted in the command blocks.

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
SRC=/dev/nvme0n1

echo "DOCK=$DOCK"

echo
echo "===== DOCK MOUNT ====="
findmnt --mountpoint "$DOCK"

echo
echo "===== DOCK FILESYSTEM ====="
df -Th "$DOCK"

echo
echo "===== DOCK FREE BYTES ====="
DOCK_FREE=$(df -B1 --output=avail "$DOCK" | tail -1 | tr -d ' ')
SRC_BYTES=$(sudo blockdev --getsize64 "$SRC")

echo "source_bytes=$SRC_BYTES"
echo "dock_free_bytes=$DOCK_FREE"

NEED_BYTES=$((SRC_BYTES + SRC_BYTES / 20))
echo "recommended_minimum_free_bytes=$NEED_BYTES"

if [ "$DOCK_FREE" -ge "$NEED_BYTES" ]; then
    echo "PASS: dock has at least source size plus 5 percent"
else
    echo "STOP: insufficient worst-case dock capacity"
fi

The 5-percent allowance ensures that the procedure does not depend on the source being compressible.

STOP unless the dock mount and capacity are correct.

4.8 Create the migration directory[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"

sudo mkdir -p "$MIG"
sudo touch "$MIG/.write-test"
sudo rm -f "$MIG/.write-test"

echo "MIG=$MIG"
sudo ls -ld "$MIG"

4.9 Capture the pre-migration identity and recovery metadata[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
SRC=/dev/nvme0n1
PV=/dev/nvme0n1p1

PTTYPE=$(sudo blkid -p -s PTTYPE -o value "$SRC")

sudo blockdev --getsize64 "$SRC" |
sudo tee "$MIG/source.bytes"

sudo blockdev --getss "$SRC" |
sudo tee "$MIG/source.logical-sector-size"

printf '%s\n' "$PTTYPE" |
sudo tee "$MIG/source.pttype"

sudo sfdisk --dump "$SRC" |
sudo tee "$MIG/source.sfdisk"

sudo lsblk -no START "$PV" |
tr -d ' ' |
sudo tee "$MIG/partition-start.before"

sudo blkid -s PARTUUID -o value "$PV" |
sudo tee "$MIG/partuuid.before"

sudo pvs --noheadings -o pv_uuid "$PV" |
tr -d ' ' |
sudo tee "$MIG/pvid.before"

sudo vgs --noheadings -o vg_uuid T0 |
tr -d ' ' |
sudo tee "$MIG/vguid.before"

sudo lvs --noheadings --separator '|' -o lv_name,lv_uuid T0 |
sed 's/^[[:space:]]*//;s/[[:space:]]*$//' |
sort |
sudo tee "$MIG/lvuuids.before"

sudo blkid -s UUID -o value /dev/T0/kvm_save |
sudo tee "$MIG/kvm-save-fsuuid.before"

sudo findmnt --mountpoint /var/lib/libvirt/qemu/save \
  -o SOURCE,TARGET,FSTYPE,OPTIONS |
sudo tee "$MIG/kvm-save-mount.before"

sudo cp -a /etc/fstab "$MIG/fstab.before"

sudo vgcfgbackup -f "$MIG/T0.vgcfg" T0

if sudo test -d /etc/lvm/devices; then
    sudo cp -a /etc/lvm/devices "$MIG/lvm-devices.before"
fi

if [ "$PTTYPE" = "gpt" ]; then
    sudo sgdisk --backup="$MIG/source.gpt" "$SRC"

    sudo sgdisk -p "$SRC" |
    awk '/Disk identifier \(GUID\):/{print $4}' |
    sudo tee "$MIG/disk-guid.before"

    sudo sgdisk -v "$SRC"
fi

For GPT, sgdisk -v must not report existing structural corruption.

STOP if the source partition table or LVM metadata already appears damaged.

5 Stage 1 — quiesce all Wort VMs[edit | edit source]

5.1 Arm the next boot for zero VMs[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control status

Verify that the next boot is configured for off mode.

5.2 Run the accepted production shutdown sequence[edit | edit source]

Stopping libvirt-guests.service invokes the KitsNet production wrapper. The wrapper performs managed-save processing for ordinary guests and then invokes the owner-aware NAS shutdown mechanism.

for i in {1..10}; do echo; done
sudo systemctl stop libvirt-guests.service
RC=$?

echo
echo "libvirt-guests stop rc=$RC"
echo
systemctl status libvirt-guests.service --no-pager -l

Required result:

libvirt-guests stop rc=0

STOP if the stop operation fails.

Do not manually destroy a guest to force the procedure onward.

5.3 Prove that no VM remains running[edit | edit source]

for i in {1..10}; do echo; done
echo "===== RUNNING LIBVIRT GUESTS ====="
sudo virsh list --state-running --name

echo
echo "===== RUNNING GUEST COUNT ====="
RUNNING=$(sudo virsh list --state-running --name |
  sed '/^[[:space:]]*$/d' |
  wc -l)

echo "running_guests=$RUNNING"

echo
echo "===== QEMU PROCESSES ====="
ps -eo pid,args |
grep -E '[q]emu-system|[q]emu-kvm' || true

Required result:

running_guests=0

STOP if any guest or QEMU process remains active.

5.4 Record the managed-save files before unmounting T0/kvm_save[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"

sudo find /var/lib/libvirt/qemu/save \
  -maxdepth 1 \
  -type f \
  -printf '%f|%s\n' |
sort |
sudo tee "$MIG/managed-save.before"

The exact count depends on which ordinary guests were running before shutdown. The file list will later be compared after restoration.

6 Stage 2 — remove the libvirt save mount from the maintenance boot path[edit | edit source]

6.1 Unmount the save filesystem[edit | edit source]

for i in {1..10}; do echo; done
sudo umount /var/lib/libvirt/qemu/save
RC=$?

echo "umount rc=$RC"

if mountpoint -q /var/lib/libvirt/qemu/save; then
    echo "STOP: /var/lib/libvirt/qemu/save is still mounted"
else
    echo "PASS: save filesystem is unmounted"
fi

STOP unless umount rc=0 and the filesystem is no longer a mountpoint.

6.2 Verify exactly one active fstab entry exists before commenting it[edit | edit source]

for i in {1..10}; do echo; done
ACTIVE_FSTAB=$(
  awk '
    $0 !~ /^[[:space:]]*#/ &&
    $2 == "/var/lib/libvirt/qemu/save" { n++ }
    END { print n+0 }
  ' /etc/fstab
)

echo "active_save_fstab_entries=$ACTIVE_FSTAB"
grep -n '/var/lib/libvirt/qemu/save' /etc/fstab

Required result:

active_save_fstab_entries=1

STOP otherwise.

6.3 Comment the save mount[edit | edit source]

for i in {1..10}; do echo; done
sudo sed -i \
  '\|^[[:space:]]*[^#].*[[:space:]]/var/lib/libvirt/qemu/save[[:space:]]| s|^|# KN-T0-NVME-MIGRATION |' \
  /etc/fstab

echo "===== RESULTING FSTAB ENTRY ====="
grep -n '/var/lib/libvirt/qemu/save' /etc/fstab

echo
echo "===== ACTIVE ENTRIES REMAINING ====="
awk '
  $0 !~ /^[[:space:]]*#/ &&
  $2 == "/var/lib/libvirt/qemu/save" { print }
' /etc/fstab

echo
sudo systemctl daemon-reload
sudo findmnt --verify --verbose

There must be no active fstab entry for the save mount.

The original line must remain present with the prefix:

# KN-T0-NVME-MIGRATION

7 Stage 3 — deactivate T0 completely[edit | edit source]

for i in {1..10}; do echo; done
sudo vgchange -an T0
RC=$?

echo
echo "vgchange rc=$RC"

echo
echo "===== T0 LV STATE ====="
sudo lvs -o vg_name,lv_name,lv_active,lv_attr,lv_size T0

echo
ACTIVE=$(
  sudo lvs --noheadings -o lv_active T0 |
  awk '$1 == "active" { n++ } END { print n+0 }'
)

echo "active_T0_LVs=$ACTIVE"

Required results:

vgchange rc=0
active_T0_LVs=0

STOP if T0 cannot be completely deactivated.

Do not image a partially active T0.

8 Stage 4 — create the compressed whole-NVMe image[edit | edit source]

8.1 Revalidate source and destination immediately before imaging[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
SRC=/dev/nvme0n1
PV=/dev/nvme0n1p1
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"

echo "===== SOURCE ====="
sudo lsblk -o NAME,PATH,SIZE,TYPE,FSTYPE,MODEL,SERIAL,LOG-SEC,PHY-SEC "$SRC"

echo
echo "===== SOURCE PV ====="
sudo pvs "$PV"

echo
echo "===== IMAGE DESTINATION ====="
echo "$IMG"

if sudo test -e "$IMG"; then
    echo "STOP: image file already exists; it will NOT be overwritten automatically"
else
    echo "PASS: image pathname is unused"
fi

echo
df -h "$DOCK"

STOP if the source device is not the known T0 NVMe or if the image filename already exists unexpectedly.

8.2 Create the image[edit | edit source]

This command:

  • reads the whole source NVMe using dd;
  • feeds the raw byte stream into pigz;
  • allows pigz to use up to 12 compression threads;
  • uses normal gzip compression level 6; and
  • enables pipeline failure propagation.

No conv=noerror option is used. A source read error is a migration failure and must not be silently padded or ignored.

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
SRC=/dev/nvme0n1
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"

sudo env SRC="$SRC" IMG="$IMG" \
bash -o pipefail -c '
    if [ -e "$IMG" ]; then
        echo "STOP: refusing to overwrite existing image: $IMG"
        exit 2
    fi

    dd if="$SRC" bs=16M iflag=fullblock status=progress |
    pigz -p 12 -6 > "$IMG"
'

RC=$?

echo
echo "image_pipeline_rc=$RC"

sudo sync

echo
sudo ls -lh "$IMG"

Required result:

image_pipeline_rc=0

STOP otherwise.

9 Stage 5 — validate the image before removing the original NVMe[edit | edit source]

9.1 Test the gzip stream[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"

sudo pigz -t "$IMG"
RC=$?

echo "pigz_test_rc=$RC"

Required result:

pigz_test_rc=0

9.2 Compare the decompressed image byte-for-byte with the original NVMe[edit | edit source]

This is the definitive pre-swap verification.

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
SRC=/dev/nvme0n1
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"

sudo env SRC="$SRC" IMG="$IMG" \
bash -o pipefail -c '
    pigz -dc "$IMG" |
    cmp - "$SRC"
'

RC=$?

echo
echo "source_image_cmp_rc=$RC"

cmp normally prints nothing when the inputs are identical.

Required result:

source_image_cmp_rc=0

STOP if the comparison is anything other than zero.

9.3 Record a checksum of the compressed recovery artifact[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"

sudo sha256sum "$IMG" |
sudo tee "$IMG.sha256"

sudo ls -lh "$IMG" "$IMG.sha256"

At this point there are two independent recovery assets:

  1. the untouched original physical NVMe; and
  2. the byte-verified compressed whole-device image.

10 Stage 6 — power off Wort and replace the NVMe[edit | edit source]

Re-arm protected startup immediately before shutdown, even though it should already be armed.

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
sudo /usr/local/sbin/kitsnet-vm-startup-control status

Verify off for the next boot.

Then power off:

for i in {1..10}; do echo; done
sudo systemctl poweroff

After Wort is completely powered off:

  1. Remove the original T0 NVMe.
  2. Label and retain it intact.
  3. Install the replacement NVMe.
  4. Do not connect the original NVMe simultaneously with its clone.
  5. Keep the dock recovery disk connected.

11 Stage 7 — first protected boot with the replacement NVMe[edit | edit source]

Wort should boot from T0h even though T0 is currently absent from the blank replacement NVMe.

11.1 Immediately protect the following boot[edit | edit source]

The off request that produced this boot has now been consumed.

The first maintenance action is therefore:

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control status
echo
echo "running_guests=$(sudo virsh list --state-running --name 2>/dev/null | sed '/^[[:space:]]*$/d' | wc -l)"

Required result:

running_guests=0

The following boot must also show off as armed.

11.2 Identify the replacement NVMe[edit | edit source]

Do not assume that the Linux device name is correct merely because it is nvme0n1.

for i in {1..10}; do echo; done
sudo lsblk -d -o NAME,PATH,SIZE,MODEL,SERIAL,WWN,LOG-SEC,PHY-SEC

echo
if command -v nvme >/dev/null 2>&1; then
    sudo nvme list
fi

Physically correlate model, serial number and capacity with the installed replacement.

Only after positive identification set:

for i in {1..10}; do echo; done
NEW=/dev/nvme0n1

sudo lsblk -o NAME,PATH,SIZE,TYPE,FSTYPE,MOUNTPOINTS,MODEL,SERIAL,WWN,LOG-SEC,PHY-SEC "$NEW"

STOP if there is any uncertainty regarding the replacement device.

11.3 Verify replacement capacity and logical sector size[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
NEW=/dev/nvme0n1

OLD_BYTES=$(sudo cat "$MIG/source.bytes")
OLD_LOGSEC=$(sudo cat "$MIG/source.logical-sector-size")

NEW_BYTES=$(sudo blockdev --getsize64 "$NEW")
NEW_LOGSEC=$(sudo blockdev --getss "$NEW")

echo "old_bytes=$OLD_BYTES"
echo "new_bytes=$NEW_BYTES"
echo
echo "old_logical_sector_size=$OLD_LOGSEC"
echo "new_logical_sector_size=$NEW_LOGSEC"

echo
if [ "$NEW_BYTES" -ge "$OLD_BYTES" ]; then
    echo "PASS: replacement is large enough"
else
    echo "STOP: replacement is smaller than source"
fi

if [ "$NEW_LOGSEC" -eq "$OLD_LOGSEC" ]; then
    echo "PASS: logical sector sizes match"
else
    echo "STOP: logical sector size mismatch"
fi

Both tests must pass.

A logical-sector-size mismatch is a STOP condition for this whole-device cloning procedure.

11.4 Verify that nothing on the replacement NVMe is mounted[edit | edit source]

for i in {1..10}; do echo; done
NEW=/dev/nvme0n1

echo "===== REPLACEMENT DEVICE TREE ====="
sudo lsblk -o NAME,PATH,SIZE,TYPE,FSTYPE,MOUNTPOINTS "$NEW"

echo
echo "===== NONEMPTY MOUNTPOINTS ====="
sudo lsblk -nr -o MOUNTPOINTS "$NEW" |
sed '/^[[:space:]]*$/d'

The second section must be empty.

STOP if any replacement-NVMe filesystem is mounted.

11.5 Verify the recovery image again before destructive writing[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"

sudo sha256sum -c "$IMG.sha256"
RC=$?

echo "image_sha256_rc=$RC"

Required result:

image_sha256_rc=0

12 Stage 8 — restore the whole-device image to the replacement NVMe[edit | edit source]

12.1 Remove stale partition-table structures from the replacement only[edit | edit source]

This prevents a previously used replacement device from retaining an unrelated backup GPT at its physical end.

This command is destructive.

Run it only after positive identification of NEW.

for i in {1..10}; do echo; done
NEW=/dev/nvme0n1

echo "ABOUT TO DESTROY EXISTING PARTITION TABLES ON:"
sudo lsblk -d -o NAME,PATH,SIZE,MODEL,SERIAL,WWN "$NEW"

echo
read -r -p 'Type WIPE_REPLACEMENT_NVME to continue: ' ANSWER

if [ "$ANSWER" = "WIPE_REPLACEMENT_NVME" ]; then
    sudo sgdisk --zap-all "$NEW"
    echo "sgdisk zap rc=$?"
else
    echo "ABORTED: replacement NVMe was not modified"
fi

Do not continue unless the zap operation completed successfully.

12.2 Restore the image[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"
NEW=/dev/nvme0n1

echo "IMAGE=$IMG"
echo "DESTINATION=$NEW"

echo
sudo lsblk -d -o NAME,PATH,SIZE,MODEL,SERIAL,WWN "$NEW"

echo
read -r -p 'Type RESTORE_T0_IMAGE to overwrite the replacement NVMe: ' ANSWER

if [ "$ANSWER" = "RESTORE_T0_IMAGE" ]; then
    sudo env IMG="$IMG" NEW="$NEW" \
    bash -o pipefail -c '
        pigz -dc "$IMG" |
        dd of="$NEW" bs=16M iflag=fullblock status=progress conv=fsync
    '

    RC=$?
    echo
    echo "restore_pipeline_rc=$RC"
    sudo sync
else
    echo "ABORTED: image was not restored"
fi

Required result:

restore_pipeline_rc=0

STOP otherwise.

12.3 Byte-verify the restored portion of the larger NVMe[edit | edit source]

Because the replacement is larger, compare exactly the number of bytes that existed on the original NVMe.

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
IMG="$MIG/wort-T0-nvme0n1.full.img.gz"
NEW=/dev/nvme0n1
OLD_BYTES=$(sudo cat "$MIG/source.bytes")

sudo env IMG="$IMG" NEW="$NEW" OLD_BYTES="$OLD_BYTES" \
bash -o pipefail -c '
    pigz -dc "$IMG" |
    cmp -n "$OLD_BYTES" - "$NEW"
'

RC=$?

echo
echo "restored_image_cmp_rc=$RC"

Required result:

restored_image_cmp_rc=0

This proves that every byte belonging to the original disk image was written identically to the replacement.

STOP otherwise.

12.4 Do not expand the disk yet[edit | edit source]

At this point the replacement contains an exact clone of the old NVMe.

If the source uses GPT, a verification command may report that the backup GPT is not at the physical end of the larger device. That condition is expected at this stage.

No correction or expansion is performed until after a clean reboot proves that LVM sees the clone correctly.

12.5 Reboot protected[edit | edit source]

The following boot should already be armed off from the beginning of this maintenance boot. Verify before rebooting:

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control status
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
sudo /usr/local/sbin/kitsnet-vm-startup-control status

Then:

for i in {1..10}; do echo; done
sudo systemctl reboot

13 Stage 9 — prove the exact clone before expansion[edit | edit source]

13.1 Immediately protect the following boot[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control status
echo
echo "running_guests=$(sudo virsh list --state-running --name 2>/dev/null | sed '/^[[:space:]]*$/d' | wc -l)"

Required:

running_guests=0

13.2 Let normal boot discovery run first[edit | edit source]

Do not issue LVM repair commands preemptively.

Begin with read-only inspection:

for i in {1..10}; do echo; done
sudo pvs
echo
sudo vgs
echo
sudo lvs T0 2>&1 || true

The expected outcome is that /dev/nvme0n1p1 appears as the existing T0 PV.

If T0 appears normally, do not run any additional scan/refresh command.

13.3 Only if T0 is missing: refresh the LVM device identity[edit | edit source]

The replacement NVMe has a different physical WWID/serial number. If Wort uses /etc/lvm/devices/system.devices, the old hardware ID may prevent immediate discovery even though the cloned PVID is correct.

First inspect:

for i in {1..10}; do echo; done
sudo lvmconfig --type current devices/use_devicesfile
echo
sudo lvmdevices 2>&1 || true

If use_devicesfile=1, perform the supported device-ID refresh:

for i in {1..10}; do echo; done
sudo lvmdevices --check --refresh
echo
sudo lvmdevices --update --refresh
RC=$?

echo
echo "lvmdevices_refresh_rc=$RC"

echo
sudo pvscan --cache /dev/nvme0n1p1

echo
sudo pvs

If the devices file is not in use, a normal cache refresh may be used:

for i in {1..10}; do echo; done
sudo pvscan --cache /dev/nvme0n1p1
echo
sudo pvs

STOP if T0 still does not appear.

Do not initialize or repair the PV merely because discovery failed.

13.4 Check LVM metadata without repairing it[edit | edit source]

for i in {1..10}; do echo; done
PV=/dev/nvme0n1p1

sudo pvck "$PV"
PVCK_RC=$?

echo
sudo vgck T0
VGCK_RC=$?

echo
echo "pvck_rc=$PVCK_RC"
echo "vgck_rc=$VGCK_RC"

Do not use pvck --repair or vgck --updatemetadata during normal migration validation.

STOP on unexplained metadata errors.

13.5 Compare the cloned PV UUID[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
PV=/dev/nvme0n1p1

sudo pvs --noheadings -o pv_uuid "$PV" |
tr -d ' ' > /tmp/T0-pvid.after

echo "===== BEFORE ====="
sudo cat "$MIG/pvid.before"

echo "===== AFTER ====="
cat /tmp/T0-pvid.after

echo "===== DIFF ====="
sudo diff -u "$MIG/pvid.before" /tmp/T0-pvid.after

Required: no difference.

13.6 Compare the cloned VG UUID[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"

sudo vgs --noheadings -o vg_uuid T0 |
tr -d ' ' > /tmp/T0-vguid.after

sudo diff -u "$MIG/vguid.before" /tmp/T0-vguid.after

Required: no difference.

13.7 Compare all LV UUIDs[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"

sudo lvs --noheadings --separator '|' -o lv_name,lv_uuid T0 |
sed 's/^[[:space:]]*//;s/[[:space:]]*$//' |
sort > /tmp/T0-lvuuids.after

sudo diff -u "$MIG/lvuuids.before" /tmp/T0-lvuuids.after

Required: no difference.

13.8 Compare the partition UUID and start sector[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
PV=/dev/nvme0n1p1

sudo blkid -s PARTUUID -o value "$PV" > /tmp/partuuid.after

sudo lsblk -no START "$PV" |
tr -d ' ' > /tmp/partition-start.after

echo "===== PARTUUID DIFF ====="
sudo diff -u "$MIG/partuuid.before" /tmp/partuuid.after

echo
echo "===== START-SECTOR DIFF ====="
sudo diff -u "$MIG/partition-start.before" /tmp/partition-start.after

Both comparisons must show no difference.

13.9 For GPT: compare the cloned disk GUID[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
NEW=/dev/nvme0n1
PTTYPE=$(sudo cat "$MIG/source.pttype")

if [ "$PTTYPE" = "gpt" ]; then
    sudo sgdisk -p "$NEW" |
    awk '/Disk identifier \(GUID\):/{print $4}' > /tmp/disk-guid.after

    sudo diff -u "$MIG/disk-guid.before" /tmp/disk-guid.after
else
    echo "Source partition table is $PTTYPE; GPT disk-GUID test does not apply"
fi

For GPT, the GUID must match.

13.10 Activate T0 temporarily and verify kvm_save filesystem identity[edit | edit source]

Do not mount it yet.

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"

sudo vgchange -ay T0
RC=$?

echo "vgchange_activate_rc=$RC"

sudo blkid -s UUID -o value /dev/T0/kvm_save > /tmp/kvm-save-fsuuid.after

echo
sudo diff -u "$MIG/kvm-save-fsuuid.before" /tmp/kvm-save-fsuuid.after

echo
mountpoint -q /var/lib/libvirt/qemu/save &&
    echo "STOP: save filesystem unexpectedly mounted" ||
    echo "PASS: save filesystem remains unmounted"

The filesystem UUID must be identical and the filesystem must remain unmounted.

At this point the clone has been proven.

14 Stage 10 — expand the replacement partition[edit | edit source]

Before modifying the partition table, deactivate T0 again:

for i in {1..10}; do echo; done
sudo vgchange -an T0
RC=$?

echo "vgchange_deactivate_rc=$RC"

ACTIVE=$(
  sudo lvs --noheadings -o lv_active T0 |
  awk '$1 == "active" { n++ } END { print n+0 }'
)

echo "active_T0_LVs=$ACTIVE"

Required:

vgchange_deactivate_rc=0
active_T0_LVs=0

14.1 If the source is GPT, move the backup GPT to the real end of the replacement[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
NEW=/dev/nvme0n1
PTTYPE=$(sudo cat "$MIG/source.pttype")

if [ "$PTTYPE" = "gpt" ]; then
    echo "Moving backup GPT to the actual end of $NEW"
    sudo sgdisk -e "$NEW"
    RC=$?
    echo "sgdisk_move_header_rc=$RC"

    echo
    sudo sgdisk -v "$NEW"
elif [ "$PTTYPE" = "dos" ]; then
    echo "Source uses DOS/MBR; no GPT backup-header relocation is required"
else
    echo "STOP: unexpected partition-table type: $PTTYPE"
fi

For GPT, sgdisk_move_header_rc must be zero.

14.2 Dry-run partition growth[edit | edit source]

for i in {1..10}; do echo; done
NEW=/dev/nvme0n1

sudo growpart -N "$NEW" 1
RC=$?

echo
echo "growpart_dry_run_rc=$RC"

The dry run should show that partition 1 can grow into the additional capacity.

STOP if the proposed change is not exactly an extension of partition 1 with the original starting sector preserved.

14.3 Grow partition 1[edit | edit source]

for i in {1..10}; do echo; done
NEW=/dev/nvme0n1

sudo growpart "$NEW" 1
RC=$?

echo
echo "growpart_rc=$RC"

echo
sudo sfdisk --dump "$NEW"

Required:

growpart_rc=0

14.4 Verify that the partition start and PARTUUID did not change[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
PV=/dev/nvme0n1p1

sudo lsblk -no START "$PV" |
tr -d ' ' > /tmp/partition-start.grown

sudo blkid -s PARTUUID -o value "$PV" > /tmp/partuuid.grown

echo "===== START-SECTOR DIFF ====="
sudo diff -u "$MIG/partition-start.before" /tmp/partition-start.grown

echo
echo "===== PARTUUID DIFF ====="
sudo diff -u "$MIG/partuuid.before" /tmp/partuuid.grown

Both must show no difference.

For GPT, verify structure again:

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
NEW=/dev/nvme0n1
PTTYPE=$(sudo cat "$MIG/source.pttype")

if [ "$PTTYPE" = "gpt" ]; then
    sudo sgdisk -v "$NEW"
fi

Do not run pvresize yet.

The next reboot establishes the enlarged partition geometry from a clean kernel/device discovery cycle.

14.5 Reboot protected[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
sudo /usr/local/sbin/kitsnet-vm-startup-control status

Then:

for i in {1..10}; do echo; done
sudo systemctl reboot

15 Stage 11 — resize the existing T0 PV[edit | edit source]

15.1 Immediately protect the following boot[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off

echo
sudo /usr/local/sbin/kitsnet-vm-startup-control status

echo
echo "running_guests=$(sudo virsh list --state-running --name 2>/dev/null | sed '/^[[:space:]]*$/d' | wc -l)"

Required:

running_guests=0

15.2 Confirm the kernel sees the enlarged partition[edit | edit source]

for i in {1..10}; do echo; done
NEW=/dev/nvme0n1
PV=/dev/nvme0n1p1

sudo lsblk -o NAME,PATH,SIZE,TYPE,FSTYPE,PTTYPE,START,LOG-SEC,PHY-SEC "$NEW"

echo
sudo pvs -o pv_name,pv_uuid,vg_name,pv_size,pv_free "$PV"

echo
sudo vgs -o vg_name,vg_uuid,vg_size,vg_free T0

The partition should now show the new size.

The LVM PV may still show the original size. That is expected until pvresize is executed.

15.3 Deactivate T0 before changing PV geometry[edit | edit source]

for i in {1..10}; do echo; done
sudo vgchange -an T0
RC=$?

echo "vgchange_deactivate_rc=$RC"

Required:

vgchange_deactivate_rc=0

15.4 Resize the existing PV[edit | edit source]

This changes the usable size of the existing PVID. It does not create or replace the PV.

for i in {1..10}; do echo; done
PV=/dev/nvme0n1p1

sudo pvresize "$PV"
RC=$?

echo
echo "pvresize_rc=$RC"

echo
sudo pvs -o pv_name,pv_uuid,vg_name,pv_size,pv_free "$PV"

echo
sudo vgs -o vg_name,vg_uuid,vg_size,vg_free T0

Required:

pvresize_rc=0

T0 should now show substantial new free space corresponding to the larger NVMe.

15.5 Revalidate LVM metadata and identity[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"
PV=/dev/nvme0n1p1

sudo pvck "$PV"
echo
sudo vgck T0

echo
sudo pvs --noheadings -o pv_uuid "$PV" |
tr -d ' ' > /tmp/T0-pvid.final

sudo vgs --noheadings -o vg_uuid T0 |
tr -d ' ' > /tmp/T0-vguid.final

sudo lvs --noheadings --separator '|' -o lv_name,lv_uuid T0 |
sed 's/^[[:space:]]*//;s/[[:space:]]*$//' |
sort > /tmp/T0-lvuuids.final

echo "===== PVID ====="
sudo diff -u "$MIG/pvid.before" /tmp/T0-pvid.final

echo
echo "===== VG UUID ====="
sudo diff -u "$MIG/vguid.before" /tmp/T0-vguid.final

echo
echo "===== LV UUIDS ====="
sudo diff -u "$MIG/lvuuids.before" /tmp/T0-lvuuids.final

All UUID comparisons must remain identical.

16 Stage 12 — restore the libvirt save filesystem[edit | edit source]

16.1 Activate T0[edit | edit source]

for i in {1..10}; do echo; done
sudo vgchange -ay T0
RC=$?

echo "vgchange_activate_rc=$RC"

echo
sudo lvs -o vg_name,lv_name,lv_active,lv_attr,lv_size T0

Required:

vgchange_activate_rc=0

16.2 Restore only the fstab line that this procedure commented[edit | edit source]

First verify that the marker exists exactly as expected:

for i in {1..10}; do echo; done
grep -n '^# KN-T0-NVME-MIGRATION ' /etc/fstab

There should be exactly one line, and it must be the original /var/lib/libvirt/qemu/save entry.

Restore it:

for i in {1..10}; do echo; done
sudo sed -i \
  's|^# KN-T0-NVME-MIGRATION ||' \
  /etc/fstab

echo "===== RESTORED ENTRY ====="
grep -n '/var/lib/libvirt/qemu/save' /etc/fstab

echo
sudo systemctl daemon-reload
sudo findmnt --verify --verbose

16.3 Mount the save filesystem manually[edit | edit source]

for i in {1..10}; do echo; done
sudo mount /var/lib/libvirt/qemu/save
RC=$?

echo "mount rc=$RC"

echo
sudo findmnt --mountpoint /var/lib/libvirt/qemu/save \
  -o SOURCE,TARGET,FSTYPE,SIZE,USED,AVAIL,OPTIONS

Required:

mount rc=0

The source must resolve to the restored T0/kvm_save LV.

16.4 Verify the filesystem UUID[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"

sudo blkid -s UUID -o value /dev/T0/kvm_save > /tmp/kvm-save-fsuuid.final

sudo diff -u \
  "$MIG/kvm-save-fsuuid.before" \
  /tmp/kvm-save-fsuuid.final

Required: no difference.

16.5 Verify the managed-save inventory survived[edit | edit source]

for i in {1..10}; do echo; done
DOCK=/mnt/REPLACE_WITH_ACTUAL_DOCK_MOUNT
MIG="$DOCK/wort-T0-nvme-migration"

sudo find /var/lib/libvirt/qemu/save \
  -maxdepth 1 \
  -type f \
  -printf '%f|%s\n' |
sort > /tmp/managed-save.final

sudo diff -u \
  "$MIG/managed-save.before" \
  /tmp/managed-save.final

Required: no difference.

17 Stage 13 — final no-VM validation reboot[edit | edit source]

This reboot tests the final intended storage configuration:

  • replacement NVMe installed;
  • clone proven;
  • partition expanded;
  • T0 PV expanded;
  • T0 identities unchanged;
  • fstab restored; and
  • /var/lib/libvirt/qemu/save restored.

No VM is yet permitted to start.

17.1 Arm the protected boot[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
sudo /usr/local/sbin/kitsnet-vm-startup-control status

Then:

for i in {1..10}; do echo; done
sudo systemctl reboot

18 Stage 14 — final validation with all VMs still off[edit | edit source]

18.1 Immediately preserve protection against an unexpected additional reboot[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot off
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control status

18.2 Verify zero VMs[edit | edit source]

for i in {1..10}; do echo; done
echo "running_guests=$(sudo virsh list --state-running --name 2>/dev/null | sed '/^[[:space:]]*$/d' | wc -l)"

Required:

running_guests=0

18.3 Verify final T0 configuration[edit | edit source]

for i in {1..10}; do echo; done
sudo pvs -o pv_name,pv_uuid,vg_name,pv_size,pv_free
echo
sudo vgs -o vg_name,vg_uuid,pv_count,lv_count,vg_size,vg_free T0
echo
sudo lvs -o vg_name,lv_name,lv_uuid,lv_size,lv_attr T0

Confirm:

  • T0 has exactly one PV;
  • the PVID is unchanged;
  • the VG UUID is unchanged;
  • all expected LVs exist;
  • their UUIDs are unchanged; and
  • the additional replacement-NVMe space now appears as T0 free space.

18.4 Verify the save mount survived a clean boot[edit | edit source]

for i in {1..10}; do echo; done
sudo findmnt --mountpoint /var/lib/libvirt/qemu/save \
  -o SOURCE,TARGET,FSTYPE,SIZE,USED,AVAIL,OPTIONS

echo
sudo find /var/lib/libvirt/qemu/save \
  -maxdepth 1 \
  -type f \
  -printf '%f|%s\n' |
sort

18.5 Final metadata checks[edit | edit source]

for i in {1..10}; do echo; done
sudo pvck /dev/nvme0n1p1
PVCK_RC=$?

echo
sudo vgck T0
VGCK_RC=$?

echo
echo "pvck_rc=$PVCK_RC"
echo "vgck_rc=$VGCK_RC"

18.6 Check host health[edit | edit source]

for i in {1..10}; do echo; done
echo "===== FAILED SYSTEMD UNITS ====="
systemctl --failed --no-pager

echo
echo "===== T0 ====="
sudo pvs /dev/nvme0n1p1
sudo vgs T0

echo
echo "===== SAVE MOUNT ====="
sudo findmnt --mountpoint /var/lib/libvirt/qemu/save

echo
echo "===== STARTUP CONTROL ====="
sudo /usr/local/sbin/kitsnet-vm-startup-control status

Do not return to production if there is an unexplained storage, LVM, mount or host failure.

19 Stage 15 — return KitsNet to normal automatic startup[edit | edit source]

Only after every preceding validation has passed should the next boot be changed from off to auto.

19.1 Explicitly select normal automatic startup[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot auto
echo
sudo /usr/local/sbin/kitsnet-vm-startup-control status

Verify that the next boot is configured for normal automatic controlled startup.

19.2 Perform the production boot[edit | edit source]

for i in {1..10}; do echo; done
sudo systemctl reboot

The normal Wort controller now owns startup sequencing.

Do not manually start Fox, Anchor, mgr1, workers, ordinary VMs or socat listeners unless the established startup-control procedure reports a failure requiring diagnosis.

20 Stage 16 — post-return production validation[edit | edit source]

20.1 Startup-controller status[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-vm-startup-control status

echo
systemctl status kitsnet-vm-auto-start.service --no-pager -l

echo
sudo journalctl -b -u kitsnet-vm-auto-start.service --no-pager

20.2 NAS status[edit | edit source]

for i in {1..10}; do echo; done
sudo /usr/local/sbin/kitsnet-nas-ha-status

The expected normal state is HEALTHY_FOX unless a legitimate accepted HEALTHY_ANCHOR state exists.

20.3 VM counts[edit | edit source]

for i in {1..10}; do echo; done
echo "running_guests=$(sudo virsh list --state-running --name | sed '/^[[:space:]]*$/d' | wc -l)"
echo "autostart_links=$(sudo find /etc/libvirt/qemu/autostart -maxdepth 1 -type l 2>/dev/null | wc -l)"
echo "managed_save_files=$(sudo find /var/lib/libvirt/qemu/save -maxdepth 1 -type f 2>/dev/null | wc -l)"
echo "held_autostarts=$(sudo find /var/lib/kitsnet-vm-startup/autostart-hold -maxdepth 1 -type l 2>/dev/null | wc -l)"

At the accepted Wort production baseline the normal complete state is:

running_guests=19
autostart_links=17
managed_save_files=0
held_autostarts=0

If the intentionally-running VM inventory has changed since this runbook was written, compare with the current documented production inventory rather than forcing those historical counts.

20.4 Completion target and remote consoles[edit | edit source]

for i in {1..10}; do echo; done
systemctl status kitsnet-vm-startup-complete.target --no-pager

echo
cat /proc/sys/kernel/random/boot_id

echo
sudo cat /run/kitsnet-vm-startup/vm-startup-complete

echo
systemctl list-units 'socat-kvm@*.service' --no-pager

The completion-marker boot ID must match the current Wort boot ID.

21 Rollback[edit | edit source]

21.1 Before physical NVMe replacement[edit | edit source]

If any source-image creation or verification step fails:

  1. Do not remove the original NVMe.
  2. Do not proceed to hardware replacement.
  3. Correct the image/dock problem or abandon the maintenance.

To abandon after T0 has already been deactivated:

for i in {1..10}; do echo; done
sudo vgchange -ay T0

sudo sed -i \
  's|^# KN-T0-NVME-MIGRATION ||' \
  /etc/fstab

sudo systemctl daemon-reload
sudo findmnt --verify --verbose

sudo mount /var/lib/libvirt/qemu/save

sudo /usr/local/sbin/kitsnet-vm-startup-control next-boot auto
sudo /usr/local/sbin/kitsnet-vm-startup-control status

A controlled normal reboot may then be used to restore the accepted VM environment.

21.2 After physical replacement but before successful production validation[edit | edit source]

The original NVMe remains the primary rollback object.

If the replacement cannot pass a STOP checkpoint:

  1. Do not repair or reinitialize T0 speculatively.
  2. If Wort is operational, arm next-boot off.
  3. Power Wort off.
  4. Remove the replacement NVMe.
  5. Reinstall the untouched original NVMe.
  6. Boot protected.
  7. Validate T0 using its original device.
  8. Restore the save-mount fstab entry if it is still commented.
  9. Validate /var/lib/libvirt/qemu/save.
  10. Only then return to normal auto startup.

21.3 Duplicate-PV warning[edit | edit source]

The original NVMe and the successfully cloned replacement NVMe contain the same:

  • partition identities;
  • LVM PVID;
  • T0 VG UUID; and
  • LV UUIDs.

They must not later be attached to Wort simultaneously as normally scanned LVM devices.

Doing so would create a duplicate-PV/duplicate-VG condition.

22 Commands that are intentionally absent[edit | edit source]

A successful migration does not require any of the following:

pvcreate
vgcreate
vgextend
vgreduce
vgimport
vgimportclone
pvchange --uuid
vgchange --uuid
vgcfgrestore
pvck --repair
vgck --updatemetadata

Their unexpected apparent necessity is a reason to stop and diagnose the clone rather than continue.

23 Completion criteria[edit | edit source]

The T0 NVMe replacement is complete only when all of the following are true:

  • the replacement NVMe is installed;
  • the original NVMe remains preserved as rollback media;
  • the compressed source image remains available on the dock;
  • the image was byte-verified against the original source before removal;
  • the restored portion of the replacement was byte-verified against the image;
  • source and replacement logical sector sizes matched;
  • T0 retains the original PVID;
  • T0 retains the original VG UUID;
  • all T0 LVs retain their original LV UUIDs;
  • the partition start sector is unchanged;
  • the partition identity is unchanged;
  • the partition is expanded to the replacement capacity;
  • pvresize has exposed the added capacity as free space in T0;
  • pvck and vgck show no unexplained errors;
  • /var/lib/libvirt/qemu/save mounts normally from T0/kvm_save;
  • the managed-save inventory survived the migration;
  • a complete no-VM boot succeeded using the final storage configuration;
  • the subsequent automatic KitsNet startup completed successfully;
  • NAS ownership is coherent;
  • the Wort startup-completion target is active; and
  • expected VMs and console listeners are operational.

Only after a suitable observation period should the original NVMe or dock recovery image be considered available for reuse.