No edit summary |
m Psmode moved page KVM:host:Wort Operations:CentOS 8 to KVM:host:Wort Operations:prior-1 (CentOS 8) without leaving a redirect |
(No difference)
| |
Revision as of 13:01, 6 July 2024
Operational history of wort during its CentOS 8 incarnation
1 Specification[edit | edit source]
| Name | wort | |
|---|---|---|
| CPU/threads | 6/12 | |
| RAM | 128 GB | |
| MAC | eno1
Intel4p.f0 Intel4p.f1 enp10s0f2 enp10s0f3 Intel2p.f0 Intel2p.f1 wlp14s0 |
18:c0:4d:2f:39:78
80:61:5f:05:f3:46 80:61:5f:05:f3:47 00:1b:21:a7:66:f6 00:1b:21:a7:66:f7 80:61:5f:05:f3:46 80:61:5f:05:f3:47 cc:d9:ac:d3:b7:bf |
| Mountpoint | Size (GB) | Guest VG | Partition Size (GB) | Device | Device Size (GB) |
|---|---|---|---|---|---|
| /boot | 1.0 | 1.0 | uv033-disk0 | 14.0 | |
| / | 7.0 | V0uv033 | 13.0 | ||
| /tmp | 0.5 | ||||
| swap | 3.0 | ||||
| /var | 2.5 | V3uv033 | 4.0 | uv033-disk1 | 4.0 |
2 Build Day -2[edit | edit source]
Collect_lsi will screw up backups!
2.1 Info Collection[edit | edit source]
- Dump out all parted, LVM (LV, vg, pv), hdparm, df, HDSentinel, and smartctl info
- Capture RAID controller software config with
collect_lsi.sh -
yum list installedpackages and alsorpm -qa - Look for RPM based installs and save copies of kits. Find their project directories.
- Save
/repo/KitsNet/External_Packages/MR/*(see https://www.broadcom.com/support/download-search?pg=Legacy+Products&pf=Legacy+RAID+Controllers&pn=MegaRAID+SAS+9260-8i&pa=&po=&dk=&pl=)
- Save
- Preserve
yum.log -
crontab -land dump allcron.*directories - Dump all kvm.xml and save copies to Lafite
- Get list of running KVMs and document kvm ID against hostname
- Dump out network information
[psmode@wort ~]$ sudo ip addr [psmode@wort ~]$ sudo ip link [psmode@wort ~]$ sudo ovs-vsctl show [psmode@wort ~]$ sudo ovs-vsctl list port [psmode@wort ~]$ sudo lspci [psmode@wort ~]$ sudo nmcli device status [psmode@wort ~]$ sudo brctl show [psmode@wort ~]$ sudo ovs-dpctl show -s [psmode@wort ~]$ cat /proc/1/task/1/net/bonding/bond0 [psmode@wort ~]$ cat /proc/1/net/bonding/bond1 [psmode@wort ~]$ virsh net-list --all [psmode@wort ~]$ virsh net-info ovs-network [psmode@wort ~]$ virsh net-dumpxml ovs-network [psmode@wort ~]$ ifconfig [psmode@wort ~]$ route
- Copy network startup scripts
- Record grub config
- Save encryption key file
- Save
/root/*scripts - Save
/etc/passwdand/etc/group - Save
/etc/fstaband/etc/exports - Save
/etc/smartmontools/smartd.conf - Save
/etc/apcupsd/* - Save
/etc/logrotate.d/* - Save
/etc/rc.d/init.d/socat-kvm - Save
/usr/local/* - Capture httpd configuration
- Capture
.sshdirectory forpsmode - Reschedule Veeam backup of Cristal (KNADA AD DC) for 23:00
- Reschedule Veeam backup of Camus (KNADA AD DC) for 03:00
- Schedule backup of Backup LV to external (set 1) for 05:00 Day -1
- Schedule Wort and LV backup to external (set 1) for 05:00 Day -1
3 Build Day -1[edit | edit source]
3.1 Backups[edit | edit source]
- Create thumb drive images of Rocky Linux 8.5 install and Boot ISOs
- Trigger Veeam backup of Cristal (KNADA AD DC)
- Trigger Veeam backup of Camus (KNADA AD DC)
- Stop backup client jobs on all client systems and disable Backup share
- lafite
- tabasco
- sriracha
- kaliber
- stoli
- minime
- Shutdown NFS server on Jameson
- Create file addressable archive of all media, using Windows Backup share as a waypoint (????)
- Get extra copy of JES and VMS backups
- Move archives and library stuff to Lafite (???)
- Backup Backup LV to external (set 2)
- Zip
/psmodeand/rootdirectories from wort into archives and copy to Lafite - Get copy of this document and Linux:KitsNet Standard
4 Build Day 0[edit | edit source]
4.1 Data Protection[edit | edit source]
- Full shutdown VMScluster on DONQ and MORGAN
- Full shutdown all VMs
- Disable auto start for all VMs
- Wort and LV backup to external (set 2)
- Wait for backups to external to complete
- Check for wort dependencies on T1 and T3. For repo, dismount and remove from
/etc/fstabonce cron job disabled. For other mounts and swap, remove from/etc/fstab.
5 Build OS Upgrade[edit | edit source]
5.1 Pre-install[edit | edit source]
- vgexport T0, T1 and T3
- Shutdown wort
- Remove RAID controller
- Install new PV for T0h (using new 500 GB SanDisk 3D SATA III)
- backup BIOS and capture all screens
- install BIOS F13 https://download.gigabyte.com/FileList/BIOS/mb_bios_b550-aorus-pro-ac_f13.zip?v=f1ddca3f3143f81662da95ba1fc68715
- Boot Rocky Linux 8.5 Boot image from thumb drive into rescue mode
- Verify only disks present are thumb drive and new PV
- Test access to external disks - requires encryption key file
- Shutdown
5.2 Install[edit | edit source]
- See also https://www.tecmint.com/install-kvm-in-centos-8/
- Boot Rocky Linux 8.5 DVD1 image from thumb drive
- Install with:
- NIC teaming(balanced-xor with LAG on switch?)
- Complete Linux:KitsNet Standard installation
| Mountpoint | Size (GB) | VG | Partition SIze (GB) | Device | Device Size (GB) |
|---|---|---|---|---|---|
| /boot | 1.0 | 1.0 | /dev/sdd | 500.0 | |
| / | 50.0 | T0h | 499.0 | ||
| swap | 16.0 | ||||
| /var | 8.0 | ||||
| /kvm_save | 64.0 | ||||
| /repo | 453.0 | T3 | 10.0 | /dev/sdb | 10.0 |
- Exit to first boot
- Install known needed kits:
- VeraCrypt (https://www.veracrypt.fr/en/Downloads.html)
- apcupsd
- zip
- unzip
- hdparm (should be there)
- smartmontools
- iperf3
- sysstat (should be there)
- socat
- OpenRGB (https://gitlab.com/CalcProgrammer1/OpenRGB)
- lm_sensors and lm_sensors-sensord
- Virtualization (old)
- libvirt
- qemu-kvm
- KVM auto-hibernate?
- Virtualization (new)
-
dnf module install virt -
dnf install virt-install virt-viewer - and more from Step 2 of https://www.tecmint.com/install-kvm-in-centos-8/
-
- Test powerfail operation of apcupsd
5.3 Network Checkout[edit | edit source]
- Validate that boot was on bonded link. (Mode 6, regular ports on switch)
- Confirm wort IP address was assigned from DHCP reservation
- Test MTU to/from multiple peers.
- Test with iperf to/from multiple peers.
- Construct another bonded link. This will be used by kvm for most data interfaces (https://docs.cumulusnetworks.com/cumulus-linux/Layer-2/Bonding-Link-Aggregation/ and https://www.tecmint.com/configure-network-bonding-or-teaming-in-rhel-centos-7/) mode 0 LAG on switch
- Prepare another interface for use by fileserver for backup share only
5.4 Storage recovery[edit | edit source]
- Setup mount points for backup to external device
- Recover encryption key file
- Test access problem with password spec in script - wants to prompt
- Shutdown
- Reconnect RAID controller
- Reconnect NVME
- Boot
- pvscan
- vgimport T0
- Validate vg and LVs
- vgimport T1
- Validate vg and LVs
- vgimport T3
- Validate vg and LVs
- MegaRaid install. Revisit storcli version
- Setup smartd (reuse configuration file from previous system)
- Verify state of controller and volumes.
- Create T3 mountpounts
- Update
/etc/fstaband test mounts - Reboot
5.5 KVM recovery[edit | edit source]
- Validate all tools required for KVM are in place
- Recreate
/KVMdirectory structure - Recreate socat/kvm startup job
- Start VMs. For each:
- Inspect xml to find network device and change this to the second bonded interface
- Verify LVs are available
- See if xml includes VNC console. Prepare to connect if it does
- Check netstat for collisions on vnc or socat console
- Boot and rapidly connect to serial console to monitor boot
- Attempt ssh in to VM
- Attempt console connection to VM
- Jameson: add second interface to VM definition for use by backup client connections
- Quadsec: Verify auto-hibernate (for this one and the rest)
- Camus: validate visibility to/from domain
- Jameson: validate availability of shares. Then validate mapping of personal network drive to K: at login
- Reestablish VM autostart
| Name | Title | |
|---|---|---|
| uv029 | hendrick (Samba AD build factory) | |
| uv025 | chandon (ConsoleWorks) | |
| uv026 | carling (Zabbix) | |
| uv038 | zoco (new Zabbix monitoring server) | |
| uv031 | camus (Samba AD DC) | |
| uv033 | cristal (Samba AD DC) | |
| uv032 | chivas (Wiki server) | |
| uv035 | quadsec (KitsNet bastion host) | |
| uv017 | screech / mailroom (PostfixAdmin/DoveCot email server) | |
| uv023 | prosecco / mailgw (eFa Email Filter Appliance) | |
| uv019 | jameson / nas (Samba NAS) | |
| uv034 | campari (new Samba NAS) | |
| uv020 | daniel / media (Plex media server) | |
| uv037 | ripple (simh VAX simulator) | |
| wv902 | crownroyal (FreeAXP simulator) |
5.6 Post-install[edit | edit source]
- Reactivate cron jobs:
- repo updates (two jobs)
- sensors data collection
- monthly backups
- enable TRIM service
- Restore schedule of Veeam backup of Cristal (KNADA AD DC) to 05:30
- Restore schedule of Veeam backup of Camus (KNADA AD DC) to 16:30
6 Updates[edit | edit source]
6.1 12/31/2021[edit | edit source]
6.1.1 Start work to enable /tmp on a tmpfs filesystem[edit | edit source]
I started the work to enable /tmp on a tmpfs filesystem and make it ephemeral. See https://unix.stackexchange.com/questions/352042/systemd-backed-tmpfs-how-to-specify-tmp-size-manually for details. As root, executed systemctl edit tmp.mount and entered:
[Mount]
Options=mode=1777,strictatime,nosuid,nodev,size=536870912
This created the supporting directory structure and the file /etc/systemd/system/tmp.mount.d/override.conf with this content. At this point, there is no change to /tmp which already contains files. What happens during reboot will be interesting.
6.1.2 Allowed NTP through firewall[edit | edit source]
[psmode@wort ~]$ sudo firewall-cmd --permanent --add-service=ntp
success
[psmode@wort ~]$ sudo firewall-cmd --reload
success
[psmode@wort ~]$ sudo firewall-cmd --list-all
public (active)
target: default
icmp-block-inversion: no
interfaces: bond0 eno1 enp4s0f0 enp5s0f1 enp6s0f0 enp6s0f1 kvm-bkup kvm-guests nm-bond
sources:
services: cockpit dhcpv6-client http https ntp ssh
ports: 3071/tcp 7902-7939/tcp
protocols:
forward: no
masquerade: no
forward-ports:
source-ports:
icmp-blocks:
rich rules:
rule family="ipv4" source address="192.168.15.64/32" port port="10050" protocol="tcp" accept
rule family="ipv4" source address="192.168.15.54/32" port port="10050" protocol="tcp" accept
6.2 01/02/2022[edit | edit source]
6.2.1 Fixed glances[edit | edit source]
Fixed glances with code mod as per https://mangolassi.it/topic/19117/how-can-i-show-disk-io-in-glances/7 while the citation calls for sed, I did it manually in vi:
# sed -i 's#fields_len == 14:#fields_len == 14 or fields_len == 18:#g' /usr/lib64/python3.6/site-packages/psutil/_pslinux.py
I also disabled ipv6 on bond0. It appears ipv6 on this interface is hopeless for now.
6.3 1/7/2002[edit | edit source]
6.3.1 Get sensors to work for Zabbix[edit | edit source]
- Installed acpid
- Zabbix is picking up errors getting temperatures (wort metric Information about the lm-sensors). Problem is with a temperature reading for the
Adapter: Virtual device. (see https://github.com/nojhan/liquidprompt/issues/445) Solution to suppressing the error for the thing we do not care about is to suppress taking the reading. To do this, create a text file named/etc/sensors.d/iwlwifi_1-virtual-0and populate it with:chip "iwlwifi-virtual-*" ignore temp1 - Still not working. Edited
/etc/zabbix/zabbix_agent2.d/userparameter_sensors.confto call for read of specific chipsUserParameter=sensors,sensors -j -A nouveau-pci-0400 it8792-isa-0a60 k10temp-pci-00c3
- After restart of the agent, the data going back looks good, but there is not yet creation of any new items or graphs.
6.3.2 Installing support for Zabbix virbix[edit | edit source]
- See (https://share.zabbix.com/virtualization/kvm/virbix and https://github.com/sergiotocalini/virbix)
- Installed ksh (required for deployment)
- did the git clone and updated
ZABBIX_INC="${ZABBIX_DIR}/zabbix_agentd.dreference toZABBIX_INC="${ZABBIX_DIR}/zabbix_agentd2.d" - Deployed
[root@wort src]# sudo ./virbix/deploy_zabbix.sh -u "qemu:///system" './virbix/virbix/virbix.sh' -> '/etc/zabbix/scripts/agentd/virbix/virbix.sh' './virbix/virbix/scripts' -> '/etc/zabbix/scripts/agentd/virbix/scripts' './virbix/virbix/scripts/domain_check.sh' -> '/etc/zabbix/scripts/agentd/virbix/scripts/domain_check.sh' './virbix/virbix/scripts/domain_list.sh' -> '/etc/zabbix/scripts/agentd/virbix/scripts/domain_list.sh' './virbix/virbix/scripts/functions.sh' -> '/etc/zabbix/scripts/agentd/virbix/scripts/functions.sh' './virbix/virbix/scripts/net_check.sh' -> '/etc/zabbix/scripts/agentd/virbix/scripts/net_check.sh' './virbix/virbix/scripts/net_list.sh' -> '/etc/zabbix/scripts/agentd/virbix/scripts/net_list.sh' './virbix/virbix/scripts/pool_check.sh' -> '/etc/zabbix/scripts/agentd/virbix/scripts/pool_check.sh' './virbix/virbix/scripts/pool_list.sh' -> '/etc/zabbix/scripts/agentd/virbix/scripts/pool_list.sh' './virbix/virbix/scripts/report.sh' -> '/etc/zabbix/scripts/agentd/virbix/scripts/report.sh' './virbix/virbix/scripts/report_domains.sh' -> '/etc/zabbix/scripts/agentd/virbix/scripts/report_domains.sh' './virbix/virbix/scripts/report_nets.sh' -> '/etc/zabbix/scripts/agentd/virbix/scripts/report_nets.sh' './virbix/virbix/scripts/report_node.sh' -> '/etc/zabbix/scripts/agentd/virbix/scripts/report_node.sh' './virbix/virbix/scripts/report_pools.sh' -> '/etc/zabbix/scripts/agentd/virbix/scripts/report_pools.sh' './virbix/virbix/virbix.conf.example' -> '/etc/zabbix/scripts/agentd/virbix/virbix.conf.new' './virbix/virbix/zabbix_agentd.conf' -> '/etc/zabbix/zabbix_agentd2.d/virbix.conf.new'
- Test
[root@wort src]# /etc/zabbix/scripts/agentd/virbix/virbix.sh -s domain_list -j DOMID:DOMNAME:DOMUUID:DOMTYPE:DOMSTATE { "data":[ { "{#DOMID}":"63", "{#DOMNAME}":"uv038", "{#DOMUUID}":"02160641-ed2a-4b3a-a3d2-8db8c3d55551", "{#DOMTYPE}":"hvm", "{#DOMSTATE}":"running" }, { "{#DOMID}":"71", "{#DOMNAME}":"wv902", "{#DOMUUID}":"0d0ec200-b0da-3fc7-df68-eb17840f0291", "{#DOMTYPE}":"hvm", "{#DOMSTATE}":"running" }, { "{#DOMID}":"70", "{#DOMNAME}":"uv037", "{#DOMUUID}":"15b657d2-1e7f-44c1-a9e3-09110ff7320a", "{#DOMTYPE}":"hvm", "{#DOMSTATE}":"running" }, { "{#DOMID}":"-", "{#DOMNAME}":"uv019", "{#DOMUUID}":"1899841b-2cbd-a480-d5d8-aa2d9f281ec1", "{#DOMTYPE}":"hvm", "{#DOMSTATE}":"shut off" }, { "{#DOMID}":"67", "{#DOMNAME}":"uv034", "{#DOMUUID}":"2580bf9c-7d91-41f7-b204-b58b1acee4b2", "{#DOMTYPE}":"hvm", "{#DOMSTATE}":"running" }, { "{#DOMID}":"66", "{#DOMNAME}":"uv017", "{#DOMUUID}":"266b61f2-bfab-1813-6c00-7dc5c780f928", "{#DOMTYPE}":"hvm", "{#DOMSTATE}":"running" }, { "{#DOMID}":"60", "{#DOMNAME}":"uv040", "{#DOMUUID}":"2d6c7acf-2eb8-4835-87e3-89b0d9726286", "{#DOMTYPE}":"hvm", "{#DOMSTATE}":"running" }, { "{#DOMID}":"68", "{#DOMNAME}":"uv039", "{#DOMUUID}":"376d4633-2e70-4554-9ab2-32bb26f39a71", "{#DOMTYPE}":"hvm", "{#DOMSTATE}":"running" }, { "{#DOMID}":"72", "{#DOMNAME}":"uv033", "{#DOMUUID}":"4acb2cb0-5489-40dc-a662-4a08f73c45c3", "{#DOMTYPE}":"hvm", "{#DOMSTATE}":"running" }, { "{#DOMID}":"62", "{#DOMNAME}":"uv032", "{#DOMUUID}":"61e9a7b1-4ba8-4292-b1fe-8b9d897c0b44", "{#DOMTYPE}":"hvm", "{#DOMSTATE}":"running" }, { "{#DOMID}":"64", "{#DOMNAME}":"uv029", "{#DOMUUID}":"6aeb3ad2-1a5b-452a-bd38-b982b21256fa", "{#DOMTYPE}":"hvm", "{#DOMSTATE}":"running" }, { "{#DOMID}":"73", "{#DOMNAME}":"uv031", "{#DOMUUID}":"6e8cdc83-3284-4866-97b6-deccf2c14f65", "{#DOMTYPE}":"hvm", "{#DOMSTATE}":"running" }, { "{#DOMID}":"65", "{#DOMNAME}":"uv023", "{#DOMUUID}":"6faae374-e311-f965-4f57-e6f2ecc22a3f", "{#DOMTYPE}":"hvm", "{#DOMSTATE}":"running" }, { "{#DOMID}":"-", "{#DOMNAME}":"uv020", "{#DOMUUID}":"9838fbec-6c80-8c51-809d-3f3079b96124", "{#DOMTYPE}":"hvm", "{#DOMSTATE}":"shut off" }, { "{#DOMID}":"-", "{#DOMNAME}":"uv026", "{#DOMUUID}":"9bf46d15-22d2-eb2c-4d25-7cf6966534ba", "{#DOMTYPE}":"hvm", "{#DOMSTATE}":"shut off" }, { "{#DOMID}":"69", "{#DOMNAME}":"uv025", "{#DOMUUID}":"b7b4795b-9334-ac91-ed77-b3ffcfd13594", "{#DOMTYPE}":"hvm", "{#DOMSTATE}":"running" }, { "{#DOMID}":"61", "{#DOMNAME}":"uv035", "{#DOMUUID}":"fb62047b-d9a7-4203-af3a-43fb993c21b3", "{#DOMTYPE}":"hvm", "{#DOMSTATE}":"running" } ] }
- Found mistaken directory reference Z
ABBIX_INC="${ZABBIX_DIR}/zabbix_agentd2.d"instead ofZABBIX_INC="${ZABBIX_DIR}/zabbix_agent2.d". I moved the config file, removed the directory and then went back to fix the deployment script.
6.3.3 Installing support for zabbix-kvm-res[edit | edit source]
- see https://share.zabbix.com/virtualization/kvm/kvm-monitoring and https://github.com/bushvin/zabbix-kvm-res
6.3.4 Fixing sudo[edit | edit source]
- Fixed
/etc/sudoers.d/zabbixfile withzabbix ALL=(ALL) NOPASSWD: /sbin/smartctl, /usr/bin/virsh
6.4 1/10/2022[edit | edit source]
6.4.1 Unable to find AMD Northbridge id[edit | edit source]
- the system message log is filling up with these messages and causing the logwatch report to grow so large that i cannot be emailed.
- This appears to be an issue in the kernel with the k10temp code. As per https://gitmemory.cn/repo/groeck/k10temp/activity, i will disable the module with
modprobe -r k10tempto make the errors stop. When the module is fixed, it will need to be reenabled.
6.5 9/11/2022[edit | edit source]
Trying to setup new 8TB Veracrypt drives turned into near disaster.
6.5.1 New 8TB kitsnet-pbd Drives[edit | edit source]
I inserted both new drives in the USB dock (Western Digital Gold S/N VR0DR3BK and VYG3DZ3R). Partitioned each with parted using a gpt disk label
[root@wort ~]# parted /dev/sde print
Model: ASMT USB 3.0 Destop H (scsi)
Disk /dev/sde: 8002GB
Sector size (logical/physical): 512B/4096B
Partition Table: gpt
Disk Flags:
Number Start End Size File system Name Flags
1 1049kB 8002GB 8002GB
With the devices partitioned, I started two screen sessions and a script logging session within each for the execution of the truecrypt encryption activity, which was going to take over seven hours.
[root@wort ~]# veracrypt -t -c --volume-type=normal /dev/sde1 --encryption=AES-Twofish-Serpent --hash=sha-512 --filesystem=none --non-interactive --pim=0 --verbose -k=truecrypt.keyfile -p=""
Both drives managed to run up to 100% usage according to diskmon.
In the morning, I was not able to log into wort. It would respond to ping, but bringing up the login prompt was very, very slow. Due to an error in selecting the correct console video output, I was not able to attempt local console login. In the end, I tried resets, but did not realize that the reboot was underway and later did a power cycle to force the system down. From here, things got worse
6.5.2 Restart Failure[edit | edit source]
Boot would fail quite early on. The system would show a grub boot menu with valid choices, but attempting to boot any of the known OSes would fail with a kernel panic and an error message resembling end Kernel panic - not syncing: VFS: Unable to mount root fs on unknown-block(0,0) . None of the available boot options would get anywhere. I theorized that spontaneous changes to the TPM settings in BIOS could cause boot failures. This proved to not be an issue as it appeared that our starting position was with TPM features disabled.
I did locate a Rocky Linux v8.6 installation thumb drive and booted into recovery mode there. However, rescue mode still said that it could not find any Linux partitions. However, upon going down to the shell, all partitions on all permanent drives (even those in the RAID controller) were found and appeared valid. LVM had started and successfully mapped all the LVs. Running xfs_repair against the LVs on T0h was generally successful, though a couple of the LVs had to be mounted to flush pending operations before the xfs_repair was successful. This still did not address the main problem, but possibly would have been needed anyway.
The problem was very odd. Grub was complaining that it could not find Linux partitions, but they were clearly there and appeared fine. Becoming concerned about the grub configuration, I looked for grub.cfg and found two of them. Doing a diff on the two returned:
# diff /boot/efi/EFI/rocky/grub.cfg /boot/grub2/grub.cfg
140c140
< set kernelopts="root=/dev/mapper/T0h-root ro crashkernel=auto resume=/dev/mapper/T0h-swap rd.lvm.lv=T0h/root rd.lvm.lv=T0h/swap rhgb quiet "
---
> set kernelopts="root=/dev/mapper/T0h-root ro crashkernel=auto resume=/dev/mapper/T0h-swap rd.lvm.lv=T0h/root rd.lvm.lv=T0h/swap rhgb usbcore.autosuspend=-1 acpi_enforce_resources=lax mem_encrypt=on kvm_amd.sev=1 "
I looked up the meaning of the kernel options that were unique to /boot/grub2/grub.cfg (the one actually used for booting). Most all of the options were really required, but mem_encrypt=on did not seem to absolutely required. This setting is part of Secure Memory Encryption (SME) . SME and Secure Encrypted Virtualization (SEV) are features found on AMD processors. Doing some googling on this turned up more than a couple articles complaining about stability issues with this. In particular, Enable Secure Memory Encryption (SME) - kernel parameter mem_encrypt - by default? argued that the feature was very buggy and that mem_encrypt should be set to off lest there be booting problems.
Next boot, I edited the kernel options conversationally to omit the mem_encrypt=on clause. With this change, booting was successful.
After this, I updated the GRUB_CMDLINE_LINUX line in /etc/default/grub to change the mem_encrypt=on clause to mem_encrypt=off Next was to build a new grub configuration with grub2-mkconfig -o /boot/grub2/grub.cfg This new boot configuration will need to be tested after the completion of the encryption operation against the first drive.
6.5.3 New 8TB kitsnet-pbd Drives - Take 2[edit | edit source]
With the system back in operation, time to restart the drive encryption. For this attempt, only one drive will be done at a time. Restarting the session for /dev/sde1 was successful. Perhaps unsurprisingly, throughput may be faster now. Prior attempt saw 195 MB/s sustained writes, while this attempt is running at around 230 MB/s. Zabbix data confirms that IO operations are up from 395 writes/s to 495 writes/s. As of 15:15, truecrypt estimated six hours remaining.
6.6 2/6/2023[edit | edit source]
6.6.1 Improve monthly backup efficiency[edit | edit source]
The monthly backup of guest virtual disks has never been fully efficient when it comes to large volumes. The code is setup to evaluate the size of the LV implementing the virtual disk and will run it through a gzip process for really large disks. The idea is that compressing out these disk images, especially for largely empty volumes or volumes with highly compressible data, will cut down on backup time overall since a much smaller number of blocks need be encrypted and written to the single output volume. The problem is that throughput of these operations is constrained by the single threaded throughput of gzip. It turns out there is a multi-threaded implementation available, pigz. While the option to build from source is available, I will start off by using the packaged version available from the Anaconda repo that wort already references. While this may be a few versions behind, automated maintenance would be a very good thing. So good in fact, that the utility is already installed!
The default in version 2.4 of pigz is to use eight processors; later versions increased the default to all online processors. This may be overridden with the --processes argument.
The reference to gzip is hardcoded in /usr/local/sbin/make_guest_bck_script.awk, which is distributed as part of the KNkvmHostBackup package on KitsNet. This awk script is called in /usr/local/sbin/KVMbackup to dynamically generate the script which will then be used to snap and backup all the LVs implementing virtual disks for all the guests on the host. I am prototyping an update to the code so that the awk script will conditionally replace the default gzip command with the command passed to it via a variable named ZipCmd. The calling shell script, KVMbackup, will look for the pigz command on the path and will define the ZipCmd variable to the awk script with the multithreaded pigz command. By default, the variable will be am empty string, which will cause the awk script to use the single-threaded gzip command by default.
In the end, an uncontested test run completed in 3.5 hours instead of the normal 6 hours. Systemwide idle times would range from 70% during execution of the the regular dd command to as low as 20% when pigz was running up to over 800% CPU utilization. Over the execution of the entire job, CPU Nice time (as recorded by Zabbix) ranged from 15% to 56%.
With the prototype in place, I adjusted the start time of the monthly KVMbackup from from 5:45 to 4:45. I will evaluate system behavior on the next monthly run to see if reduction in the overall elpased time of the monthly backup can be achieved by moving this start time up further.
6.7 6/11/2023[edit | edit source]
6.7.1 Planned downtime for 128 GB RAM upgrade and top up of radiator.[edit | edit source]
In preparation for RAM expansion, I expanded the T0h/kvm_save volume from 64 GB to 128 GB.
In order to eliminate co-incidental issues complicating the restart of wort after the memory upgrade, i performed a simple reboot of wort from the physical console. I discovered the wireless keyboard batteries were dead and replaced these. Switching the left monitor to input HDMI1 allowed me to see the console display. For additional boot prep, I ensured that the USB backup drives were dismounted and the dock powered down. I also shutdown the es40 emulator within alphakronik, since it was spinning the CPUs assigned to it. I also noticed that the clock on the instance was way off showing 5/31 (probably not worth running this). All other instances were left running, including nested simulators.
Wort came back up relatively smoothly. However, DONQ appears to have crashed. It appeared that the remote console was hosed as well, so I restarted it within ConsoleWorks, though this may not have actually been necessary. The initial phases of simulator startup were very, very slow, but gained speed once the simulator instance advanced to the VMS startup sequence. All other instances (including porter running x86 VMS) came back nicely, excepting for time. Many of the guests had their time out of sync by as much as five minutes.
I did discover that wort was configured to use its own time as a source in /etc/chronyd.conf which lead it to mistakenly promote itself to stratum 0. For now, I dropped the time sevrer (CNAME for wort) as a source and replaced it with courvoisier.
There are settings defined in /etc/sysconfig/libvirt-guests that can improve the performance of shutdown/startup of libvirt and guests on the system as described in the RHEL Virtualization Administration Guide section 14.9.3. Manipulating the libvirt-guests Configuration Settings . I updated the PARALLEL_SHUTDOWN parameter from the default of zero (no parallelism) to 6. I also updated SYNC_TIME to 1 so that libvirt will try to sync guest time on domain resume. To test the efficacy of these settings, a second and third reboot test was carried out, with DONQ crashing and requiring console restart each time.
To prepare for the "real" shutdown, I used /usr/local/sbin/able-vm-startup.sh disable to make the host safe for multiple reboots.
During the downtime, I cleaned the dust from the case and fans, topped up the radiator with distilled water and ran the system with the motherboard power bypassed so the pump would run. With the coolant circulating, i was able to verify the reservoir full after tilting it back and forth and restarting it a few times. I did have a probllem with the SATA power connector to the AOC assembly getting disconnected while i was working on it, resulting in no power to the pump.
Reboot was a hassle because the memory settings in BIOS has the XMP profile of the 64 GB kit still active, so the system would not event get to the POST screen. I removed all the new member, put in one stick of the old and got into BIOS. I then disabled XMP, saved and exited. I was than able to install the new RAM and get into BIOS to re-enable XMP with the profile of the new RAM. I had to manually set the RAM voltage to 1.35V. With that, the system booted clean and I was able to re-enable libvirtd and reboot to bring the guests back up. They loaded just fine from the suspend files.
6.8 8/22/2023[edit | edit source]
6.8.1 Fix MegaRAID[edit | edit source]
It turns out that MegaRAID was never properly installed, so there was no local CLI support nor any remote GUI support to check on the status of the RAID controller and the devices connected to it. I downloaded updated versions of the MSM software and updated the report on wort. For more information, see KitsNet Operations:External Packages:Linux:MegaRAID Storage Manager. All packages associated with MSM (MegaRAID_Storage_Manager, sas_snmp and sas_ir_snmp) appeared to have most of their files missing despite showing as installed in rpm and dnf. I ended up removing all these three products in order to facilitate a clean install of the updated versions. With this version of MSM, it was also necessary to install a Java environment. AS pre the recommendation of the MSM readme file, I installed the older java-1.8.0-openjdk package via dnf. once all this was in place, I then executed the install.csh script for a complete install.
With the install completed, both the CLI and GUI versions of MS wer operational. However, the lcoation of the storcli64 image is now in the path at /usr/local/sbin/storcli64