KVM:CPU Pinning Guidance: Difference between revisions

Peter A. Smode (talk | contribs)
Created page with "This guide defines when CPU pinning is appropriate for KitsNet KVM/libvirt virtual machines, how to evaluate the need, how to implement it safely, and how to measure whether it helped. It is intended for internal operations on the KitsNet virtualization and Docker Swarm environment. The guidance distinguishes observed facts from recommendations. CPU pinning is a scheduling-control technique; it is not a substitute for fixing storage, network, guest, or hypervisor faults..."
 
Peter A. Smode (talk | contribs)
No edit summary
 
(One intermediate revision by the same user not shown)
Line 49: Line 49:
* pinning would make capacity planning or failover less flexible;
* pinning would make capacity planning or failover less flexible;
* the measurement plan cannot distinguish improvement from normal run-to-run variation.
* the measurement plan cannot distinguish improvement from normal run-to-run variation.
=== Decision threshold ===
Use the libvirt vCPU delay rate as the primary numeric screening metric:
<math>delay\_rate = \frac{\Delta vcpu.delay}{\Delta vcpu.time}</math>
For a general-purpose candidate, treat pinning as warranted when the '''per-vCPU delay rate is at least 5% in three consecutive 60-second samples''' under a representative workload, while the guest is not itself CPU-saturated. A rate of '''10% or more''' is a high-priority scheduling problem and should be investigated immediately.
For a latency-sensitive HA candidate, use the stricter event rule: '''one Keepalived timer-expiration event delayed by 2 seconds or more, or two shorter timer-expiration events in a 15-minute window, is sufficient to investigate and normally justify a pinning trial''' when the event correlates with elevated vCPU delay. A timer event alone does not prove CPU scheduling is the cause; check network, storage, and guest evidence as well.
These are KitsNet operating thresholds, not universal KVM limits. They provide a repeatable change decision and can be revised when enough local history exists. The 2026-09-05 <code>zaya</code> sample (approximately 1.5-3.2% delay rate during a Veeam SQL-dump workload, with host idle 60-80%, I/O wait 0%, and steal 0%) remained below the 5% general-purpose threshold and was left unpinned. Fox and anchor had timer delays and much higher cumulative scheduling delay, so a pinning trial was justified.


== Evaluation workflow ==
== Evaluation workflow ==
Line 58: Line 70:
Record the guest state, workload, time window, vCPU count, memory, and any relevant event timestamps. Do not compare a busy backup window with an idle window.
Record the guest state, workload, time window, vCPU count, memory, and any relevant event timestamps. Do not compare a busy backup window with an idle window.


The following is a read-only example used during the 2026-09-05 KitsNet NAS investigation. It connects as <code>psmode</code> and does not use root SSH.
Begin on the KVM host and collect one or more candidates interactively. This avoids baking host or VM names into later code blocks. The label is a human-readable tracking name; the OS hostname is used for guest SSH; and the libvirt VM/domain name is used for <code>virsh</code>.
 
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
collect_kitsnet_candidates() {
    local target_env="/tmp/kitsnet-kvm-candidates.${UID}.env"
    local candidate_count index label host vm
 
    read -r -p "Number of VM candidates (1-32): " candidate_count
    case "$candidate_count" in
        ''|*[!0-9]*) echo "Invalid candidate count" >&2; return 1 ;;
    esac
    if [ "$candidate_count" -lt 1 ] || [ "$candidate_count" -gt 32 ]; then
        echo "Candidate count must be between 1 and 32" >&2
        return 1
    fi
 
    umask 077
    {
        printf 'export KITSNET_HYPERVISOR_HOST=%q\n' "$(hostname -f)"
        printf 'export KITSNET_CANDIDATE_COUNT=%q\n' "$candidate_count"
    } > "$target_env"
 
    for index in $(seq 1 "$candidate_count"); do
        echo "===== CANDIDATE $index ====="
        read -r -p "Candidate label: " label
        read -r -p "Candidate OS hostname/FQDN: " host
        read -r -p "Candidate libvirt VM/domain name: " vm
 
        case "$label" in ''|*[!A-Za-z0-9_.-]*) echo "Invalid label" >&2; return 1 ;; esac
        case "$host" in ''|*[!A-Za-z0-9.-]*) echo "Invalid hostname" >&2; return 1 ;; esac
        case "$vm" in ''|*[!A-Za-z0-9_.-]*) echo "Invalid VM name" >&2; return 1 ;; esac
 
        if ! sudo -n virsh dominfo "$vm" >/dev/null 2>&1; then
            echo "VM/domain not found: $vm" >&2
            return 1
        fi
 
        printf 'export KITSNET_CANDIDATE_%s_LABEL=%q\n' "$index" "$label" >> "$target_env"
        printf 'export KITSNET_CANDIDATE_%s_HOST=%q\n' "$index" "$host" >> "$target_env"
        printf 'export KITSNET_CANDIDATE_%s_VM=%q\n' "$index" "$vm" >> "$target_env"
    done
 
    chmod 600 "$target_env"
    echo "target_environment=$target_env"
    cat "$target_env"
}
 
collect_kitsnet_candidates
collection_rc=$?
echo "candidate_collection=$collection_rc"
echo "The interactive shell remains open."
</syntaxhighlight>
 
The collection block is intentionally run directly in the interactive shell rather than inside a heredoc-driven child shell; this keeps prompts attached to the terminal. Later blocks source the resulting environment file. It connects as <code>psmode</code> when guest SSH is used and does not use root SSH.
 
The following baseline block is noninteractive and evaluates every collected candidate:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
Line 64: Line 132:
if bash <<'BASH'
if bash <<'BASH'
set -Eeuo pipefail
set -Eeuo pipefail
test "$(id -un)" = psmode
target_env="/tmp/kitsnet-kvm-candidates.${UID}.env"
for peer in fox anchor; do
test -r "$target_env"
     echo "===== $peer GUEST ====="
. "$target_env"
     ssh -n -o BatchMode=yes "psmode@$peer.lan.kitsnet.us" \
for index in $(seq 1 "$KITSNET_CANDIDATE_COUNT"); do
         "hostname -f; awk '/^cpu([0-9 ]|$)/ {print}' /proc/stat; uptime"
    label_var="KITSNET_CANDIDATE_${index}_LABEL"
    host_var="KITSNET_CANDIDATE_${index}_HOST"
    label="${!label_var}"
    host="${!host_var}"
     echo "===== ${label} GUEST ====="
     if ssh -n -o BatchMode=yes "psmode@${host}" \
         "hostname -f; awk '/^cpu([0-9 ]|$)/ {print}' /proc/stat; uptime"; then
        echo "guest_baseline=PASS"
    else
        echo "guest_baseline=FAIL"
    fi
done
done
BASH
BASH
Line 115: Line 193:
if bash <<'BASH'
if bash <<'BASH'
set -Eeuo pipefail
set -Eeuo pipefail
test "$(id -un)" = psmode
target_env="/tmp/kitsnet-kvm-candidates.${UID}.env"
ssh -o BatchMode=yes psmode@wort.lan.kitsnet.us 'bash -s' <<'REMOTE'
test -r "$target_env"
set -Eeuo pipefail
. "$target_env"
for domain in uv059 uv060; do
for index in $(seq 1 "$KITSNET_CANDIDATE_COUNT"); do
     echo "===== $domain ====="
    label_var="KITSNET_CANDIDATE_${index}_LABEL"
     sudo -n virsh dominfo "$domain" | grep -E '^(Name|UUID|CPU\(s\)|CPU time|State):' || true
    vm_var="KITSNET_CANDIDATE_${index}_VM"
     sudo -n virsh vcpupin "$domain"
    label="${!label_var}"
     sudo -n virsh emulatorpin "$domain"
    vm="${!vm_var}"
     sudo -n virsh schedinfo "$domain" || true
     echo "===== ${label} (${vm}) ====="
     sudo -n virsh domstats "$domain" --vcpu | \
     sudo -n virsh dominfo "$vm" | grep -E '^(Name|UUID|CPU\(s\)|CPU time|State):' || true
     sudo -n virsh vcpupin "$vm"
     sudo -n virsh emulatorpin "$vm"
     sudo -n virsh schedinfo "$vm" || true
     sudo -n virsh domstats "$vm" --vcpu | \
         grep -E '^(Domain|  vcpu\.[0-9]+\.(time|delay))' || true
         grep -E '^(Domain|  vcpu\.[0-9]+\.(time|delay))' || true
done
done
REMOTE
BASH
BASH
then
then
Line 146: Line 227:
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
if bash <<'BASH'
if bash <<'BASH'
set -Eeuo pipefail
test "$(id -un)" = psmode
ssh -o BatchMode=yes psmode@wort.lan.kitsnet.us 'bash -s' <<'REMOTE'
set -Eeuo pipefail
set -Eeuo pipefail
sudo -n virsh list --all
sudo -n virsh list --all
Line 155: Line 233:
         grep -E '^(Name|CPU\(s\)|CPU time|State):' || true
         grep -E '^(Name|CPU\(s\)|CPU time|State):' || true
done
done
REMOTE
BASH
BASH
then
then
Line 191: Line 268:
<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
if bash <<'BASH'
target_env="/tmp/kitsnet-kvm-candidates.${UID}.env"
set -Eeuo pipefail
test -r "$target_env"
test "$(id -un)" = psmode
. "$target_env"
ssh -o BatchMode=yes psmode@wort.lan.kitsnet.us 'bash -s' <<'REMOTE'
apply_candidate_pinning() {
set -Eeuo pipefail
    local index label_var vm_var label vm vcpu_count emulator_cpuset pin_line vcpu cpu
for domain in uv059 uv060; do
    local -a pin_sets
    sudo -n virsh dominfo "$domain" >/dev/null
    for index in $(seq 1 "$KITSNET_CANDIDATE_COUNT"); do
done
        label_var="KITSNET_CANDIDATE_${index}_LABEL"
 
        vm_var="KITSNET_CANDIDATE_${index}_VM"
sudo -n virsh vcpupin uv059 0 0 --live --config
        label="${!label_var}"
sudo -n virsh vcpupin uv059 1 1 --live --config
        vm="${!vm_var}"
sudo -n virsh emulatorpin uv059 8-9 --live --config
        vcpu_count="$(sudo -n virsh dominfo "$vm" | awk '$1 == "CPU(s):" {print $2}')"
 
        test "$vcpu_count" -ge 1
sudo -n virsh vcpupin uv060 0 2 --live --config
        echo "===== ${label} (${vm}) ====="
sudo -n virsh vcpupin uv060 1 3 --live --config
        read -r -p "Host CPU set for each of ${vcpu_count} vCPUs (space-separated): " pin_line
sudo -n virsh emulatorpin uv060 10-11 --live --config
        read -r -a pin_sets <<< "$pin_line"
 
        test "${#pin_sets[@]}" -eq "$vcpu_count"
for domain in uv059 uv060; do
        for vcpu in $(seq 0 $((vcpu_count - 1))); do
    echo "===== $domain ====="
            cpu="${pin_sets[$vcpu]}"
    sudo -n virsh vcpupin "$domain"
            [[ "$cpu" =~ ^[0-9,-]+$ ]]
    sudo -n virsh emulatorpin "$domain"
            sudo -n virsh vcpupin "$vm" "$vcpu" "$cpu" --live --config
done
        done
REMOTE
        read -r -p "Emulator-thread CPU set: " emulator_cpuset
BASH
        [[ "$emulator_cpuset" =~ ^[0-9,-]+$ ]]
then
        sudo -n virsh emulatorpin "$vm" "$emulator_cpuset" --live --config
    echo "cpu_placement_return_code=0"
        sudo -n virsh vcpupin "$vm"
else
        sudo -n virsh emulatorpin "$vm"
    echo "cpu_placement_return_code=$?"
    done
fi
}
apply_candidate_pinning
echo "The interactive shell remains open."
echo "The interactive shell remains open."
</syntaxhighlight>
</syntaxhighlight>
Line 236: Line 314:
=== Rollback ===
=== Rollback ===


Rollback is done by restoring an unrestricted CPU set for the affected domain. Use the exact prior placement recorded in the change log when possible. The following example removes the tested pinning from the two NAS guests while preserving the domain definitions:
Rollback is done by restoring an unrestricted CPU set for the affected domain. Use the exact prior placement recorded in the change log when possible. The following interactive block operates on every collected candidate only after an explicit confirmation:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
if bash <<'BASH'
target_env="/tmp/kitsnet-kvm-candidates.${UID}.env"
set -Eeuo pipefail
test -r "$target_env"
test "$(id -un)" = psmode
. "$target_env"
ssh -o BatchMode=yes psmode@wort.lan.kitsnet.us 'bash -s' <<'REMOTE'
read -r -p "Type ROLLBACK to clear pinning for all collected candidates: " confirmation
set -Eeuo pipefail
test "$confirmation" = ROLLBACK
sudo -n virsh vcpupin uv059 0 0-15 --live --config
host_cpu_count="$(nproc)"
sudo -n virsh vcpupin uv059 1 0-15 --live --config
host_cpu_set="0-$((host_cpu_count - 1))"
sudo -n virsh emulatorpin uv059 0-15 --live --config
for index in $(seq 1 "$KITSNET_CANDIDATE_COUNT"); do
sudo -n virsh vcpupin uv060 0 0-15 --live --config
    vm_var="KITSNET_CANDIDATE_${index}_VM"
sudo -n virsh vcpupin uv060 1 0-15 --live --config
    vm="${!vm_var}"
sudo -n virsh emulatorpin uv060 0-15 --live --config
    sudo -n virsh vcpupin "$vm" --config --live 0 "$host_cpu_set"
REMOTE
    vcpu_count="$(sudo -n virsh dominfo "$vm" | awk '$1 == "CPU(s):" {print $2}')"
BASH
    for vcpu in $(seq 1 $((vcpu_count - 1))); do
then
        sudo -n virsh vcpupin "$vm" --config --live "$vcpu" "$host_cpu_set"
    echo "rollback=PASS"
    done
else
    sudo -n virsh emulatorpin "$vm" "$host_cpu_set" --config --live
    echo "rollback=$?"
done
fi
echo "rollback=PASS"
echo "The interactive shell remains open."
echo "The interactive shell remains open."
</syntaxhighlight>
</syntaxhighlight>
Line 280: Line 358:
for i in {1..10}; do echo; done
for i in {1..10}; do echo; done
if bash <<'BASH'
if bash <<'BASH'
set -Eeuo pipefail
test "$(id -un)" = psmode
ssh -o BatchMode=yes psmode@wort.lan.kitsnet.us 'bash -s' <<'REMOTE'
set -Eeuo pipefail
set -Eeuo pipefail
echo "===== SAMPLE 1 ====="
echo "===== SAMPLE 1 ====="
sudo -n virsh domstats uv059 uv060 --vcpu | \
target_env="/tmp/kitsnet-kvm-candidates.${UID}.env"
    grep -E '^(Domain|  vcpu\.[01]\.(time|delay))' || true
test -r "$target_env"
. "$target_env"
for index in $(seq 1 "$KITSNET_CANDIDATE_COUNT"); do
    label_var="KITSNET_CANDIDATE_${index}_LABEL"
    vm_var="KITSNET_CANDIDATE_${index}_VM"
    label="${!label_var}"
    vm="${!vm_var}"
    echo "===== ${label} (${vm}) SAMPLE 1 ====="
    sudo -n virsh domstats "$vm" --vcpu | \
        grep -E '^(Domain|  vcpu\.[0-9]+\.(time|delay))' || true
done
sleep 30
sleep 30
echo "===== SAMPLE 2 ====="
for index in $(seq 1 "$KITSNET_CANDIDATE_COUNT"); do
sudo -n virsh domstats uv059 uv060 --vcpu | \
    label_var="KITSNET_CANDIDATE_${index}_LABEL"
    grep -E '^(Domain|  vcpu\.[01]\.(time|delay))' || true
    vm_var="KITSNET_CANDIDATE_${index}_VM"
REMOTE
    label="${!label_var}"
    vm="${!vm_var}"
    echo "===== ${label} (${vm}) SAMPLE 2 ====="
    sudo -n virsh domstats "$vm" --vcpu | \
        grep -E '^(Domain|  vcpu\.[0-9]+\.(time|delay))' || true
done
echo "===== KEEPALIVED TIMER EVENTS ====="
echo "===== KEEPALIVED TIMER EVENTS ====="
for peer in fox anchor; do
for index in $(seq 1 "$KITSNET_CANDIDATE_COUNT"); do
     echo "===== $peer ====="
    label_var="KITSNET_CANDIDATE_${index}_LABEL"
     ssh -n -o BatchMode=yes "psmode@$peer.lan.kitsnet.us" \
    host_var="KITSNET_CANDIDATE_${index}_HOST"
    label="${!label_var}"
    host="${!host_var}"
     echo "===== ${label} ====="
     ssh -n -o BatchMode=yes "psmode@${host}" \
         "sudo -n journalctl -u keepalived --since '10 minutes ago' --no-pager | grep -E 'timer expired|Entering|Leaving' || true"
         "sudo -n journalctl -u keepalived --since '10 minutes ago' --no-pager | grep -E 'timer expired|Entering|Leaving' || true"
done
done
Line 372: Line 466:
|-
|-
| wrk2 || Docker Swarm worker; host/domain to be recorded || To be recorded || Unpinned || Not required || Benchmark workload only; pinning would alter methodology || No pinning test performed || Only for a separately controlled benchmark
| wrk2 || Docker Swarm worker; host/domain to be recorded || To be recorded || Unpinned || Not required || Benchmark workload only; pinning would alter methodology || No pinning test performed || Only for a separately controlled benchmark
|-
| Zabbix || wort / <code>uv052</code> || 4 || vCPUs and emulator allowed on 0-15 || Not applied || 1.5-3.2% per-vCPU delay rate during Veeam SQL-dump I/O; host idle 60-80%, I/O wait 0%, steal 0% || 2026-09-05 30-second stress sample || Repeat after backup or during a heavier phase
|}
|}



Latest revision as of 10:40, 5 September 2026

This guide defines when CPU pinning is appropriate for KitsNet KVM/libvirt virtual machines, how to evaluate the need, how to implement it safely, and how to measure whether it helped. It is intended for internal operations on the KitsNet virtualization and Docker Swarm environment.

The guidance distinguishes observed facts from recommendations. CPU pinning is a scheduling-control technique; it is not a substitute for fixing storage, network, guest, or hypervisor faults.

1 Scope and operating rules[edit | edit source]

This guide applies to KVM/libvirt guests hosted on KitsNet hypervisors such as wort. It does not require pinning Docker Swarm workers merely because a benchmark or service runs in a container. A Swarm worker should be pinned only when a separate, controlled performance experiment requires it.

The following operating rules apply:

  • Use the authenticated psmode account for outbound SSH from mgr1.
  • Do not SSH or SCP into root on a KitsNet system, except radix.kitsnet.us. Outbound SSH or SCP from a system as root is permitted when required by an approved operation.
  • Do not reconstruct SSH credentials with a root shell on mgr1; the loaded keys are associated with psmode.
  • Use explicit host names and sudo -n on the remote host. Do not use unrestricted root SSH.
  • For production NAS storage ownership, use the KitsNet storage transition and fencing workflow. Do not use ad-hoc virsh attach/detach operations to manipulate the shared production LV.
  • CPU placement changes may use the approved libvirt management path, but must be recorded, reviewed, and reversible.
  • Operational source belongs in /srv/git/KNdkr. This design and operations guide remains documentation outside that repository unless a separate documentation decision is made.

2 Design rationale[edit | edit source]

CPU utilization reported inside a guest is not the same as prompt vCPU scheduling. A guest can report a low load average while a runnable vCPU waits in the host scheduler. Useful evidence includes:

  • guest %steal and load;
  • libvirt vcpu.*.delay counters;
  • host CPU topology and run-queue pressure;
  • vCPU, emulator-thread, and I/O-thread placement;
  • Keepalived timer-expiration messages or other latency-sensitive symptoms.

Pinning gives a guest a defined set of host logical CPUs. It does not automatically reserve those CPUs: other guests may still be allowed to run there unless their placement is also constrained or host-level isolation is configured. Pinning therefore improves placement determinism, but the resulting isolation must be verified.

2.1 When pinning is usually justified[edit | edit source]

Consider pinning when one or more of these conditions are present and ordinary tuning is insufficient:

  • latency-sensitive HA, real-time, packet-processing, or control-plane software shows missed timers or unacceptable wake-up latency;
  • libvirt vCPU delay is materially increasing while guest CPU utilization remains low;
  • a guest has repeatable tail-latency failures during host activity, storage I/O, or backup windows;
  • a VM is a dedicated appliance whose performance must be reproducible;
  • a controlled benchmark needs repeatable CPU placement and the changed methodology is documented;
  • multiple important VMs interfere with one another and host-level CPU capacity is available for separation.

2.2 When pinning is usually not justified[edit | edit source]

Do not pin by default for ordinary servers, Docker Swarm workers, or general-purpose guests. Avoid it when:

  • the host is CPU-saturated and there are no spare physical cores;
  • the guest is routinely resized or migrated across hosts with different topology;
  • the observed problem is actually storage latency, packet loss, memory pressure, or a guest process fault;
  • pinning would make capacity planning or failover less flexible;
  • the measurement plan cannot distinguish improvement from normal run-to-run variation.

2.3 Decision threshold[edit | edit source]

Use the libvirt vCPU delay rate as the primary numeric screening metric:

For a general-purpose candidate, treat pinning as warranted when the per-vCPU delay rate is at least 5% in three consecutive 60-second samples under a representative workload, while the guest is not itself CPU-saturated. A rate of 10% or more is a high-priority scheduling problem and should be investigated immediately.

For a latency-sensitive HA candidate, use the stricter event rule: one Keepalived timer-expiration event delayed by 2 seconds or more, or two shorter timer-expiration events in a 15-minute window, is sufficient to investigate and normally justify a pinning trial when the event correlates with elevated vCPU delay. A timer event alone does not prove CPU scheduling is the cause; check network, storage, and guest evidence as well.

These are KitsNet operating thresholds, not universal KVM limits. They provide a repeatable change decision and can be revised when enough local history exists. The 2026-09-05 zaya sample (approximately 1.5-3.2% delay rate during a Veeam SQL-dump workload, with host idle 60-80%, I/O wait 0%, and steal 0%) remained below the 5% general-purpose threshold and was left unpinned. Fox and anchor had timer delays and much higher cumulative scheduling delay, so a pinning trial was justified.

3 Evaluation workflow[edit | edit source]

Evaluate a candidate in five stages: baseline, topology, correlation, controlled change, and post-change comparison.

3.1 1. Establish a baseline[edit | edit source]

Record the guest state, workload, time window, vCPU count, memory, and any relevant event timestamps. Do not compare a busy backup window with an idle window.

Begin on the KVM host and collect one or more candidates interactively. This avoids baking host or VM names into later code blocks. The label is a human-readable tracking name; the OS hostname is used for guest SSH; and the libvirt VM/domain name is used for virsh.

for i in {1..10}; do echo; done
collect_kitsnet_candidates() {
    local target_env="/tmp/kitsnet-kvm-candidates.${UID}.env"
    local candidate_count index label host vm

    read -r -p "Number of VM candidates (1-32): " candidate_count
    case "$candidate_count" in
        ''|*[!0-9]*) echo "Invalid candidate count" >&2; return 1 ;;
    esac
    if [ "$candidate_count" -lt 1 ] || [ "$candidate_count" -gt 32 ]; then
        echo "Candidate count must be between 1 and 32" >&2
        return 1
    fi

    umask 077
    {
        printf 'export KITSNET_HYPERVISOR_HOST=%q\n' "$(hostname -f)"
        printf 'export KITSNET_CANDIDATE_COUNT=%q\n' "$candidate_count"
    } > "$target_env"

    for index in $(seq 1 "$candidate_count"); do
        echo "===== CANDIDATE $index ====="
        read -r -p "Candidate label: " label
        read -r -p "Candidate OS hostname/FQDN: " host
        read -r -p "Candidate libvirt VM/domain name: " vm

        case "$label" in ''|*[!A-Za-z0-9_.-]*) echo "Invalid label" >&2; return 1 ;; esac
        case "$host" in ''|*[!A-Za-z0-9.-]*) echo "Invalid hostname" >&2; return 1 ;; esac
        case "$vm" in ''|*[!A-Za-z0-9_.-]*) echo "Invalid VM name" >&2; return 1 ;; esac

        if ! sudo -n virsh dominfo "$vm" >/dev/null 2>&1; then
            echo "VM/domain not found: $vm" >&2
            return 1
        fi

        printf 'export KITSNET_CANDIDATE_%s_LABEL=%q\n' "$index" "$label" >> "$target_env"
        printf 'export KITSNET_CANDIDATE_%s_HOST=%q\n' "$index" "$host" >> "$target_env"
        printf 'export KITSNET_CANDIDATE_%s_VM=%q\n' "$index" "$vm" >> "$target_env"
    done

    chmod 600 "$target_env"
    echo "target_environment=$target_env"
    cat "$target_env"
}

collect_kitsnet_candidates
collection_rc=$?
echo "candidate_collection=$collection_rc"
echo "The interactive shell remains open."

The collection block is intentionally run directly in the interactive shell rather than inside a heredoc-driven child shell; this keeps prompts attached to the terminal. Later blocks source the resulting environment file. It connects as psmode when guest SSH is used and does not use root SSH.

The following baseline block is noninteractive and evaluates every collected candidate:

for i in {1..10}; do echo; done
if bash <<'BASH'
set -Eeuo pipefail
target_env="/tmp/kitsnet-kvm-candidates.${UID}.env"
test -r "$target_env"
. "$target_env"
for index in $(seq 1 "$KITSNET_CANDIDATE_COUNT"); do
    label_var="KITSNET_CANDIDATE_${index}_LABEL"
    host_var="KITSNET_CANDIDATE_${index}_HOST"
    label="${!label_var}"
    host="${!host_var}"
    echo "===== ${label} GUEST ====="
    if ssh -n -o BatchMode=yes "psmode@${host}" \
        "hostname -f; awk '/^cpu([0-9 ]|$)/ {print}' /proc/stat; uptime"; then
        echo "guest_baseline=PASS"
    else
        echo "guest_baseline=FAIL"
    fi
done
BASH
then
    echo "baseline_collection=PASS"
else
    echo "baseline_collection=$?"
fi
echo "The interactive shell remains open."

The guest CPU line has the standard Linux fields: user, nice, system, idle, iowait, irq, softirq, steal, guest, and guest_nice. Take two samples over a known interval if a current steal percentage is required; the cumulative line alone is not a complete time-series measurement.

3.2 2. Map the host topology[edit | edit source]

Identify physical cores and SMT siblings before selecting CPU sets. On a single-socket AMD Ryzen 7 5800X, the observed topology was eight physical cores and sixteen logical CPUs, with logical pairs 0/8, 1/9, through 7/15.

for i in {1..10}; do echo; done
if bash <<'BASH'
set -Eeuo pipefail
test "$(id -un)" = psmode
ssh -o BatchMode=yes psmode@wort.lan.kitsnet.us 'bash -s' <<'REMOTE'
set -Eeuo pipefail
hostname -f
lscpu -e=CPU,CORE,SOCKET,NODE
lscpu | grep -E '^(CPU\(s\)|Thread|Core|Socket|NUMA|Model name):' || true
REMOTE
BASH
then
    echo "topology_collection=PASS"
else
    echo "topology_collection=$?"
fi
echo "The interactive shell remains open."

Do not assume CPU numbering on another host. Repeat the topology query for every hypervisor.

3.3 3. Correlate libvirt scheduling evidence[edit | edit source]

Inspect domain placement, scheduler settings, and vCPU delay. The vcpu.*.delay value is cumulative nanoseconds for time a vCPU was runnable but not scheduled; compare deltas over equal intervals rather than comparing only absolute totals.

for i in {1..10}; do echo; done
if bash <<'BASH'
set -Eeuo pipefail
target_env="/tmp/kitsnet-kvm-candidates.${UID}.env"
test -r "$target_env"
. "$target_env"
for index in $(seq 1 "$KITSNET_CANDIDATE_COUNT"); do
    label_var="KITSNET_CANDIDATE_${index}_LABEL"
    vm_var="KITSNET_CANDIDATE_${index}_VM"
    label="${!label_var}"
    vm="${!vm_var}"
    echo "===== ${label} (${vm}) ====="
    sudo -n virsh dominfo "$vm" | grep -E '^(Name|UUID|CPU\(s\)|CPU time|State):' || true
    sudo -n virsh vcpupin "$vm"
    sudo -n virsh emulatorpin "$vm"
    sudo -n virsh schedinfo "$vm" || true
    sudo -n virsh domstats "$vm" --vcpu | \
        grep -E '^(Domain|  vcpu\.[0-9]+\.(time|delay))' || true
done
BASH
then
    echo "libvirt_baseline=PASS"
else
    echo "libvirt_baseline=$?"
fi
echo "The interactive shell remains open."

During the NAS incident, both uv059 (fox) and uv060 (anchor) had every vCPU and emulator thread allowed on CPUs 0-15. Anchor's cumulative vCPU delay was approximately 417 seconds per vCPU over approximately 1,527 seconds of vCPU runtime; fox's was approximately 55 seconds over approximately 673 seconds. This was strong evidence of scheduling delay despite near-zero guest load, but it was not by itself proof of the root cause.

3.4 4. Check competing capacity[edit | edit source]

List all running domains and their vCPU counts. A host with sixteen logical CPUs and substantially more runnable vCPUs may need placement priority, but pinning every guest can create artificial contention. Inspect before changing any domain.

for i in {1..10}; do echo; done
if bash <<'BASH'
set -Eeuo pipefail
sudo -n virsh list --all
for domain in $(sudo -n virsh list --name | sed '/^$/d'); do
    sudo -n virsh dominfo "$domain" | \
        grep -E '^(Name|CPU\(s\)|CPU time|State):' || true
done
BASH
then
    echo "capacity_inventory=PASS"
else
    echo "capacity_inventory=$?"
fi
echo "The interactive shell remains open."

4 Implementation[edit | edit source]

4.1 Select CPU sets[edit | edit source]

Select physical cores, not two SMT siblings of the same core, when the goal is physical-core separation. Keep each HA VM on a separate pair when possible. Also decide where its emulator and virtio/I/O threads will run. Pinning only vCPUs while leaving an overloaded emulator thread unrestricted may not solve the problem.

For the 2026-09-05 NAS test, the selected temporary placement was:

VM libvirt domain vCPU 0 vCPU 1 emulator Reason
fox uv059 CPU 0 CPU 1 CPUs 8-9 Separate physical cores; emulator on SMT siblings
anchor uv060 CPU 2 CPU 3 CPUs 10-11 Separate physical cores from fox

Table 1. Example HA VM CPU placement used for the KitsNet test.

This arrangement improves determinism for the two HA guests, but it does not fully reserve CPUs 0-3 because other guests were still allowed on the full 0-15 set. Full isolation would require a broader, host-wide capacity plan.

4.2 Apply live and persistent placement[edit | edit source]

The following commands change the running domain and persistent libvirt definition. Confirm domain names and topology first. Do not apply this pattern to production storage attach/detach operations.

for i in {1..10}; do echo; done
target_env="/tmp/kitsnet-kvm-candidates.${UID}.env"
test -r "$target_env"
. "$target_env"
apply_candidate_pinning() {
    local index label_var vm_var label vm vcpu_count emulator_cpuset pin_line vcpu cpu
    local -a pin_sets
    for index in $(seq 1 "$KITSNET_CANDIDATE_COUNT"); do
        label_var="KITSNET_CANDIDATE_${index}_LABEL"
        vm_var="KITSNET_CANDIDATE_${index}_VM"
        label="${!label_var}"
        vm="${!vm_var}"
        vcpu_count="$(sudo -n virsh dominfo "$vm" | awk '$1 == "CPU(s):" {print $2}')"
        test "$vcpu_count" -ge 1
        echo "===== ${label} (${vm}) ====="
        read -r -p "Host CPU set for each of ${vcpu_count} vCPUs (space-separated): " pin_line
        read -r -a pin_sets <<< "$pin_line"
        test "${#pin_sets[@]}" -eq "$vcpu_count"
        for vcpu in $(seq 0 $((vcpu_count - 1))); do
            cpu="${pin_sets[$vcpu]}"
            [[ "$cpu" =~ ^[0-9,-]+$ ]]
            sudo -n virsh vcpupin "$vm" "$vcpu" "$cpu" --live --config
        done
        read -r -p "Emulator-thread CPU set: " emulator_cpuset
        [[ "$emulator_cpuset" =~ ^[0-9,-]+$ ]]
        sudo -n virsh emulatorpin "$vm" "$emulator_cpuset" --live --config
        sudo -n virsh vcpupin "$vm"
        sudo -n virsh emulatorpin "$vm"
    done
}
apply_candidate_pinning
echo "The interactive shell remains open."

The resulting persistent XML should contain entries equivalent to:

<vcpu placement='static'>2</vcpu>
<vcpupin vcpu='0' cpuset='0'/>
<vcpupin vcpu='1' cpuset='1'/>
<emulatorpin cpuset='8-9'/>

Do not hand-edit generated libvirt XML while the domain is running. Use the approved libvirt management command, then verify the resulting XML and live placement.

4.3 Rollback[edit | edit source]

Rollback is done by restoring an unrestricted CPU set for the affected domain. Use the exact prior placement recorded in the change log when possible. The following interactive block operates on every collected candidate only after an explicit confirmation:

for i in {1..10}; do echo; done
target_env="/tmp/kitsnet-kvm-candidates.${UID}.env"
test -r "$target_env"
. "$target_env"
read -r -p "Type ROLLBACK to clear pinning for all collected candidates: " confirmation
test "$confirmation" = ROLLBACK
host_cpu_count="$(nproc)"
host_cpu_set="0-$((host_cpu_count - 1))"
for index in $(seq 1 "$KITSNET_CANDIDATE_COUNT"); do
    vm_var="KITSNET_CANDIDATE_${index}_VM"
    vm="${!vm_var}"
    sudo -n virsh vcpupin "$vm" --config --live 0 "$host_cpu_set"
    vcpu_count="$(sudo -n virsh dominfo "$vm" | awk '$1 == "CPU(s):" {print $2}')"
    for vcpu in $(seq 1 $((vcpu_count - 1))); do
        sudo -n virsh vcpupin "$vm" --config --live "$vcpu" "$host_cpu_set"
    done
    sudo -n virsh emulatorpin "$vm" "$host_cpu_set" --config --live
done
echo "rollback=PASS"
echo "The interactive shell remains open."

4.4 Resource and migration considerations[edit | edit source]

Pinning must be reviewed against live migration, host maintenance, and recovery. A CPU set valid on one hypervisor may not exist on another. If a guest can migrate, define an equivalent placement on every destination or remove pinning before migration. Do not enable host-wide isolcpus or real-time scheduling solely from one incident without a capacity and recovery review.

5 Evaluating effectiveness[edit | edit source]

Evaluate both latency and service correctness. A successful change should show a lower rate of vCPU delay accumulation, fewer timer-expiration events, and no new failover or storage errors under a comparable workload.

5.1 Immediate validation[edit | edit source]

Verify the live and persistent settings and compare two domstats samples over 30-60 seconds. Calculate:

Also check guest %steal, load, and service logs during the same interval. The post-pinning NAS sample showed only approximately 0.1 seconds of new delay per vCPU over 30 seconds while idle, and no new fox Keepalived timer-expiration messages.

for i in {1..10}; do echo; done
if bash <<'BASH'
set -Eeuo pipefail
echo "===== SAMPLE 1 ====="
target_env="/tmp/kitsnet-kvm-candidates.${UID}.env"
test -r "$target_env"
. "$target_env"
for index in $(seq 1 "$KITSNET_CANDIDATE_COUNT"); do
    label_var="KITSNET_CANDIDATE_${index}_LABEL"
    vm_var="KITSNET_CANDIDATE_${index}_VM"
    label="${!label_var}"
    vm="${!vm_var}"
    echo "===== ${label} (${vm}) SAMPLE 1 ====="
    sudo -n virsh domstats "$vm" --vcpu | \
        grep -E '^(Domain|  vcpu\.[0-9]+\.(time|delay))' || true
done
sleep 30
for index in $(seq 1 "$KITSNET_CANDIDATE_COUNT"); do
    label_var="KITSNET_CANDIDATE_${index}_LABEL"
    vm_var="KITSNET_CANDIDATE_${index}_VM"
    label="${!label_var}"
    vm="${!vm_var}"
    echo "===== ${label} (${vm}) SAMPLE 2 ====="
    sudo -n virsh domstats "$vm" --vcpu | \
        grep -E '^(Domain|  vcpu\.[0-9]+\.(time|delay))' || true
done
echo "===== KEEPALIVED TIMER EVENTS ====="
for index in $(seq 1 "$KITSNET_CANDIDATE_COUNT"); do
    label_var="KITSNET_CANDIDATE_${index}_LABEL"
    host_var="KITSNET_CANDIDATE_${index}_HOST"
    label="${!label_var}"
    host="${!host_var}"
    echo "===== ${label} ====="
    ssh -n -o BatchMode=yes "psmode@${host}" \
        "sudo -n journalctl -u keepalived --since '10 minutes ago' --no-pager | grep -E 'timer expired|Entering|Leaving' || true"
done
BASH
then
    echo "post_pin_validation=PASS"
else
    echo "post_pin_validation=$?"
fi
echo "The interactive shell remains open."

5.2 Controlled workload validation[edit | edit source]

Repeat the original workload or a deliberately bounded test, recording the exact image, workload, duration, storage path, and host. For HA guests, test failover separately from application benchmarking. A benchmark can expose a scheduling problem, but a failover test must verify exclusive storage ownership, VIP movement, mount state, and service transition logs.

In the 2026-09-05 test, the NAS sequence was:

  1. Fox was primary with the VIP and /dkr/data1 mounted.
  2. Fox Keepalived was stopped after a completed backup.
  3. Anchor acquired the production LV, mounted /dev/vdb1, started NFS, and claimed the VIP in approximately six seconds.
  4. Anchor was stopped; fox reacquired the LV and VIP immediately on the observed poll.
  5. Anchor was restarted and rejoined as BACKUP with no VIP and no production mount.

That result validates HA transition behavior after pinning and transition serialization; it does not establish a benchmark performance improvement by itself.

5.3 Failure interpretation[edit | edit source]

  • If guest %steal and libvirt delay fall, but timer failures remain, inspect guest scheduling, Keepalived priority, network loss, and transition scripts.
  • If delay remains high after pinning, inspect host contention, emulator/I/O-thread placement, storage stalls, and host power-management behavior.
  • If delay improves but application latency worsens, the pinned cores may be contended by other guests or the workload may have changed.
  • If HA state or storage ownership fails, stop benchmarking and preserve fencing safeguards. Do not suppress alerts to make the test appear successful.

6 Monitoring and alerting[edit | edit source]

At minimum, monitor:

  • Keepalived state and timer expired messages;
  • VIP presence on exactly one node;
  • production filesystem source and mount state;
  • exclusive production-LV ownership through the NAS fencing workflow;
  • vCPU delay-rate samples on the hypervisor;
  • guest steal time and host CPU pressure;
  • ConsoleWorks event codes.

KitsNet ConsoleWorks mappings for NAS failover logs are:

Syslog PRI Facility/severity Intended event class
<138> local1.crit Critical
<139> local1.err Error
<140> local1.warning Warning
<141> local1.notice Notice

Table 2. KitsNet NAS ConsoleWorks local1 event-code mapping.

The nas-failover logger must specify local1 explicitly. A bare logger -t nas-failover defaults to the user facility and produces <13> for notice, which bypasses the intended local1 scans.

7 Tracking table[edit | edit source]

Use this table as the living record. Update it after each baseline, placement change, rollback, and validation window.

Candidate Hypervisor/domain vCPUs Current placement Pinning status Evidence/trigger Last validation Next review
fox wort / uv059 2 vCPU 0→0, vCPU 1→1; emulator→8-9 Applied Keepalived timer delays; elevated cumulative vCPU delay 2026-09-05 controlled failover passed After next HA observation window
anchor wort / uv060 2 vCPU 0→2, vCPU 1→3; emulator→10-11 Applied Keepalived timer delays; higher cumulative vCPU delay 2026-09-05 controlled failover and rejoin passed After next HA observation window
wrk1 Docker Swarm worker; host/domain to be recorded To be recorded Unpinned Not required Benchmark workload only; pinning would alter methodology No pinning test performed Only for a separately controlled benchmark
wrk2 Docker Swarm worker; host/domain to be recorded To be recorded Unpinned Not required Benchmark workload only; pinning would alter methodology No pinning test performed Only for a separately controlled benchmark
Zabbix wort / uv052 4 vCPUs and emulator allowed on 0-15 Not applied 1.5-3.2% per-vCPU delay rate during Veeam SQL-dump I/O; host idle 60-80%, I/O wait 0%, steal 0% 2026-09-05 30-second stress sample Repeat after backup or during a heavier phase

Table 3. Initial KitsNet CPU-pinning candidate register.

For each additional candidate, record the host CPU topology, vCPU and emulator placement, baseline delay rate, workload, guest steal percentage, host pressure, change timestamp, rollback point, and post-change result. A candidate remains Applied only while the measured benefit persists and the placement is compatible with migration and recovery.

8 Change record for the 2026-09-05 NAS incident[edit | edit source]

The following facts are verified from the recorded command output:

  • Hypervisor: wort; CPU: AMD Ryzen 7 5800X, eight physical cores and sixteen logical CPUs.
  • Domains: uv059=fox and uv060=anchor, each with two vCPUs.
  • Initial placement: all vCPUs and emulator threads allowed on CPUs 0-15.
  • Guest load was near zero, while cumulative libvirt vCPU delay was significant, especially on anchor.
  • Persistent live-and-config placement was applied as shown in Table 1.
  • The NAS failover test passed in both directions. Fox ended primary; anchor ended active BACKUP.
  • SELinux hostname permissions were narrowed to keepalived_t and hostname_exec_t; no new AVCs appeared after installation.
  • NAS transition scripts were updated to use local1 logging, correct storage-operation return status, and a host-local transition lock.

These changes explain why pinning was justified for fox and anchor. They do not imply that every KitsNet VM requires pinning.

9 Operational checklist[edit | edit source]

  1. Establish a time-bounded guest and host baseline.
  2. Map physical cores and SMT siblings.
  3. Inventory all domain vCPUs and current placement.
  4. Correlate delay, steal, timer, storage, and network evidence.
  5. Document the candidate, selected CPU sets, and rollback.
  6. Apply live and persistent placement through approved libvirt management.
  7. Verify placement and syntax/configuration state.
  8. Compare delay-rate and service behavior under a comparable workload.
  9. Test migration, failover, and recovery implications.
  10. Update Table 3 and retain the change backup/checksum record.

Do not declare pinning effective solely because a single command succeeded. It is effective only when the relevant latency or correctness symptom improves without creating a capacity, migration, or recovery regression.