KVM:CPU Pinning Guidance: Difference between revisions

Peter A. Smode (talk | contribs)
Created page with "This guide defines when CPU pinning is appropriate for KitsNet KVM/libvirt virtual machines, how to evaluate the need, how to implement it safely, and how to measure whether it helped. It is intended for internal operations on the KitsNet virtualization and Docker Swarm environment. The guidance distinguishes observed facts from recommendations. CPU pinning is a scheduling-control technique; it is not a substitute for fixing storage, network, guest, or hypervisor faults..."
 
Peter A. Smode (talk | contribs)
Line 66: Line 66:
test "$(id -un)" = psmode
test "$(id -un)" = psmode
for peer in fox anchor; do
for peer in fox anchor; do
     echo "===== $peer GUEST ====="
     echo "===== ${peer} GUEST ====="
     ssh -n -o BatchMode=yes "psmode@$peer.lan.kitsnet.us" \
     ssh -n -o BatchMode=yes "psmode@${peer}.lan.kitsnet.us" \
         "hostname -f; awk '/^cpu([0-9 ]|$)/ {print}' /proc/stat; uptime"
         "hostname -f; awk '/^cpu([0-9 ]|$)/ {print}' /proc/stat; uptime"
done
done

Revision as of 10:04, 5 September 2026

This guide defines when CPU pinning is appropriate for KitsNet KVM/libvirt virtual machines, how to evaluate the need, how to implement it safely, and how to measure whether it helped. It is intended for internal operations on the KitsNet virtualization and Docker Swarm environment.

The guidance distinguishes observed facts from recommendations. CPU pinning is a scheduling-control technique; it is not a substitute for fixing storage, network, guest, or hypervisor faults.

1 Scope and operating rules[edit | edit source]

This guide applies to KVM/libvirt guests hosted on KitsNet hypervisors such as wort. It does not require pinning Docker Swarm workers merely because a benchmark or service runs in a container. A Swarm worker should be pinned only when a separate, controlled performance experiment requires it.

The following operating rules apply:

  • Use the authenticated psmode account for outbound SSH from mgr1.
  • Do not SSH or SCP into root on a KitsNet system, except radix.kitsnet.us. Outbound SSH or SCP from a system as root is permitted when required by an approved operation.
  • Do not reconstruct SSH credentials with a root shell on mgr1; the loaded keys are associated with psmode.
  • Use explicit host names and sudo -n on the remote host. Do not use unrestricted root SSH.
  • For production NAS storage ownership, use the KitsNet storage transition and fencing workflow. Do not use ad-hoc virsh attach/detach operations to manipulate the shared production LV.
  • CPU placement changes may use the approved libvirt management path, but must be recorded, reviewed, and reversible.
  • Operational source belongs in /srv/git/KNdkr. This design and operations guide remains documentation outside that repository unless a separate documentation decision is made.

2 Design rationale[edit | edit source]

CPU utilization reported inside a guest is not the same as prompt vCPU scheduling. A guest can report a low load average while a runnable vCPU waits in the host scheduler. Useful evidence includes:

  • guest %steal and load;
  • libvirt vcpu.*.delay counters;
  • host CPU topology and run-queue pressure;
  • vCPU, emulator-thread, and I/O-thread placement;
  • Keepalived timer-expiration messages or other latency-sensitive symptoms.

Pinning gives a guest a defined set of host logical CPUs. It does not automatically reserve those CPUs: other guests may still be allowed to run there unless their placement is also constrained or host-level isolation is configured. Pinning therefore improves placement determinism, but the resulting isolation must be verified.

2.1 When pinning is usually justified[edit | edit source]

Consider pinning when one or more of these conditions are present and ordinary tuning is insufficient:

  • latency-sensitive HA, real-time, packet-processing, or control-plane software shows missed timers or unacceptable wake-up latency;
  • libvirt vCPU delay is materially increasing while guest CPU utilization remains low;
  • a guest has repeatable tail-latency failures during host activity, storage I/O, or backup windows;
  • a VM is a dedicated appliance whose performance must be reproducible;
  • a controlled benchmark needs repeatable CPU placement and the changed methodology is documented;
  • multiple important VMs interfere with one another and host-level CPU capacity is available for separation.

2.2 When pinning is usually not justified[edit | edit source]

Do not pin by default for ordinary servers, Docker Swarm workers, or general-purpose guests. Avoid it when:

  • the host is CPU-saturated and there are no spare physical cores;
  • the guest is routinely resized or migrated across hosts with different topology;
  • the observed problem is actually storage latency, packet loss, memory pressure, or a guest process fault;
  • pinning would make capacity planning or failover less flexible;
  • the measurement plan cannot distinguish improvement from normal run-to-run variation.

3 Evaluation workflow[edit | edit source]

Evaluate a candidate in five stages: baseline, topology, correlation, controlled change, and post-change comparison.

3.1 1. Establish a baseline[edit | edit source]

Record the guest state, workload, time window, vCPU count, memory, and any relevant event timestamps. Do not compare a busy backup window with an idle window.

The following is a read-only example used during the 2026-09-05 KitsNet NAS investigation. It connects as psmode and does not use root SSH.

for i in {1..10}; do echo; done
if bash <<'BASH'
set -Eeuo pipefail
test "$(id -un)" = psmode
for peer in fox anchor; do
    echo "===== ${peer} GUEST ====="
    ssh -n -o BatchMode=yes "psmode@${peer}.lan.kitsnet.us" \
        "hostname -f; awk '/^cpu([0-9 ]|$)/ {print}' /proc/stat; uptime"
done
BASH
then
    echo "baseline_collection=PASS"
else
    echo "baseline_collection=$?"
fi
echo "The interactive shell remains open."

The guest CPU line has the standard Linux fields: user, nice, system, idle, iowait, irq, softirq, steal, guest, and guest_nice. Take two samples over a known interval if a current steal percentage is required; the cumulative line alone is not a complete time-series measurement.

3.2 2. Map the host topology[edit | edit source]

Identify physical cores and SMT siblings before selecting CPU sets. On a single-socket AMD Ryzen 7 5800X, the observed topology was eight physical cores and sixteen logical CPUs, with logical pairs 0/8, 1/9, through 7/15.

for i in {1..10}; do echo; done
if bash <<'BASH'
set -Eeuo pipefail
test "$(id -un)" = psmode
ssh -o BatchMode=yes psmode@wort.lan.kitsnet.us 'bash -s' <<'REMOTE'
set -Eeuo pipefail
hostname -f
lscpu -e=CPU,CORE,SOCKET,NODE
lscpu | grep -E '^(CPU\(s\)|Thread|Core|Socket|NUMA|Model name):' || true
REMOTE
BASH
then
    echo "topology_collection=PASS"
else
    echo "topology_collection=$?"
fi
echo "The interactive shell remains open."

Do not assume CPU numbering on another host. Repeat the topology query for every hypervisor.

3.3 3. Correlate libvirt scheduling evidence[edit | edit source]

Inspect domain placement, scheduler settings, and vCPU delay. The vcpu.*.delay value is cumulative nanoseconds for time a vCPU was runnable but not scheduled; compare deltas over equal intervals rather than comparing only absolute totals.

for i in {1..10}; do echo; done
if bash <<'BASH'
set -Eeuo pipefail
test "$(id -un)" = psmode
ssh -o BatchMode=yes psmode@wort.lan.kitsnet.us 'bash -s' <<'REMOTE'
set -Eeuo pipefail
for domain in uv059 uv060; do
    echo "===== $domain ====="
    sudo -n virsh dominfo "$domain" | grep -E '^(Name|UUID|CPU\(s\)|CPU time|State):' || true
    sudo -n virsh vcpupin "$domain"
    sudo -n virsh emulatorpin "$domain"
    sudo -n virsh schedinfo "$domain" || true
    sudo -n virsh domstats "$domain" --vcpu | \
        grep -E '^(Domain|  vcpu\.[0-9]+\.(time|delay))' || true
done
REMOTE
BASH
then
    echo "libvirt_baseline=PASS"
else
    echo "libvirt_baseline=$?"
fi
echo "The interactive shell remains open."

During the NAS incident, both uv059 (fox) and uv060 (anchor) had every vCPU and emulator thread allowed on CPUs 0-15. Anchor's cumulative vCPU delay was approximately 417 seconds per vCPU over approximately 1,527 seconds of vCPU runtime; fox's was approximately 55 seconds over approximately 673 seconds. This was strong evidence of scheduling delay despite near-zero guest load, but it was not by itself proof of the root cause.

3.4 4. Check competing capacity[edit | edit source]

List all running domains and their vCPU counts. A host with sixteen logical CPUs and substantially more runnable vCPUs may need placement priority, but pinning every guest can create artificial contention. Inspect before changing any domain.

for i in {1..10}; do echo; done
if bash <<'BASH'
set -Eeuo pipefail
test "$(id -un)" = psmode
ssh -o BatchMode=yes psmode@wort.lan.kitsnet.us 'bash -s' <<'REMOTE'
set -Eeuo pipefail
sudo -n virsh list --all
for domain in $(sudo -n virsh list --name | sed '/^$/d'); do
    sudo -n virsh dominfo "$domain" | \
        grep -E '^(Name|CPU\(s\)|CPU time|State):' || true
done
REMOTE
BASH
then
    echo "capacity_inventory=PASS"
else
    echo "capacity_inventory=$?"
fi
echo "The interactive shell remains open."

4 Implementation[edit | edit source]

4.1 Select CPU sets[edit | edit source]

Select physical cores, not two SMT siblings of the same core, when the goal is physical-core separation. Keep each HA VM on a separate pair when possible. Also decide where its emulator and virtio/I/O threads will run. Pinning only vCPUs while leaving an overloaded emulator thread unrestricted may not solve the problem.

For the 2026-09-05 NAS test, the selected temporary placement was:

VM libvirt domain vCPU 0 vCPU 1 emulator Reason
fox uv059 CPU 0 CPU 1 CPUs 8-9 Separate physical cores; emulator on SMT siblings
anchor uv060 CPU 2 CPU 3 CPUs 10-11 Separate physical cores from fox

Table 1. Example HA VM CPU placement used for the KitsNet test.

This arrangement improves determinism for the two HA guests, but it does not fully reserve CPUs 0-3 because other guests were still allowed on the full 0-15 set. Full isolation would require a broader, host-wide capacity plan.

4.2 Apply live and persistent placement[edit | edit source]

The following commands change the running domain and persistent libvirt definition. Confirm domain names and topology first. Do not apply this pattern to production storage attach/detach operations.

for i in {1..10}; do echo; done
if bash <<'BASH'
set -Eeuo pipefail
test "$(id -un)" = psmode
ssh -o BatchMode=yes psmode@wort.lan.kitsnet.us 'bash -s' <<'REMOTE'
set -Eeuo pipefail
for domain in uv059 uv060; do
    sudo -n virsh dominfo "$domain" >/dev/null
done

sudo -n virsh vcpupin uv059 0 0 --live --config
sudo -n virsh vcpupin uv059 1 1 --live --config
sudo -n virsh emulatorpin uv059 8-9 --live --config

sudo -n virsh vcpupin uv060 0 2 --live --config
sudo -n virsh vcpupin uv060 1 3 --live --config
sudo -n virsh emulatorpin uv060 10-11 --live --config

for domain in uv059 uv060; do
    echo "===== $domain ====="
    sudo -n virsh vcpupin "$domain"
    sudo -n virsh emulatorpin "$domain"
done
REMOTE
BASH
then
    echo "cpu_placement_return_code=0"
else
    echo "cpu_placement_return_code=$?"
fi
echo "The interactive shell remains open."

The resulting persistent XML should contain entries equivalent to:

<vcpu placement='static'>2</vcpu>
<vcpupin vcpu='0' cpuset='0'/>
<vcpupin vcpu='1' cpuset='1'/>
<emulatorpin cpuset='8-9'/>

Do not hand-edit generated libvirt XML while the domain is running. Use the approved libvirt management command, then verify the resulting XML and live placement.

4.3 Rollback[edit | edit source]

Rollback is done by restoring an unrestricted CPU set for the affected domain. Use the exact prior placement recorded in the change log when possible. The following example removes the tested pinning from the two NAS guests while preserving the domain definitions:

for i in {1..10}; do echo; done
if bash <<'BASH'
set -Eeuo pipefail
test "$(id -un)" = psmode
ssh -o BatchMode=yes psmode@wort.lan.kitsnet.us 'bash -s' <<'REMOTE'
set -Eeuo pipefail
sudo -n virsh vcpupin uv059 0 0-15 --live --config
sudo -n virsh vcpupin uv059 1 0-15 --live --config
sudo -n virsh emulatorpin uv059 0-15 --live --config
sudo -n virsh vcpupin uv060 0 0-15 --live --config
sudo -n virsh vcpupin uv060 1 0-15 --live --config
sudo -n virsh emulatorpin uv060 0-15 --live --config
REMOTE
BASH
then
    echo "rollback=PASS"
else
    echo "rollback=$?"
fi
echo "The interactive shell remains open."

4.4 Resource and migration considerations[edit | edit source]

Pinning must be reviewed against live migration, host maintenance, and recovery. A CPU set valid on one hypervisor may not exist on another. If a guest can migrate, define an equivalent placement on every destination or remove pinning before migration. Do not enable host-wide isolcpus or real-time scheduling solely from one incident without a capacity and recovery review.

5 Evaluating effectiveness[edit | edit source]

Evaluate both latency and service correctness. A successful change should show a lower rate of vCPU delay accumulation, fewer timer-expiration events, and no new failover or storage errors under a comparable workload.

5.1 Immediate validation[edit | edit source]

Verify the live and persistent settings and compare two domstats samples over 30-60 seconds. Calculate:

Also check guest %steal, load, and service logs during the same interval. The post-pinning NAS sample showed only approximately 0.1 seconds of new delay per vCPU over 30 seconds while idle, and no new fox Keepalived timer-expiration messages.

for i in {1..10}; do echo; done
if bash <<'BASH'
set -Eeuo pipefail
test "$(id -un)" = psmode
ssh -o BatchMode=yes psmode@wort.lan.kitsnet.us 'bash -s' <<'REMOTE'
set -Eeuo pipefail
echo "===== SAMPLE 1 ====="
sudo -n virsh domstats uv059 uv060 --vcpu | \
    grep -E '^(Domain|  vcpu\.[01]\.(time|delay))' || true
sleep 30
echo "===== SAMPLE 2 ====="
sudo -n virsh domstats uv059 uv060 --vcpu | \
    grep -E '^(Domain|  vcpu\.[01]\.(time|delay))' || true
REMOTE
echo "===== KEEPALIVED TIMER EVENTS ====="
for peer in fox anchor; do
    echo "===== $peer ====="
    ssh -n -o BatchMode=yes "psmode@$peer.lan.kitsnet.us" \
        "sudo -n journalctl -u keepalived --since '10 minutes ago' --no-pager | grep -E 'timer expired|Entering|Leaving' || true"
done
BASH
then
    echo "post_pin_validation=PASS"
else
    echo "post_pin_validation=$?"
fi
echo "The interactive shell remains open."

5.2 Controlled workload validation[edit | edit source]

Repeat the original workload or a deliberately bounded test, recording the exact image, workload, duration, storage path, and host. For HA guests, test failover separately from application benchmarking. A benchmark can expose a scheduling problem, but a failover test must verify exclusive storage ownership, VIP movement, mount state, and service transition logs.

In the 2026-09-05 test, the NAS sequence was:

  1. Fox was primary with the VIP and /dkr/data1 mounted.
  2. Fox Keepalived was stopped after a completed backup.
  3. Anchor acquired the production LV, mounted /dev/vdb1, started NFS, and claimed the VIP in approximately six seconds.
  4. Anchor was stopped; fox reacquired the LV and VIP immediately on the observed poll.
  5. Anchor was restarted and rejoined as BACKUP with no VIP and no production mount.

That result validates HA transition behavior after pinning and transition serialization; it does not establish a benchmark performance improvement by itself.

5.3 Failure interpretation[edit | edit source]

  • If guest %steal and libvirt delay fall, but timer failures remain, inspect guest scheduling, Keepalived priority, network loss, and transition scripts.
  • If delay remains high after pinning, inspect host contention, emulator/I/O-thread placement, storage stalls, and host power-management behavior.
  • If delay improves but application latency worsens, the pinned cores may be contended by other guests or the workload may have changed.
  • If HA state or storage ownership fails, stop benchmarking and preserve fencing safeguards. Do not suppress alerts to make the test appear successful.

6 Monitoring and alerting[edit | edit source]

At minimum, monitor:

  • Keepalived state and timer expired messages;
  • VIP presence on exactly one node;
  • production filesystem source and mount state;
  • exclusive production-LV ownership through the NAS fencing workflow;
  • vCPU delay-rate samples on the hypervisor;
  • guest steal time and host CPU pressure;
  • ConsoleWorks event codes.

KitsNet ConsoleWorks mappings for NAS failover logs are:

Syslog PRI Facility/severity Intended event class
<138> local1.crit Critical
<139> local1.err Error
<140> local1.warning Warning
<141> local1.notice Notice

Table 2. KitsNet NAS ConsoleWorks local1 event-code mapping.

The nas-failover logger must specify local1 explicitly. A bare logger -t nas-failover defaults to the user facility and produces <13> for notice, which bypasses the intended local1 scans.

7 Tracking table[edit | edit source]

Use this table as the living record. Update it after each baseline, placement change, rollback, and validation window.

Candidate Hypervisor/domain vCPUs Current placement Pinning status Evidence/trigger Last validation Next review
fox wort / uv059 2 vCPU 0→0, vCPU 1→1; emulator→8-9 Applied Keepalived timer delays; elevated cumulative vCPU delay 2026-09-05 controlled failover passed After next HA observation window
anchor wort / uv060 2 vCPU 0→2, vCPU 1→3; emulator→10-11 Applied Keepalived timer delays; higher cumulative vCPU delay 2026-09-05 controlled failover and rejoin passed After next HA observation window
wrk1 Docker Swarm worker; host/domain to be recorded To be recorded Unpinned Not required Benchmark workload only; pinning would alter methodology No pinning test performed Only for a separately controlled benchmark
wrk2 Docker Swarm worker; host/domain to be recorded To be recorded Unpinned Not required Benchmark workload only; pinning would alter methodology No pinning test performed Only for a separately controlled benchmark

Table 3. Initial KitsNet CPU-pinning candidate register.

For each additional candidate, record the host CPU topology, vCPU and emulator placement, baseline delay rate, workload, guest steal percentage, host pressure, change timestamp, rollback point, and post-change result. A candidate remains Applied only while the measured benefit persists and the placement is compatible with migration and recovery.

8 Change record for the 2026-09-05 NAS incident[edit | edit source]

The following facts are verified from the recorded command output:

  • Hypervisor: wort; CPU: AMD Ryzen 7 5800X, eight physical cores and sixteen logical CPUs.
  • Domains: uv059=fox and uv060=anchor, each with two vCPUs.
  • Initial placement: all vCPUs and emulator threads allowed on CPUs 0-15.
  • Guest load was near zero, while cumulative libvirt vCPU delay was significant, especially on anchor.
  • Persistent live-and-config placement was applied as shown in Table 1.
  • The NAS failover test passed in both directions. Fox ended primary; anchor ended active BACKUP.
  • SELinux hostname permissions were narrowed to keepalived_t and hostname_exec_t; no new AVCs appeared after installation.
  • NAS transition scripts were updated to use local1 logging, correct storage-operation return status, and a host-local transition lock.

These changes explain why pinning was justified for fox and anchor. They do not imply that every KitsNet VM requires pinning.

9 Operational checklist[edit | edit source]

  1. Establish a time-bounded guest and host baseline.
  2. Map physical cores and SMT siblings.
  3. Inventory all domain vCPUs and current placement.
  4. Correlate delay, steal, timer, storage, and network evidence.
  5. Document the candidate, selected CPU sets, and rollback.
  6. Apply live and persistent placement through approved libvirt management.
  7. Verify placement and syntax/configuration state.
  8. Compare delay-rate and service behavior under a comparable workload.
  9. Test migration, failover, and recovery implications.
  10. Update Table 3 and retain the change backup/checksum record.

Do not declare pinning effective solely because a single command succeeded. It is effective only when the relevant latency or correctness symptom improves without creating a capacity, migration, or recovery regression.