Troubleshooting
This section provides solutions to common problems you may encounter.
Index
- Troubleshooting
- Index
- Illegal Instruction Error with SIMD Access in QEMU
- Could not access KVM kernel module: Permission denied
cxlPackage IssuedaxctlPackage Issuedaxctl reconfigure-devicefails withDevice or resource busyxcena_cliExecution Errors- Install skips the driver: no kernel headers
- Example Test Failures
- Collecting Troubleshooting Logs
- pxcc Compiler Issues
Illegal Instruction Error with SIMD Access in QEMU
Problem
When accessing CXL memory allocated by PXL’s memAlloc API using SIMD instructions within QEMU, an illegal instruction error may occur. This issue can also arise when accessing the memory using float* or double* pointers. This is a known issue in QEMU Issue #3075.
Solution
To resolve this issue, run QEMU in no-KVM mode. This can be done by adding the --no-kvm option when starting QEMU:
./run.sh --no-kvm
Note Disabling KVM may reduce guest OS performance.
Could not access KVM kernel module: Permission denied
Problem
In order to launch QEMU, this error occurs when the current user lacks the necessary permissions to access the KVM (Kernel-based Virtual Machine) kernel module.
Solution
- Add User to KVM Group.
sudo usermod -a -G kvm $USER - Log out and log back in (or open a new terminal session) to apply the group change.
- Run QEMU again.
cxl Package Issue
Problem
The cxl list command is not working.
Solution
- Ensure the environment is a Docker container:
- Check the
cxlpackage version:- If the version is incorrect, certain features may not be visible.
cxl version # Expected version: 72.1+
- If the version is incorrect, certain features may not be visible.
- If issues persist, try reinstalling Docker.
daxctl Package Issue
Problem
The daxctl list command is not working.
Solution
- Ensure the environment is a Docker container.
- Verify that the CXL device is attached as a
.Memdevice:lspci | grep CXL # Check the BDF (Bus Device Function), lspci -D -d 20a6: # or match on the XCENA vendor ID. lspci -vvs <BDF> | grep CXLCtl # Ensure "Mem+" is included.
daxctl reconfigure-device fails with Device or resource busy
Problem
sudo daxctl reconfigure-device --mode=devdax dax0.0 cannot offline the CXL memory.
Solution
The CXL memory was onlined as system-ram (check with free -h; the total includes the device capacity). Once ZONE_NORMAL pages land on it, it cannot be offlined without a reboot.
- Ensure
memhp_default_state=offlineis on the kernel cmdline (cat /proc/cmdline). - Every supported distro ships a udev rule that force-onlines hotplug memory even with that parameter set (Ubuntu
90-daxctl-device.rules, openSUSE80-hotplug-cpu-mem.rules, RHEL 940-redhat.rules) — runsudo ./scripts/setup_host.sh(shipped with the SDK) to override the rule per distro and install the boot-time conversion service. - Reboot, then re-run the reconfigure command (or let
xcena-devdax.servicedo it).
xcena_cli Execution Errors
Problem
The output is 0 in the xcena_cli num-device command.
Solution
- Ensure the required Python modules are installed:
pip list | grep -E 'click|pandas|pyvcd|serial|pexpect' - Check if Docker was started with the
--privilegedflag:ls /sys/fs/cgroup/ # If successful, Docker is running in privileged mode. - Confirm that the
mx_dmamodule is loaded into the kernel:lsmod | grep mx_dma - (Re-)Install and reload the
mx_dmamodule if necessary:docker cp xcena_sdk:/work/driver /tmp/mx_dma docker stop xcena_sdk docker rm xcena_sdk cd /tmp/mx_dma/mxdriver sudo ./install.sh reboot lsmod | grep mx_dma # If mx_dma is not loaded, reload the modules: sudo rmmod cxl_pmem cxl_acpi cxl_pci cxl_core mx_dma sudo insmod /lib/modules/5.15.0-43-generic/extra/mx_dma.ko sudo insmod /lib/modules/5.15.0-43-generic/extra/cxl_5.15/core/cxl_core.ko sudo insmod /lib/modules/5.15.0-43-generic/extra/cxl_5.15/cxl_pci.ko sudo insmod /lib/modules/5.15.0-43-generic/extra/cxl_5.15/cxl_acpi.ko sudo insmod /lib/modules/5.15.0-43-generic/extra/cxl_5.15/cxl_pmem.ko # Re-run a Docker container
Install skips the driver: no kernel headers
Problem
install.sh finishes, but the summary reports the driver step as skipped:
[WARN] Skipping driver: no kernel headers for 6.8.0-136-generic (/lib/modules/6.8.0-136-generic/build missing).
DKMS cannot build mx_dma without headers for the booted kernel. Every other SDK component installs normally, but the device is unusable until the module is built. install_dependencies.sh tries to install the matching headers first and prints its own warning when it cannot resolve them — usually because the booted kernel is older than what the enabled repositories carry.
Solution
- Install the headers for the booted kernel:
sudo apt-get install -y "linux-headers-$(uname -r)" # Ubuntu sudo zypper install -y "kernel-default-devel" # openSUSE Leap sudo dnf install -y "kernel-devel-$(uname -r)" # RHEL 9 - If that version is gone from the repositories, move to a kernel that still has headers, then reboot:
sudo apt-get install -y linux-generic && sudo reboot # Ubuntu sudo zypper -n update kernel-default && sudo reboot # openSUSE Leap sudo dnf -y update kernel && sudo reboot # RHEL 9 - Confirm
/lib/modules/$(uname -r)/buildexists, then build the module:sudo bash driver/mxdriver/install.sh lsmod | grep mx_dma
Example Test Failures
Problem
An example test fails with a core dump.
Steps to Troubleshoot
- Check if
num_deviceis greater than or equal to1:xcena_cli num-device- If
num_deviceis0, refer to thexcena_cliExecution Errors section.
- If
- Verify that
MSUB bitmapis non-zero:xcena_cli device-info 0- If it is
0x0, offloading cannot proceed.
- If it is
Collecting Troubleshooting Logs
If the issue persists after following the steps above, collect diagnostic logs using the troubleshooting.sh script and share them with the support team.
Run Log Collection Script
Download and run troubleshooting.sh as root to collect system and device diagnostic logs:
wget https://raw.githubusercontent.com/xcena-dev/public_sdk_release/refs/heads/main/scripts/troubleshooting.sh
sudo bash troubleshooting.sh
Note Run it as root. Without root,
dmesg,dmidecode,lspci -vv,acpidump,journalctland most of the CXL sysfs tree return nothing useful. The script re-executes itself undersudowhen it can; when it cannot, it still runs but marks the report_INCOMPLETEso the recipient can tell at a glance.
The report is written to the current directory as troubleshooting_report_YYYY-MM-DD-HH-MM-SS.log.
| Option | Effect |
|---|---|
| (none) | Collects everything needed for a diagnosis. A few very large, low-signal sources are summarised. |
--full | No summarising. Use only if XCENA support asks for it — the report grows several times larger, and the filename gains a _full tag. |
-h, --help | Usage. |
Optional Diagnostic Packages
When a tool is missing the script marks the item SKIP, or records the absence in place of the output, rather than aborting the run. The report is most useful with these installed:
sudo apt install cxl ndctl daxctl numactl pciutils dmidecode acpica-tools mokutil jq lsof ipmitool rasdaemon
Without acpica-tools the ACPI tables are still collected, as an undecoded hexdump instead of decoded ASL.
What It Collects
Every item records the command that produced it, or the source it read.
| # | Section | Contents |
|---|---|---|
| 1 | Host Validation | validate_host.sh, run from the same directory if present, otherwise fetched at a pinned revision |
| 2 | Host Platform & BIOS | BIOS, board and CPU via dmidecode, DRAM population, PCIe slot inventory with per-slot CXL capability, Secure Boot, clock sync |
| 3 | Software & Tool Versions | cxl / daxctl / ndctl / numactl / lspci versions, XCENA packages, modinfo mx_dma, MU toolchain |
| 4 | Kernel | /proc/cmdline, CONFIG_CXL_* build options, loaded modules, taint state, IOMMU, dmesg -T, current and previous boot journals |
| 5 | XCENA Runtime | mx_dma device nodes and messages, pxl_resourced status and journal, interrupt delivery, holders of /dev/dax*, memlock limits, SELinux/AppArmor |
| 6 | Memory & NUMA | /proc/iomem, NUMA topology and distances, HMAT bandwidth and latency, memory block zones, tiering, EDAC error counters |
| 7 | CXL Subsystem | Full cxl list topology, device health / partition / alert configuration, every CXL sysfs attribute value, CDAT, and the CEDT, SRAT, HMAT, SLIT, MCFG and HEST ACPI tables |
| 8 | DAX | daxctl regions and devices, sysfs attributes, devdax mode check |
| 9 | Device Firmware | xcena_cli num-device, device-info and fw-info per device |
| 10 | InfiniteMemory SMART | xcena_cli im get-smart -v for each device reported as InfiniteMemory |
| 11 | PCIe | Topology tree, physical slot of each CXL device, link speed and width from endpoint to root port, AER counters, ASPM policy, ACPI _OSC negotiation, GHES records, BMC/IPMI event log, lspci -vvv and config space |
| 12 | Summary | Automated triage — see below |
Example Output
XCENA Troubleshooting Report
yyyy-mm-dd hh:mm:ss KST
Privilege: root
Output: troubleshooting_report_yyyy-mm-dd-hh-mm-ss.log
[ 1] Host Validation
-> 1-1. validate_host.sh (local) [OK]
[ 2] Host Platform & BIOS
-> 2-1. System identity [OK]
-> 2-2. OS release [OK]
-> 2-3. Kernel (uname -a) [OK]
-> ...
[ 3] Software & Tool Versions
-> 3-1. Tool versions [OK]
-> 3-2. XCENA / CXL related packages [OK]
-> 3-3. mx_dma driver module info [OK]
-> ...
[ 4] Kernel
-> 4-1. Boot parameters [OK]
-> 4-2. Kernel build configuration (CXL/DAX/NUMA) [OK]
-> 4-3. IOMMU / DMA remapping [OK]
-> ...
[ 5] XCENA Runtime (driver / PXL daemon / access)
-> 5-1. mx_dma device nodes [OK]
-> 5-2. mx_dma kernel messages [OK]
-> 5-3. pxl_resourced service status [OK]
-> ...
[ 6] Memory & NUMA Topology
-> 6-1. /proc/iomem [OK]
-> 6-2. NUMA hardware summary [OK]
-> 6-3. NUMA statistics [OK]
-> ...
[ 7] CXL Subsystem
-> 7-1. cxl list (full topology, verbose) [OK]
-> 7-2. cxl list (including idle/disabled devices) [OK]
-> 7-3. cxl list (memdev health) [OK]
-> ...
[ 8] DAX
-> 8-1. daxctl list (regions + devices) [OK]
-> 8-2. daxctl list (including idle) [OK]
-> 8-3. DAX device nodes [OK]
-> ...
[ 9] Device Firmware (xcena_cli)
-> 9-1. xcena_cli num-device [OK]
-> 9-2. xcena_cli device-info (device 0) [OK]
-> 9-3. xcena_cli fw-info (device 0) [OK]
-> ...
[10] InfiniteMemory SMART (xcena_cli)
-> 10-1. xcena_cli im get-smart (device 0) [OK]
[11] PCIe
-> 11-1. PCI topology tree [OK]
-> 11-2. All PCI devices [OK]
-> 11-3. CXL/XCENA device discovery [OK]
-> ...
[12] Summary
[device chain]
OK XCENA/CXL PCI 0000:2a:00.0 slot: J2A16 - SLOT_C (CXL 2.0 capable)
OK CXL memdev mem0 fw=1.0.11 233.00 GiB
OK CXL region region0 mode=ram commit=1
OK DAX device dax0.0 mode=devdax target_node=6
[link]
NOTE PCIe link 0000:2a:00.0 32.0 GT/s PCIe x8 (device max 64.0 GT/s PCIe x8)
capped by upstream port 0000:29:02.0 (max 32.0 GT/s PCIe)
expected: the device outruns the platform, not a fault
... [errors], [kernel], [host-identifying data], [collection]
Done. collected OK=98 FAIL=0 SKIP=0 (20s)
Summary: 1 WARN, 3 NOTE — see the Summary section above
Report : /home/user/troubleshooting_report_yyyy-mm-dd-hh-mm-ss.log (1.7M)
Summary Section
The report ends with a triage summary, also printed to the terminal, covering the device chain (PCI to memdev to region to DAX to driver to daemon), link state, error sources and kernel configuration. Each line is one of:
OK— observed and unremarkableNOTE— worth a human’s eye, but a legitimate configurationWARN— likely to mislead, or to block something--— not determined: a tool was missing or the data was unavailable; never a judgement about the host
This is a triage aid, not a verdict. validate_host.sh is the script that passes or fails a host.
Privacy
Host-identifying fields are always masked. The report includes a table stating exactly what was masked and what was kept.
| Fields | |
|---|---|
| Masked | hostname, machine-id, system / chassis / board / CPU / DIMM serial numbers, UUIDs, asset tags, account names, MAC and IP addresses |
| Kept | BIOS vendor, version and date, system and baseboard model, PCI device inventory, physical slot labels, local filesystem sizes and mount points, CXL device serial number and firmware version |
If anything is left unmasked, the report is renamed _NOT_FULLY_MASKED.log so you can review it before sharing.
Note Masking cannot cover everything. Free-text kernel log lines may still contain identifying strings — an internal hostname printed by an application, a custom path, or an identifier in a format no rule recognises.
Reading the Report
The report is plain text, and each collected item is marked >>> <section>-<item>. <title>:
grep '^>>> ' troubleshooting_report_*.log # list the collected items
awk '/^>>> 7-3\./{p=1} /^>>> 7-4\./{p=0} p' troubleshooting_report_*.log # read one item
Share the generated .log file when reporting issues.
pxcc Compiler Issues
The sections below cover the experimental pxcc compiler.
Device C++ standard errors
Problem
The device compile fails with errors about C++20/C++23 features used in a kernel or device function, or about an unsupported -std value on the device side. Code that runs on the device (functions reachable from __pxl_kernel__) must be C++17 or earlier — the device backend does not yet support later standards. Host code is unaffected and may use C++17/20/23.
Solution
Keep device code within C++17, or set the device standard explicitly with -Xmu=. The host standard is independent and set the usual way:
pxcc++ -Xmu=-std=c++17 app.cpp -o app # device standard; host unaffected
set(CMAKE_CXX_STANDARD 20) # host may be C++20/23; device stays C++17 or lower
Kernel not found or not extracted
Problem
The build succeeds but a launch fails at runtime, or the device code is empty. This usually means the function was not annotated, so pxcc compiled it as ordinary host code instead of extracting it for the device.
Solution
Annotate the entry point with __pxl_kernel__, and make sure #include "mu/mu.hpp" is present. Confirm the extraction with -save-temps and check that the kernel appears in mu.{filename}.cpp:
pxcc++ -save-temps -c app.cpp -o app.o
grep sort_with_ptr mu.app.cpp
See the annotations table for __pxl_kernel__, __mu_device__, and __mu_shared__.
mu:: used in shared code
Problem
Errors that mu:: symbols are undeclared inside a __mu_shared__ function. __mu_shared__ code also compiles for the host, where mu:: APIs do not exist.
Solution
Move the mu::-using logic into a __mu_device__ helper, or call it only from __pxl_kernel__ / __mu_device__ code.
-lpxl not found at link
Problem
The link step fails with cannot find -lpxl or unresolved PXL symbols. pxcc adds -lpxl automatically, but the PXL library is not on the linker search path (for example, a non-standard SDK install location).
Solution
Point the linker at the PXL library directory.
pxcc++ app.o -L/usr/local/lib -o app
Verify the SDK is installed — see Install.
Host compiler not found
Problem
pxcc aborts early with Host compiler not found before any compilation runs. It could not resolve a host compiler from --host-compiler, the PXCC_HOST_CXX / PXCC_HOST_CC (or CXX / CC) environment variables, or a clang on PATH.
Solution
Point pxcc at a host compiler explicitly, or install one on PATH.
pxcc++ --host-compiler=/usr/bin/g++-12 -c app.cpp -o app.o
# or
PXCC_HOST_CXX=/usr/bin/g++-12 pxcc++ app.cpp -o app
No device code found
Problem
The build prints No device code found and produces a host-only binary. This is informational, not an error — the source has no __pxl_kernel__ (or __mu_shared__) code, so pxcc compiled it as an ordinary host translation unit.
Solution
Nothing is needed if the file is intentionally host-only. If you expected a kernel, confirm it is annotated with __pxl_kernel__ — see Kernel not found or not extracted.
Inspecting intermediate output
Problem
pxcc behaves unexpectedly and you need to see what it generated for the host and device sides.
Solution
Build with -save-temps and read the generated files:
| File | Use it to check |
|---|---|
mu.{filename}.cpp | Did the kernel get extracted? Is the body what you expect? |
host.{filename}.cpp | How the host side was transformed. |
mu_kernel.mubin | Device binary — confirm the kernel compiled successfully. |
mu_kernel.o | Embedding object — confirm it was linked into the host binary. |
If the extracted device code looks correct but the binary still misbehaves, see pxcc Compiler for the full diagnostics reference.