Kernel Execution

PXL provides two host APIs for running kernels. Both use the same task and argument model, but they differ in how much setup the application manages.

API Kernel selection Resource setup Completion
Launcher Typed C++ kernel symbol Automatic run() or runAsync()
Map Loaded Function or function name Explicit Context, Job, and Map execute() and synchronize()

See Programming Objects for how Launcher fits into the PXL object model.


Launcher Execution

Launcher is the typed path used by single-source pxcc programs. The host passes the kernel’s own arguments to execute(), sets the launch task count with .tasks(), and then runs the request.

__pxl_kernel__ void copy(mu::input_ptr<int> input,
                         mu::output_ptr<int> output)
{
    const auto index = mu::getTaskIdx();
    output[index] = input[index];
}

auto result = pxl::Launcher()
                  .execute<copy>(input, output)
                  .tasks(taskCount)
                  .run();

The no-count overload returns a pending builder. It accepts optional settings for the current request, followed by the required tasks(), but exposes no run() / runAsync() or next sequential / parallel kernel operation. Attempting to run or continue first is therefore a C++ compile error. Discarding the pending request produces a [[nodiscard]] diagnostic and executes nothing.

The positional form execute<kernel>(taskCount, args...).run() is still supported. Launcher distinguishes it from execute<kernel>(args...).tasks(taskCount) by the number of parameters in the kernel signature. In either form, pass every kernel argument explicitly; Launcher does not fill in C++ default arguments. Use a non-variadic kernel function, or a function template with explicit template arguments. New code should use .tasks(taskCount) to keep launch settings separate from kernel arguments.

Sequential and parallel stages

stage<kernel>(args...) adds a sequential kernel. Each kernel entry needs .tasks(taskCount) before execution or the next kernel. Optional settings such as batchSize() may appear before or after tasks(). Move a named stage builder explicitly to continue its chain:

auto stages = pxl::Launcher().stage<first>(input).tasks(taskCount);
auto result = std::move(stages).stage<second>(output).tasks(taskCount).run();

Moving the builder transfers its accumulated stages. Continue with the returned builder; do not reuse the moved-from builder to execute the old stages.

To run kernels in parallel, call stage() without a kernel or arguments to open a StageGroupBuilder. Only that group exposes add<kernel>(); it is not available directly after stage<kernel>().

The two copies below run in one parallel group. The final copy starts after both have completed. Use separate output buffers for independent parallel work.

auto result = pxl::Launcher()
                  .stage()
                      .add<copy>(firstInput, firstOutput).tasks(firstTaskCount)
                      .add<copy>(secondInput, secondOutput).tasks(secondTaskCount)
                  .stage<copy>(firstOutput, finalOutput).tasks(firstTaskCount)
                  .run();

A group created from a temporary or moved stage builder keeps its preceding stages alive. It can be stored and executed later. To add another kernel to a named group, move the group explicitly, just as with a named stage builder:

auto group = pxl::Launcher()
                 .stage<copy>(input, intermediate).tasks(taskCount)
                 .stage();
auto result = std::move(group)
                  .add<copy>(intermediate, output).tasks(taskCount)
                  .run();

run() waits and returns a LaunchResult; runAsync() returns a future. Keep argument buffers alive until execution completes. Launcher copies NDArray descriptors, so those descriptor objects may go out of scope after the request is built. For a complete buildable program, see Hello Sort (Heterogeneous Programming).

Directional pointer arguments

Launcher derives a pointer buffer’s synchronization direction from the kernel parameter type:

Kernel parameter Before the kernel After the kernel
mu::input_ptr<T> Synchronize host contents to the device. Skip device-to-host synchronization and host-cache invalidation.
mu::output_ptr<T> Skip ordinary host-to-device synchronization. Synchronize device results to the host.
T* Synchronize in both directions, preserving the legacy default. Synchronize in both directions, preserving the legacy default.

The host still passes the original typed pointer or pxl::NDArray; the directional type appears only in the kernel signature. input_ptr<T> exposes const element access. For output_ptr<T>, previous contents are discardable, so the kernel must write every byte that the host will observe. Unwritten bytes are not guaranteed to retain their previous host values. On copying backends such as AdaptiveIoMem, the post-launch copy can overwrite those bytes with stale or unspecified device contents. For partial output, pass an NDArray that bounds the output region; use an unmarked, bidirectional parameter when previous contents must be preserved. Directional parameters require a typed pointer, so cast an untyped void* allocation to the matching element pointer before passing it.

An NDArray may be passed to a directional pointer parameter. PXL slices it and passes each slice’s data() pointer. A kernel parameter declared as mu::NDArray<T> remains bidirectional because NDArray describes shape and slicing, not access direction. A raw pointer covers the managed allocation from that pointer, while an NDArray limits synchronization to the view size.

NDArray descriptors may be mutable, const, or temporary; Launcher copies them. For directional kernel parameters, element constness is checked: NDArray<const T> can supply an input but cannot be passed to output_ptr<T>. This does not add constness validation to unmarked legacy parameters; their accesses must still match the underlying buffer’s mutability.

Declare directional kernel parameters by value, as in mu::input_ptr<T> and mu::output_ptr<T>. Reference forms such as const mu::input_ptr<T>& are not supported and are rejected at compile time. This restriction is independent of whether the caller’s NDArray descriptor is const or temporary.

CxlMem output buffers may still be flushed before launch. On BI-off paths this prevents a delayed dirty host writeback from overwriting the kernel’s result; BI-enabled paths also retain their existing pre-launch compatibility flush. An output direction skips ordinary uploads, not all cache maintenance.

AdaptiveIoMem and BI-off CxlMem use Map-managed synchronization, including for unmarked kernels. Map’s preload and exclusive-access checks apply to these arguments. If a launch reports that Map rejected a device-memory argument, check these requirements:

  • InfiniteMemory buffers on these paths must be preloaded. For sub-chunk allocations use AllocFlag::Preload when allocating; explicit pxl::preloadMemory() requires the pointer and byte size to be aligned to the device’s v2pEntrySize.
  • Buffers owned exclusively by another Job cannot be used by Launcher’s new Job. Launch them through a Map belonging to the owning Job.

Map Execution

A Map represents a reusable, explicitly configured kernel launch. It binds a loaded Function to a task count and a set of arguments.

auto map = job->buildMap("my_kernel", taskCount);
map->execute(arg0, arg1, arg2);
map->synchronize();

Map does not infer synchronization direction from the kernel signature. Register buffers with the existing setInput() or setOutput() API before execute(); register the same buffer with both when the kernel reads its previous contents and writes a result. Register every output buffer explicitly: on AdaptiveIoMem, once any input or output is registered, other unregistered buffers are treated as input-only and are not copied back. On devices without Back-Invalidate, CxlMem output cache synchronization also requires setOutput() registration. BI-enabled devices provide hardware-coherent output visibility.


Task Distribution

Two numbers control how a launch is sized:

  • taskCount — the total number of times your kernel will be invoked. Set with Launcher .tasks(taskCount) or when a Map is built with buildMap(func, taskCount). The maximum is pxl::MaxTaskCount ((1 << 20) - 1).
  • batchSize — the number of tasks each MU core processes back-to-back. Default: 16. Set with Launcher .batchSize(batchSize) or Map::setBatchSize(batchSize).

Launcher rejects a taskCount of zero before dispatch. buildMap() accepts zero: the Map is built, and execute() completes successfully without running the kernel. Check a computed task count in the application if a no-op launch would indicate an error.

The library divides taskCount into batches and gives one batch to each MU core:

batchCount   = ceil(taskCount / batchSize)
active cores ≤ min(batchCount, available MU cores in the Job)

The active core count is an upper bound — locality mode (see below) and other dispatch constraints may leave some cores idle even when batches are available.

The unit of distribution is the MU core, not the Sub. numSub only sets the size of the available core pool. Inside the pool, work is handed out per-core.

For the example below, assume the device exposes 128 MU cores per Sub, so a Job with 4 Subs has a pool of 512 cores:

taskCount batchSize batchCount Active cores Kernel invocations
1024 16 64 64 1024
1024 1 1024 512 (pool cap) 1024
100 1 100 100 100
8192 8 1024 512 (pool cap) 8192

The kernel is invoked exactly taskCount times regardless of batchSize. A larger batchSize reduces per-core entry overhead at the cost of fewer cores being active.

Reading the task index inside a kernel

Each kernel invocation gets a unique logical index in the range [0, taskCount-1]:

#include "mu/mu.hpp"

void my_kernel(int* data, int size) {
    auto idx = mu::getTaskIdx();        // [0, taskCount-1]
    auto total = mu::getTaskCount();    // == configured taskCount
    // ...
}

For full kernel-side details (header, registration, parameter limits), see Kernel Programming Guide.


Argument Types

Launcher and Map::execute(args...) use the following argument categories, with different support for Constant values:

Kind When to use How PXL treats it
Constant Launcher: arithmetic scalars such as int and float; bool and enums are not supported. Map also accepts trivially-copyable structs. Broadcast — every kernel invocation sees the same value.
Device pointer A device allocation from Context::memAlloc() or pxl::allocateMemory<T>(). Directional parameters require a pointer of the matching element type. Broadcast — every invocation sees the same pointer. The kernel typically uses mu::getTaskIdx() to compute its slice.
NDArray Typed, shaped view over a device buffer. Sliced — PXL hands each invocation its own NDArray view.

With Map::execute(), a struct of trivial POD members counts as a single Constant parameter and is broadcast to every invocation. Launcher does not currently accept structs as by-value arguments. Launcher retains legacy void* arguments for unmarked pointer parameters.

The pxcc/MU kernel entry supports at most 9 kernel parameters. Map’s internal MaxNumArguments = 10 is a separate transport limit; it does not extend the kernel entry’s limit.

Launch with NDArray

NDArray<T> carries shape and stride information. PXL slices the array along its leading dimension and gives each MU core its slice:

Launch with NDArray

auto data = ctx->memAlloc(testCount * rowSize * sizeof(int));
auto arr  = pxl::NDArray<int>(static_cast<int*>(data), {testCount, rowSize});

auto map = job->buildMap("sort_with_ndarray", testCount);
map->execute(arr);
map->synchronize();

Launch with a device pointer

When you pass a raw device pointer or scalar, every invocation receives the same value. The kernel uses mu::getTaskIdx() to address its slice manually:

Launch with Pointer

auto map = job->buildMap("sort_with_ptr", testCount);
map->execute(static_cast<int*>(data), rowSize);

Locality Mode

Launcher .locality(mode) and Map::setLocalityMode(mode) control how tasks are distributed across the Subs and Clusters used by the launch.

  • CompactMode (default) — fill within a Sub first (cluster-first inside a Sub). Favors L2 reuse when neighbouring tasks share data.
  • SpreadMode — distribute across Subs first. Favors aggregate memory bandwidth when each task is independent and bandwidth-bound.

The tables below assume 4 Clusters per Sub for illustration:

numSub = 1:

Cluster CL0 CL1 CL2 CL3
Spread 0, 4, 8, 12 1, 5, 9, 13 2, 6, 10, 14 3, 7, 11, 15
Compact 0, 1, 2, 3 4, 5, 6, 7 8, 9, 10, 11 12, 13, 14, 15

numSub = 2:

  SUB0 / CL0 SUB0 / CL1 SUB0 / CL2 SUB0 / CL3 SUB1 / CL0 SUB1 / CL1 SUB1 / CL2 SUB1 / CL3
Spread 0, 8 2, 10 4, 12 6, 14 1, 9 3, 11 5, 13 7, 15
Compact 0, 1 2, 3 4, 5 6, 7 8, 9 10, 11 12, 13 14, 15
auto result = pxl::Launcher()
                  .execute<kernel>(data)
                  .tasks(taskCount)
                  .locality(pxl::LocalityMode::SpreadMode)
                  .run();

map->setLocalityMode(pxl::LocalityMode::SpreadMode);

The default CompactMode fits most workloads. Switch to SpreadMode if profiling shows the kernel is bandwidth-bound and would benefit from spreading evenly across Subs.


Map Execution Lifecycle

A single Map::execute() call moves the Map through a sequence of states. The current state is observable with Map::getExecuteStatus().

stateDiagram-v2
    direction LR
    [*] --> Idle
    Idle --> HostInit: execute()
    HostInit --> DeviceInit
    DeviceInit --> Request
    Request --> Waiting
    Waiting --> DeviceFinalize
    DeviceFinalize --> HostFinalize
    HostFinalize --> Completed
    Waiting --> Fail: device error
    Waiting --> Cancelled: cancel()
    Cancelled --> HostInit: execute()
    Completed --> HostInit: execute()

Completed and Cancelled are reusable — calling execute() on a Map in either state starts a new run, and so does the initial Idle. Fail is terminal: the Map should not be re-executed after a device error.

A successful run reports Completed. Earlier releases reported Idle for both “has not executed yet” and “finished cleanly”, so code that detects completion with getExecuteStatus() == Idle never sees a finished run on this release and a polling loop written that way does not exit. Compare against Completed, or use the return value of synchronize().

State What happens
HostInit Host-side argument setup and host-to-device data sync.
DeviceInit Per-execute device-side initialization.
Request Tasks are dispatched to the device.
Waiting Host waits for device-side completion.
DeviceFinalize Device-side cleanup.
HostFinalize Device-to-host data sync, callback dispatch.
Idle The Map has not executed yet.
Completed Run finished cleanly. The Map is reusable.
Fail A device or runtime error occurred.
Cancelled A cancel() request was honored.

The Map::getProgress() helper returns target / issued / done packet counts for live progress reporting.


Map Synchronization, Cancellation, and Callbacks

Map::execute() is non-blocking — it enqueues work onto the Map’s stream and returns. It returns pxl::Result::Failure if the underlying stream is torn down or its consumer thread dies; check this when robustness against runtime tear-down matters:

if (map->execute(arg0, arg1) != pxl::Result::Success) {
    // stream torn down — abort or recover
}

Use one of the following to wait for or interrupt a run:

  • synchronize() — block until the Map reaches Completed, Cancelled, or Fail.
  • cancel() — soft stop. Skips not-yet-issued batches and waits for in-flight tasks to drain. Output is not synced back. The Map is reusable afterward.
  • Callbacks — register before calling execute():
map->setCompletionCallback([](void* arg) { /* success */ }, nullptr);
map->setMessageCallback   ([](void* msg, void* arg) { /* device message */ }, nullptr);
map->setErrorCallback     ([](void* arg) { /* failure */ }, nullptr);

To pipeline multiple kernel launches, issue several execute() calls in sequence and place a single synchronize() at the end.


→ Related: Programming Objects, Streams