API

Kernel language

KernelAbstractions.@kernel — Macro
@kernel function f(args) end

Takes a function definition and generates a Kernel constructor from it. The enclosed function is allowed to contain kernel language constructs. In order to call it the kernel has first to be specialized on the backend and then invoked on the arguments.

Kernel language

Kernel constructor

After defining a kernel function f, call f(backend[, workgroupsize[, ndrange]]) to obtain a Kernel specialized for that backend. Workgroup size and ndrange can be fixed at construction time (enabling size-specific compile-time optimizations and fewer runtime checks, at the cost of recompilation when the sizes change) or supplied at launch:

f(backend)                    # dynamic workgroup size and ndrange
f(backend, 64)                # static workgroup size of 64
f(backend, 64, 1024)          # static workgroup size and ndrange
f(backend, 64, (128, 128))    # multi-dimensional ndrange

Example

using KernelAbstractions

@kernel function vecadd(A, @Const(B))
    I = @index(Global)
    @inbounds A[I] += B[I]
end

dev = CPU()
A = ones(1024)
B = rand(1024)
vecadd(dev, 64)(A, B, ndrange=length(A))
synchronize(dev)
source
@kernel config function f(args) end

This allows for the following configurations:

  1. cpu={true, false}: Deprecated in KernelAbstractions 0.11; this option is ignored.
  2. inbounds={false, true}: Enables a forced @inbounds macro around the function definition in the case the user is using too many @inbounds already in their kernel. Note that this can lead to incorrect results, crashes, etc and is fundamentally unsafe. Be careful!
  3. unsafe_indices={false, true}: Disables the implicit validation of indices, users must avoid @index(Global).
  4. generated={false, true}: Turns the kernel into a generated function, see Generated below.
Warning

This is an experimental feature.

Generated

With generated=true the kernel body is treated as a quoted expression, so $ interpolation is available and where-parameters are bound to their values, e.g. to unroll a loop $N times:

@kernel generated = true function kernel_unroll!(a, ::Val{N}) where {N}
    @unroll $N for i in 1:5
        @inbounds a[i] = i * $N
    end
end

This is meant for macros that need a literal, such as @unroll $N, Base.Cartesian.@nexprs $N or @ntuple $N; plain where-parameters are compile-time constants in every kernel already. Configuration parameters must therefore be passed as types (::Val{N}) to be usable inside $. Inside $(...) the argument names refer to the types of the arguments, not their values, as in any generated function, and the body cannot contain closures, comprehensions or generators (x -> ..., do blocks, [f(i) for i in ...]); use the Cartesian macros above instead.

source
KernelAbstractions.@Const — Macro
@Const(A)

@Const is an argument annotiation that asserts that the memory reference by A is both not written to as part of the kernel and that it does not alias any other memory in the kernel.

The kernel receives such an argument with the backend's read-only array type, so a type annotation on it has to accept that type, e.g., @Const(A::AbstractArray{T}). Which arguments are read-only is determined by the kernel method the arguments select before that, so kernel methods should not dispatch on the read-only type.

Danger

Violating those constraints will lead to arbitrary behaviour.

As an example given a kernel signature kernel(A, @Const(B)), you are not allowed to call the kernel with kernel(A, A) or kernel(A, view(A, :)).

source
KernelAbstractions.@index — Macro
@index

The @index macro can be used to give you the index of a workitem within a kernel function. It supports both the production of a linear index or a cartesian index. A cartesian index is a general N-dimensional index that is derived from the iteration space.

Index granularity

  • Global: Used to access global memory.
  • Group: The index of the workgroup.
  • Local: The within workgroup index.

Index kind

  • Linear: Produces an Int64 that can be used to linearly index into memory.
  • Cartesian: Produces a CartesianIndex{N} that can be used to index into memory.
  • NTuple: Produces a NTuple{N} that can be used to index into memory.

If the index kind is not provided it defaults to Linear, this is subject to change.

Examples

@index(Global, Linear)
@index(Global, Cartesian)
@index(Local, Cartesian)
@index(Group, Linear)
@index(Local, NTuple)
@index(Global)
source
KernelAbstractions.@private — Macro
@private T dims

Declare storage that is private to each work-item. It is preserved across @synchronize statements.

This returns a statically sized StaticArraysCore.StaticArray (a KernelAbstractions.PrivateArray) in stack storage, which supports indexing and in-place operations on non-overlapping regions. Load StaticArrays for static-array arithmetic, slicing and unrolled whole-array reductions. Without it, these fall back to generic AbstractArray methods: operations that return a new array fail to compile, and reductions may spill to local memory or fail to compile, depending on the back-end and size.

Assignment shares storage, and each @private declaration reuses its storage across loop iterations. copy, similar and other operations that allocate are not guaranteed to work, and neither is assigning between overlapping views of the same array.

source
@private var = expr

Equivalent to @uniform var = expr. Ordinary variables are already private to each work-item, and keep their value across @synchronize statements.

source
KernelAbstractions.@synchronize — Macro
@synchronize()

After a @synchronize statement all read and writes to global and local memory from each thread in the workgroup are visible in from all other threads in the workgroup.

@synchronize must be reached by all work-items of a workgroup. To ensure that padding work-items outside of the ndrange reach it too, @kernel treats it specially, so it has to appear directly in the kernel body, and not in a function called by the kernel (unless the kernel uses unsafe_indices=true). Control flow containing it must be uniform across the workgroup, see @uniform.

source
@synchronize(cond)

After a @synchronize statement all read and writes to global and local memory from each thread in the workgroup are visible in from all other threads in the workgroup. cond is not allowed to have any visible sideffects.

Platform differences

  • GPU: This synchronization will only occur if the cond evaluates.
  • CPU: This synchronization will always occur.
Warning

This variant of the @synchronize macro violates the requirement that @synchronize must be encountered by all workitems of a work-group executing the kernel or by none at all. Since v0.9.34 this version of the macro is deprecated and lowers to @synchronize()

source
KernelAbstractions.@print — Macro
@print(items...)

This is a unified print statement.

Platform differences

  • GPU: This will reorganize the items to print via @cuprintf
  • CPU: This will call print(items...)
source
KernelAbstractions.@uniform — Macro
@uniform expr

Evaluate expr on every work-item of the workgroup, including padding work-items that fall outside of the ndrange. This only applies to @uniform statements at the top level of the kernel, or directly in control flow that contains a @synchronize. Elsewhere, @uniform has no effect.

When the ndrange is not a multiple of the workgroup size, @kernel only runs the kernel body on work-items inside the ndrange. Padding work-items do still need to reach every @synchronize, so control flow that contains a @synchronize runs on all work-items. Values used by such control flow, like the bounds of a loop, must therefore be computed with @uniform:

@kernel function f(A, n)
    i = @index(Global)
    @uniform iterations = 2n
    for j in 1:iterations
        A[i] += j
        @synchronize()
    end
end

@uniform statements are hoisted to the start of the code between two @synchronize statements. expr must be safe to evaluate on padding work-items, so it should not depend on the work-item's index.

source
KernelAbstractions.@groupsize — Macro
@groupsize()

Query the workgroupsize on the backend. This function returns a tuple corresponding to kernel configuration. In order to get the total size you can use prod(@groupsize()).

source

Host language

Note

The Backend type hierarchy and most of the host-side management functions below (get_backend, allocate, synchronize, device selection, …) are defined in the KernelInterface sibling package and re-exported by KernelAbstractions, so KernelAbstractions.allocate and KernelInterface.allocate are the same function. User code can keep calling them through KernelAbstractions as before. Note that this only applies to KernelAbstractions 0.10 and later: KernelAbstractions 0.9 defines its own versions of these functions, which are distinct from the KernelInterface ones.

Backends and arrays

KernelInterface.Backend — Type
Backend

Abstract supertype for all KernelInterface backends. Backends subtype it directly.

A backend value identifies a backend and its configuration (e.g. compiler options). The device and the queue that operations use are task-local: each task selects its active device with device!. Host-side queries answer for the calling task's active device, and work is queued on the calling task's queue of that device. Use get_backend to obtain the backend of an array and allocate to create storage on a backend.

Example

backend = get_backend(A)
kernel = my_kernel(backend, 256)
kernel(A, ndrange=length(A))
synchronize(backend)
source
KernelAbstractions.CPU — Type
CPU

Type alias for POCLBackend, the CPU execution backend.

Construct with CPU() (equivalent to POCLBackend()). Kernels run on the host via POCL/OpenCL using the same programming model as GPU backends, which is useful for debugging and for running kernel code without a GPU.

Example

A = ones(Float32, 1024)
mul2_kernel(CPU(), 64)(A, ndrange=length(A))
synchronize(CPU())

Threads

Kernels run on as many threads as Julia's default thread pool has (julia -t N), up to the number of hardware threads. These are POCL's own threads, so launching a kernel doesn't occupy Julia's. To use a different number, set the JULIA_KA_CPU_THREADS environment variable before the backend is first used, e.g., JULIA_KA_CPU_THREADS=8 julia -t1. POCL's own variables (e.g., POCL_CPU_MAX_CU_COUNT) are respected too, and can also raise the number above the number of hardware threads, but they also affect other users of POCL, like OpenCL.jl. The device reports the number as its compute units: KernelAbstractions.POCL.device().max_compute_units.

Atomics

8- and 16-bit atomic operations (e.g., on Int8, UInt16 or Float16 arrays) are performed on the aligned 32-bit word that contains the value, which must not overlap another allocation or memory that is modified independently. For arrays backed by memory that Julia allocated (e.g., by allocate, Array, zeros, resize!, copy or similar), that word stays within the same allocation, because of how Julia's allocator is implemented. For arrays that wrap other memory, e.g., with unsafe_wrap, this is up to the user: make that memory extend to the 4-byte boundaries around the array, e.g., by allocating a multiple of 4 bytes at a 4-byte aligned address.

This includes adjacent elements of the same array: while an 8- or 16-bit element is updated atomically, the other elements in its aligned 32-bit word must not be modified by plain stores concurrently, only atomically.

source
KernelInterface.get_backend — Function
get_backend(A::AbstractArray)::Backend

Get a Backend instance suitable for array A.

Note

Backend implementations must provide get_backend for their custom array type. It should be the same as the return type of allocate

source
Adapt.adapt_storage — Method
adapt(backend::Backend, x)

Convert x such that its array storage lives on backend. This is an extension of Adapt.jl, and lets code move data to a backend without knowing the backend's array type:

using Adapt
x = adapt(CUDABackend(), rand(Float32, 8))  # a CuArray
y = adapt(CPU(), x)                         # an Array again
Note

Backend implementations must implement Adapt.adapt_storage(::NewBackend, x). Adapt.jl's fallback is the identity, so a backend that omits this method silently leaves data where it is. The recommended definition delegates to the backend's array type, so that adapt(backend, x) behaves exactly like adapt(BackendArray, x):

Adapt.adapt_storage(::CUDABackend, x) = adapt(CuArray, x)
KernelAbstractions 0.10

adapt(backend, x) has been supported by the GPU backends since KernelAbstractions 0.9, but is only documented, and required of every backend, since 0.10.

source
KernelInterface.allocate — Function
allocate(::Backend, Type, dims...; unified=false)::AbstractArray

Allocate an uninitialized array on the active device of the backend. unified=true allocates unified memory, accessible from the host and the device without explicit copies, if the backend supports it and throws otherwise. Use supports_unified to determine whether it is supported by a backend.

Note

Backend implementations must implement allocate(::NewBackend, T, dims::Tuple) Backend implementations should implement allocate(::NewBackend, T, dims::Tuple; unified::Bool=false)

source
KernelInterface.zeros — Function
zeros(::Backend, Type, dims...; unified=false)::AbstractArray

Allocate an array with allocate and fill it with zeros.

This is generic: backends implement allocate (and fill! for their array type).

source
KernelInterface.ones — Function
ones(::Backend, Type, dims...; unified=false)::AbstractArray

Allocate an array with allocate and fill it with ones.

This is generic: backends implement allocate (and fill! for their array type).

source
KernelInterface.copyto! — Function
copyto!(::Backend, dest::AbstractArray, src::AbstractArray)::typeof(dest)

Copy the elements of src to dest, ordered with respect to the other work on the calling task's queue: after work queued before the copy, and before work queued after it. Returns dest.

Either array can be a host array or an array of backend. dest and src must have the same length, otherwise an ArgumentError is thrown. Backends only have to support dense (contiguous) arrays with the same element type.

The copy may be asynchronous with respect to the host, but doesn't have to be: it can also block until it has completed. For a simple, synchronous copy, use Base.copyto!.

Warning

Because the copy may be asynchronous, the caller has to keep both arrays alive, and not access them from the host, until the copy has completed, e.g. by calling synchronize before using them. A GC.@preserve around copyto! only keeps them alive until the copy is queued:

arr = zeros(64)
GC.@preserve arr begin
    copyto!(backend, arr, ...)
    # other operations
    synchronize(backend)
end
Note

On some backends it may be necessary to first call pagelock! on host memory to enable fully asynchronous behavior w.r.t to the host.

Note

Backends must implement this function, for host-to-device, device-to-host and device-to-device copies.

source
KernelInterface.pagelock! — Function
pagelock!(::Backend, dest::AbstractArray)::Union{Nothing, Missing}

Pagelock (pin) a host memory buffer for a backend device. This may be necessary for copyto! to perform asynchronously with respect to the host.

This function returns nothing, or missing if not implemented.

Note

Backends may implement this function.

source
KernelInterface.unsafe_free! — Function
unsafe_free!(x::AbstractArray)

Release the memory of an array for reuse by future allocations, reducing pressure on the allocator. The array may not be used afterwards.

This is a hint: releasing the memory is allowed to do nothing.

Note

Backend implementations may implement this function for their array type, and should forward it to their own unsafe_free! if they have one. The fallback is a no-op.

source
KernelInterface.functional — Function
functional(::Backend)::Union{Bool, Missing}

Queries if the provided backend is functional. This may mean different things for different backends, but generally should mean that the necessary drivers and a compute device are available.

This function should return a Bool or missing if not implemented.

KernelAbstractions v0.9.22

This function was added in KernelAbstractions v0.9.22

source
KernelInterface.versioninfo — Function
versioninfo(io::IO=stdout, backend::Backend)::Nothing

Print information about backend to io. It is up to the backends to determine what is relevant.

Note

Backend implementations may implement this function. If they do so, they should implement versioninfo(io::IO, ::Backend)::Nothing

source
KernelInterface.supports_unified — Function
supports_unified(::Backend)::Bool

Whether allocate supports unified=true on the active device: memory that can be accessed from both the host and the device without explicit copies.

Note

Backend implementations must implement this function if they support unified memory. The fallback returns false.

source
KernelInterface.supports_atomics — Function
supports_atomics(::Backend)::Bool

Whether kernels on the active device support Atomix.jl's atomic operations: at least add and compare-and-swap on 32-bit integers and floats in global memory.

Note

Backend implementations must implement this function if they support atomics. The fallback returns false.

source
KernelInterface.supports_float64 — Function
supports_float64(::Backend)::Bool

Whether kernels on the active device support Float64 values.

Note

Backend implementations must implement this function if they support Float64. The fallback returns false.

source

Devices and execution

KernelInterface.synchronize — Function
synchronize(::Backend)

Block the calling task until all work it has queued on the active device of backend has completed.

Note

Backend implementations must implement this function cooperatively, yielding to the Julia scheduler while waiting rather than blocking inside a driver call. It must not wait for work on other queues that the current queue is not ordered after, e.g., by wait_event. See the notes for backend implementations for why.

source
KernelAbstractions.@spawn — Macro
@spawn [threadpool] backend [device=id] expr

Run expr on a new Julia task, like Threads.@spawn, and return the Task. Use it in place of Threads.@spawn to launch kernels from a task. It guarantees that

  • the task runs on the device that was active in the spawning task, or on device when that argument is given;
  • the work the task queues on backend runs after the work the spawning task had queued on backend before calling @spawn;
  • once wait(task) or fetch(task) returns successfully, all work the task queued on backend has completed, so its results may be used from any task. fetch(task) returns the value of expr. If expr throws, the task is not synchronized: its queued work may still be running when the exception surfaces.

Everything else works as for Threads.@spawn: the optional threadpool argument (:default or :interactive) is forwarded, $x captures the value of x at spawn time, and an enclosing @sync waits for the task.

Example

A = KernelAbstractions.ones(backend, Float32, 1024)
mul2_kernel(backend, 64)(A, ndrange = length(A))   # queued by the current task

task = KernelAbstractions.@spawn backend begin
    mul2_kernel(backend, 64)(A, ndrange = length(A))  # ordered after the launch above
    sum(A)
end
fetch(task) == 4 * length(A)

Choosing the device

Which device a task started with plain Threads.@spawn uses depends on the backend, and need not be the spawning task's. @spawn selects the spawning task's device, or the one given by device, an index into 1:ndevices(backend):

task = KernelAbstractions.@spawn backend device=2 begin
    mul2_kernel(backend, 64)(B, ndrange = length(B))
end

The task's work on that device still runs after the spawning task's earlier work on its own device.

Note

@spawn does not order the task's work against work the spawning task queues afterwards. Order conflicting uses of shared data by waiting for the task, or by spawning a new task after that work.

Note

Queued work is ordered, but the spawning task's earlier work need not have completed when expr starts. Before passing that work's results to a consumer outside that ordering, e.g., an MPI call on a GPU buffer, call synchronize(backend) in expr, which also waits for that work. Some backends synchronize implicitly when the buffer's pointer is taken, but portable code should not rely on that. @spawn does not wait for asynchronous operations outside the backend, like MPI.Isend.

Note

Prefer device= over calling device! in expr. To keep work ordered across a manual switch, call record_event before it and wait_event after it, as shown for wait_event.

Backend authors: see the notes for backend implementations for the protocol behind these guarantees.

source
KernelInterface.record_event — Function
record_event(backend::Backend)

Capture the work the calling task has queued on backend's active device before this call, and return a handle for wait_event. Recording need not wait for that work to complete. The handle is only meant for wait_event.

Note

The default implementation calls synchronize and returns nothing. Backends whose queue is task-local may override this to return an event recorded on the current task's queue instead, without blocking the host. Such a backend must then also implement wait_event for the returned type. See the notes for backend implementations.

source
KernelInterface.wait_event — Function
wait_event(backend::Backend, event)

Order the work the calling task subsequently queues on backend's active device after the work captured by event, which was returned by record_event. Returning does not mean that the captured work has completed.

The wait applies to the queue of the device that is active when wait_event is called; switching devices adds no ordering. To order work across a device switch, select the device first and wait afterwards:

event = record_event(backend)   # captures work on the current device
device!(backend, 2)
wait_event(backend, event)      # orders this task's work on device 2 after it
Note

wait_event(::Backend, ::Nothing) is a no-op, matching the default record_event. A backend that returns another event type must implement wait_event for it, either by adding a dependency to the current task's queue or by waiting cooperatively as synchronize does. A backend with more than one device must accept an event recorded on another device, waiting cooperatively if the driver cannot add a cross-device dependency. See the notes for backend implementations for how this ordering should interact with implicit synchronization.

source
KernelInterface.device — Function
device(backend::Backend)::Int

Return the 1-based index of the currently active device for backend.

Note

Backends supporting multiple devices must implement device(backend::Backend)::Int, along with ndevices, device! and device(backend, A). The fallback only works for a single device, and throws if ndevices reports more.

source
device(backend::Backend, A::AbstractArray)::Int

Return the 1-based index of the device that owns the memory of A, independently of the currently active device.

Note

Backends supporting multiple devices must implement this for their array type. The fallback only works for a single device, and throws if ndevices reports more.

source
KernelInterface.ndevices — Function
ndevices(backend::Backend)::Int

Return the number of devices available to backend.

Note

Backends supporting multiple devices must implement ndevices(backend::Backend)::Int, along with device and device!. The fallback returns 1.

source
KernelInterface.device! — Function
device!(backend::Backend, id::Int)::Nothing

Select the active device for backend. id is a 1-based device index; an id outside 1:ndevices(backend) throws an ArgumentError.

device! is not a synchronization point: work queued before the switch is not ordered with respect to work queued after it. To order across a switch, either synchronize beforehand, or bracket the switch with record_event and wait_event.

Example

device!(CUDABackend(), 2)  # use the second CUDA device
Note

Backends supporting multiple devices must implement device!(backend::Backend, id::Int), along with ndevices and device. The fallback only works for a single device, and throws if ndevices reports more.

source
KernelInterface.priority! — Function
priority!(::Backend, prio::Symbol)::Nothing

Set the priority for the backend stream/queue. This is an optional feature that backends may or may not implement. If a backend shall support priorities it must accept :high, :normal, :low. Where :normal is the default.

Note

Backend implementations may implement this function.

source

Kernel handles

KernelAbstractions.Kernel — Type
Kernel{Backend, WorkgroupSize, NDRange, Func}

Host-side handle for a kernel specialized on a backend, workgroup size, and ndrange.

Kernels are created by calling a @kernel function on a backend, for example my_kernel(CUDABackend(), 256). The returned object is callable:

kernel = my_kernel(backend, 64)
kernel(A, B, ndrange=length(A))   # launch asynchronously
synchronize(backend)

Use workgroupsize, ndrange, and backend to inspect a kernel's static configuration.

Kernels are launched on any backend that implements KernelInterface; see the notes for backend implementations.

source

Index loops

KernelAbstractions.foreach_index — Function
foreach_index(f, A::AbstractArray, Bs::AbstractArray...)
foreach_index(f, backend::Backend, indices)

Call f(i) once for every index i of the arrays A, Bs..., or for every index i in indices, with one work item per index. Returns nothing; the iterations run asynchronously.

This is a for loop over indices without a kernel to write out: the body is an ordinary Julia function, which becomes the body of a kernel.

function scale!(y, x)
    foreach_index(y, x) do i
        @inbounds y[i] = 2 * x[i] + 1
    end
    return y
end

scale!(y, x)
synchronize(get_backend(y))

The first form runs on the backend of the arrays, which must all have the same one, over eachindex(A, Bs...): f receives the index that a for i in eachindex(A, Bs...) loop would, a linear index if the arrays have IndexLinear style and a CartesianIndex otherwise. Pass every array that the body indexes with i, so that the index is valid for each of them. Ranges, CartesianIndices and LinearIndices, and views or reshapes of them, have no backend and run on that of the other arrays; if none of the arrays has a backend, use the second form.

The second form runs on backend, over the given indices: a range of integers such as 1:n or axes(A, 2), for which f receives an Int, or a CartesianIndices of such ranges, for which f receives a CartesianIndex. Indices that do not start at 1 are passed on as they are, so this iterates, e.g., the interior of a 2-D array A:

foreach_index(get_backend(A), CartesianIndices((2:size(A, 1)-1, 2:size(A, 2)-1))) do I
    ...
end

The iterations run concurrently and in no particular order, so they must not race on the same memory. The value f returns is ignored. Like any other kernel launch, foreach_index returns before the iterations have finished: call synchronize before reading their results on the host. Bounds checks are not elided, so write @inbounds in the body where it is warranted, as in a hand-written kernel.

The keyword argument workgroupsize sets the workgroup size of the launch; by default the backend chooses it.

Extended help

f is subject to the same restrictions as a kernel: every value it captures must be of a known type. Closing over a variable of the enclosing global scope leaves its type unknown and fails to compile (typically with unsupported dynamic function invocation), which is why the example above wraps the loop in a function.

For the same reason f cannot capture a variable that is assigned to after f is created, or in f, as Julia then boxes the variable. Bind the value to a new variable (e.g. with let) for f to capture, and accumulate results into an array, with an atomic update if the indices race.

On the CPU backend foreach_index also launches a kernel, compiled for every new f. For a loop that runs once, or over few indices, a threaded loop (Threads.@threads) is cheaper.

See also @kernel to write the kernel out, which is what to reach for when the body needs more of the kernel language than an index (workgroup-level indices, local memory, or synchronization).

source

Reflection

To look at the code a backend actually generates, wrap a kernel launch in one of the @device_code_* macros below. They work the same on the CPU backend and on GPU backends, and they are public, but not exported, so you must call them qualified:

KernelAbstractions.@device_code_llvm mul2(backend, 64)(A, ndrange=length(A))
KernelAbstractions.@device_code_lowered — Macro
KernelAbstractions.@device_code_lowered [kwargs...] ex

Evaluate ex and, for every device kernel compiled along the way, show the lowered IR.

This is GPUCompiler.@device_code_lowered, re-exposed for convenience; see its documentation for the supported keyword arguments. It applies to any GPUCompiler-based backend, so wrapping a kernel launch works on the CPU backend and on GPU backends alike. Note that ex is really evaluated: the kernels it launches are compiled and run.

Examples

KernelAbstractions.@device_code_lowered my_kernel(backend, 64)(A, ndrange=length(A))
source
KernelAbstractions.@device_code_typed — Macro
KernelAbstractions.@device_code_typed [kwargs...] ex

Evaluate ex and, for every device kernel compiled along the way, show the type-inferred IR.

This is GPUCompiler.@device_code_typed, re-exposed for convenience; see its documentation for the supported keyword arguments. It applies to any GPUCompiler-based backend, so wrapping a kernel launch works on the CPU backend and on GPU backends alike. Note that ex is really evaluated: the kernels it launches are compiled and run.

Examples

KernelAbstractions.@device_code_typed my_kernel(backend, 64)(A, ndrange=length(A))
source
KernelAbstractions.@device_code_warntype — Macro
KernelAbstractions.@device_code_warntype [kwargs...] ex

Evaluate ex and, for every device kernel compiled along the way, show the type-inferred IR, highlighting type instabilities.

This is GPUCompiler.@device_code_warntype, re-exposed for convenience; see its documentation for the supported keyword arguments. It applies to any GPUCompiler-based backend, so wrapping a kernel launch works on the CPU backend and on GPU backends alike. Note that ex is really evaluated: the kernels it launches are compiled and run.

Examples

KernelAbstractions.@device_code_warntype my_kernel(backend, 64)(A, ndrange=length(A))
source
KernelAbstractions.@device_code_llvm — Macro
KernelAbstractions.@device_code_llvm [kwargs...] ex

Evaluate ex and, for every device kernel compiled along the way, show the generated LLVM IR.

This is GPUCompiler.@device_code_llvm, re-exposed for convenience; see its documentation for the supported keyword arguments. It applies to any GPUCompiler-based backend, so wrapping a kernel launch works on the CPU backend and on GPU backends alike. Note that ex is really evaluated: the kernels it launches are compiled and run.

Examples

KernelAbstractions.@device_code_llvm my_kernel(backend, 64)(A, ndrange=length(A))
source
KernelAbstractions.@device_code_native — Macro
KernelAbstractions.@device_code_native [kwargs...] ex

Evaluate ex and, for every device kernel compiled along the way, show the generated machine code.

This is GPUCompiler.@device_code_native, re-exposed for convenience; see its documentation for the supported keyword arguments. It applies to any GPUCompiler-based backend, so wrapping a kernel launch works on the CPU backend and on GPU backends alike. Note that ex is really evaluated: the kernels it launches are compiled and run.

Examples

KernelAbstractions.@device_code_native my_kernel(backend, 64)(A, ndrange=length(A))
source
KernelAbstractions.@device_code — Macro
KernelAbstractions.@device_code [dir=...] [...] ex

Evaluate ex and dump all forms of code generated for the device kernels it compiles to the directory dir, or to a temporary directory if none is given.

This is GPUCompiler.@device_code, re-exposed for convenience; see its documentation for the supported keyword arguments. Like the other @device_code_* macros it applies to any GPUCompiler-based backend, and really evaluates ex.

source

Internal

The functionalities in this section are considered internal and not part of the public API contract. They are only documented here for developers and contributors of KernelAbstractions.jl, but should not be used by end users (and if they do, they should expect breakage without notice).

KernelAbstractions.partition — Function
partition(kernel, ndrange, workgroupsize)

Partition the iteration space of kernel into workgroups.

Returns the blocked iteration space and whether dynamic bounds-checking is required for the last (possibly partial) workgroup. Primarily used by backend implementations and tests.

source
KernelAbstractions.@context — Macro
@context()

Access the hidden context object used by KernelAbstractions.

KernelAbstractions 0.10

@context is supported on all backends since KernelAbstractions 0.10.

function f(@context, a)
    I = @index(Global, Linear)
    a[I]
end

@kernel function my_kernel(a)
    f(@context, a)
end
source
KernelAbstractions.argconvert — Function
argconvert(kernel::Kernel, arg)

Convert arg to the device-side representation expected by kernel's backend.

Backend implementations define methods for their array and scalar types. This is called automatically when a kernel is launched.

source
KernelAbstractions.NDIteration.StaticSize — Type
StaticSize{S}

Marker type encoding a compile-time workgroup size or ndrange as a tuple S. Each entry of S is an Int extent or, for an ndrange axis whose indices do not start at 1, a UnitRange{Int}.

source
KernelAbstractions.NDIteration.NDRange — Type
NDRange

Encodes a blocked iteration space. The mapping field relates blocked indices to ndrange indices: nothing for the identity, or a StaticOffset/DynamicOffset for an ndrange whose indices do not start at 1.

Example

ndrange = NDRange{2, DynamicSize, DynamicSize}(CartesianIndices((256, 256)), CartesianIndices((32, 32)))
for block in ndrange
    for items in workitems(ndrange)
        I = expand(ndrange, block, items)
        checkbounds(Bool, A, I) || continue
        @inbounds A[I] = 2*A[I]
    end
end
source
KernelAbstractions.NDIteration.linear_index — Function
linear_index(ndrange::CartesianIndices, I::CartesianIndex)

Column-major position of I within ndrange, counted from 1.

source
linear_index(iterspace::NDRange, ndrange, groupidx::CartesianIndex, idx::CartesianIndex)

Linear index of work item idx of workgroup groupidx, as returned by @index(Global, Linear). Defaults to the position of expand(iterspace, groupidx, idx) within ndrange. A custom mapping whose linear index is not a function of the expanded index alone, e.g. a list of indices whose linear index is the position in the list, can specialize this on its NDRange type.

source
KernelAbstractions.LinearLaunch — Type
LinearLaunch{T}()

Launch configuration for a kernel launched on a 1-D grid: the x-components of the hardware group and local ids are the column-major positions in blocks(iterspace) and workitems(iterspace), like with the default launch (nothing).

@index computes in T (Int32 or Int), which has to hold the number of work-items in the padded iteration space. It still returns Ints.

source
KernelAbstractions.NDLaunch — Type
NDLaunch{T}()

Launch configuration for a kernel launched on an N-d grid, where N = ndims(iterspace) is at most the number of grid dimensions of the backend (the length of KI.max_work_group_dims, i.e. 3): the grid consists of size(blocks(iterspace)) groups of size(workitems(iterspace)) work-items, so the x, y and z-components of the hardware group and local ids are the positions along the first N dimensions of the iteration space. This avoids decomposing linear ids into Cartesian positions, i.e., divisions.

The linear group and local indices are computed x-fastest, so they agree with the ordering of a LinearLaunch (and with how GPUs typically form sub-groups).

@index computes in T (Int32 or Int), which has to hold the number of work-items in the padded iteration space. It still returns Ints.

source
KernelAbstractions.select_launch — Function
select_launch(kernel::Kernel, workgroupsize, iterspace)::Union{LinearLaunch, NDLaunch}

Choose how to launch kernel with the (possibly preliminary) iteration space iterspace and the given workgroupsize (nothing if it will be tuned), as returned by launch_config: an NDLaunch if the iteration space fits the grid of backend(kernel), i.e. its number of dimensions and per-dimension limits, and a LinearLaunch otherwise, both computing indices in Int32 if possible.

If the workgroup size will be tuned, the choice holds for any workgroup size the backend tunes afterwards, so the context and thus the compiled kernel are the same before and after tuning. That requires tuning with launch_workgroupsize and with at most KI.max_work_group_size(backend) work-items, and assumes that iterspace covers the padded iteration space with a single workgroup, as launch_config does.

Throws an ArgumentError if the iteration space has more than typemax(Int) work-items.

The result is not inferred concretely: launching with it takes a dynamic dispatch, unless the caller checks for the common case of an NDLaunch{Int32} first.

source
KernelAbstractions.compiler_options — Function
compiler_options(kernel::Kernel)::NamedTuple

Backend-specific compiler options for compiling kernel with KI.kernel_function, e.g. a hint derived from its static workgroup size (CUDA.jl passes maxthreads). Backends may implement this for their backend type; the default is no options.

source
KernelAbstractions.PrivateArray — Type
PrivateArray{T,S,N,L} <: StaticArray{S,T,N}

Fixed-size array in per-work-item stack storage, as returned by @private. The storage is uninitialized, lives until the kernel returns, and is shared by all copies of the array object. Like a local array in C, every @private declaration has a single allocation: arrays created by the same declaration, e.g. in different iterations of a loop, share storage.

The type only implements indexing; everything else uses the generic AbstractArray implementations, or StaticArrays' if that package is loaded. See @private for what that means in a kernel. It cannot be constructed from values, so copy, zero and other methods that construct a new array of the same type are not supported.

source