API
Kernel language
KernelAbstractions.@kernel — Macro
@kernel function f(args) endTakes a function definition and generates a Kernel constructor from it. The enclosed function is allowed to contain kernel language constructs. In order to call it the kernel has first to be specialized on the backend and then invoked on the arguments.
Kernel language
Kernel constructor
After defining a kernel function f, call f(backend[, workgroupsize[, ndrange]]) to obtain a Kernel specialized for that backend. Workgroup size and ndrange can be fixed at construction time (enabling size-specific compile-time optimizations and fewer runtime checks, at the cost of recompilation when the sizes change) or supplied at launch:
f(backend) # dynamic workgroup size and ndrange
f(backend, 64) # static workgroup size of 64
f(backend, 64, 1024) # static workgroup size and ndrange
f(backend, 64, (128, 128)) # multi-dimensional ndrangeExample
using KernelAbstractions
@kernel function vecadd(A, @Const(B))
I = @index(Global)
@inbounds A[I] += B[I]
end
dev = CPU()
A = ones(1024)
B = rand(1024)
vecadd(dev, 64)(A, B, ndrange=length(A))
synchronize(dev)@kernel config function f(args) endThis allows for the following configurations:
cpu={true, false}: Deprecated in KernelAbstractions 0.11; this option is ignored.inbounds={false, true}: Enables a forced@inboundsmacro around the function definition in the case the user is using too many@inboundsalready in their kernel. Note that this can lead to incorrect results, crashes, etc and is fundamentally unsafe. Be careful!unsafe_indices={false, true}: Disables the implicit validation of indices, users must avoid@index(Global).generated={false, true}: Turns the kernel into a generated function, see Generated below.
Generated
With generated=true the kernel body is treated as a quoted expression, so $ interpolation is available and where-parameters are bound to their values, e.g. to unroll a loop $N times:
@kernel generated = true function kernel_unroll!(a, ::Val{N}) where {N}
@unroll $N for i in 1:5
@inbounds a[i] = i * $N
end
endThis is meant for macros that need a literal, such as @unroll $N, Base.Cartesian.@nexprs $N or @ntuple $N; plain where-parameters are compile-time constants in every kernel already. Configuration parameters must therefore be passed as types (::Val{N}) to be usable inside $. Inside $(...) the argument names refer to the types of the arguments, not their values, as in any generated function, and the body cannot contain closures, comprehensions or generators (x -> ..., do blocks, [f(i) for i in ...]); use the Cartesian macros above instead.
KernelAbstractions.@Const — Macro
@Const(A)@Const is an argument annotiation that asserts that the memory reference by A is both not written to as part of the kernel and that it does not alias any other memory in the kernel.
The kernel receives such an argument with the backend's read-only array type, so a type annotation on it has to accept that type, e.g., @Const(A::AbstractArray{T}). Which arguments are read-only is determined by the kernel method the arguments select before that, so kernel methods should not dispatch on the read-only type.
KernelAbstractions.@index — Macro
@indexThe @index macro can be used to give you the index of a workitem within a kernel function. It supports both the production of a linear index or a cartesian index. A cartesian index is a general N-dimensional index that is derived from the iteration space.
Index granularity
Global: Used to access global memory.Group: The index of theworkgroup.Local: The withinworkgroupindex.
Index kind
Linear: Produces anInt64that can be used to linearly index into memory.Cartesian: Produces aCartesianIndex{N}that can be used to index into memory.NTuple: Produces aNTuple{N}that can be used to index into memory.
If the index kind is not provided it defaults to Linear, this is subject to change.
Examples
@index(Global, Linear)
@index(Global, Cartesian)
@index(Local, Cartesian)
@index(Group, Linear)
@index(Local, NTuple)
@index(Global)KernelAbstractions.@localmem — Macro
@localmem T dimsDeclare storage that is local to a workgroup.
Like @uniform, the allocation is also executed by padding work-items that fall outside of the ndrange.
KernelAbstractions.@private — Macro
@private T dimsDeclare storage that is private to each work-item. It is preserved across @synchronize statements.
This returns a statically sized StaticArraysCore.StaticArray (a KernelAbstractions.PrivateArray) in stack storage, which supports indexing and in-place operations on non-overlapping regions. Load StaticArrays for static-array arithmetic, slicing and unrolled whole-array reductions. Without it, these fall back to generic AbstractArray methods: operations that return a new array fail to compile, and reductions may spill to local memory or fail to compile, depending on the back-end and size.
Assignment shares storage, and each @private declaration reuses its storage across loop iterations. copy, similar and other operations that allocate are not guaranteed to work, and neither is assigning between overlapping views of the same array.
@private var = exprEquivalent to @uniform var = expr. Ordinary variables are already private to each work-item, and keep their value across @synchronize statements.
KernelAbstractions.@synchronize — Macro
@synchronize()After a @synchronize statement all read and writes to global and local memory from each thread in the workgroup are visible in from all other threads in the workgroup.
@synchronize must be reached by all work-items of a workgroup. To ensure that padding work-items outside of the ndrange reach it too, @kernel treats it specially, so it has to appear directly in the kernel body, and not in a function called by the kernel (unless the kernel uses unsafe_indices=true). Control flow containing it must be uniform across the workgroup, see @uniform.
@synchronize(cond)After a @synchronize statement all read and writes to global and local memory from each thread in the workgroup are visible in from all other threads in the workgroup. cond is not allowed to have any visible sideffects.
Platform differences
GPU: This synchronization will only occur if thecondevaluates.CPU: This synchronization will always occur.
KernelAbstractions.@print — Macro
@print(items...)This is a unified print statement.
Platform differences
GPU: This will reorganize the items to print via@cuprintfCPU: This will callprint(items...)
KernelAbstractions.@uniform — Macro
@uniform exprEvaluate expr on every work-item of the workgroup, including padding work-items that fall outside of the ndrange. This only applies to @uniform statements at the top level of the kernel, or directly in control flow that contains a @synchronize. Elsewhere, @uniform has no effect.
When the ndrange is not a multiple of the workgroup size, @kernel only runs the kernel body on work-items inside the ndrange. Padding work-items do still need to reach every @synchronize, so control flow that contains a @synchronize runs on all work-items. Values used by such control flow, like the bounds of a loop, must therefore be computed with @uniform:
@kernel function f(A, n)
i = @index(Global)
@uniform iterations = 2n
for j in 1:iterations
A[i] += j
@synchronize()
end
end@uniform statements are hoisted to the start of the code between two @synchronize statements. expr must be safe to evaluate on padding work-items, so it should not depend on the work-item's index.
KernelAbstractions.@groupsize — Macro
@groupsize()Query the workgroupsize on the backend. This function returns a tuple corresponding to kernel configuration. In order to get the total size you can use prod(@groupsize()).
KernelAbstractions.@ndrange — Macro
@ndrange()Query the ndrange on the backend. This function returns a tuple corresponding to kernel configuration.
Host language
The Backend type hierarchy and most of the host-side management functions below (get_backend, allocate, synchronize, device selection, …) are defined in the KernelInterface sibling package and re-exported by KernelAbstractions, so KernelAbstractions.allocate and KernelInterface.allocate are the same function. User code can keep calling them through KernelAbstractions as before. Note that this only applies to KernelAbstractions 0.10 and later: KernelAbstractions 0.9 defines its own versions of these functions, which are distinct from the KernelInterface ones.
Backends and arrays
KernelInterface.Backend — Type
BackendAbstract supertype for all KernelInterface backends. Backends subtype it directly.
A backend value identifies a backend and its configuration (e.g. compiler options). The device and the queue that operations use are task-local: each task selects its active device with device!. Host-side queries answer for the calling task's active device, and work is queued on the calling task's queue of that device. Use get_backend to obtain the backend of an array and allocate to create storage on a backend.
Example
backend = get_backend(A)
kernel = my_kernel(backend, 256)
kernel(A, ndrange=length(A))
synchronize(backend)KernelAbstractions.CPU — Type
CPUType alias for POCLBackend, the CPU execution backend.
Construct with CPU() (equivalent to POCLBackend()). Kernels run on the host via POCL/OpenCL using the same programming model as GPU backends, which is useful for debugging and for running kernel code without a GPU.
Example
A = ones(Float32, 1024)
mul2_kernel(CPU(), 64)(A, ndrange=length(A))
synchronize(CPU())Threads
Kernels run on as many threads as Julia's default thread pool has (julia -t N), up to the number of hardware threads. These are POCL's own threads, so launching a kernel doesn't occupy Julia's. To use a different number, set the JULIA_KA_CPU_THREADS environment variable before the backend is first used, e.g., JULIA_KA_CPU_THREADS=8 julia -t1. POCL's own variables (e.g., POCL_CPU_MAX_CU_COUNT) are respected too, and can also raise the number above the number of hardware threads, but they also affect other users of POCL, like OpenCL.jl. The device reports the number as its compute units: KernelAbstractions.POCL.device().max_compute_units.
Atomics
8- and 16-bit atomic operations (e.g., on Int8, UInt16 or Float16 arrays) are performed on the aligned 32-bit word that contains the value, which must not overlap another allocation or memory that is modified independently. For arrays backed by memory that Julia allocated (e.g., by allocate, Array, zeros, resize!, copy or similar), that word stays within the same allocation, because of how Julia's allocator is implemented. For arrays that wrap other memory, e.g., with unsafe_wrap, this is up to the user: make that memory extend to the 4-byte boundaries around the array, e.g., by allocating a multiple of 4 bytes at a 4-byte aligned address.
This includes adjacent elements of the same array: while an 8- or 16-bit element is updated atomically, the other elements in its aligned 32-bit word must not be modified by plain stores concurrently, only atomically.
KernelAbstractions.POCL.POCLKernels.POCLBackend — Type
POCLBackend()CPU backend that compiles kernels to OpenCL via POCL and executes them on the host. This is the concrete type behind the CPU alias.
KernelInterface.get_backend — Function
get_backend(A::AbstractArray)::BackendGet a Backend instance suitable for array A.
Backend implementations must provide get_backend for their custom array type. It should be the same as the return type of allocate
Adapt.adapt_storage — Method
adapt(backend::Backend, x)Convert x such that its array storage lives on backend. This is an extension of Adapt.jl, and lets code move data to a backend without knowing the backend's array type:
using Adapt
x = adapt(CUDABackend(), rand(Float32, 8)) # a CuArray
y = adapt(CPU(), x) # an Array againBackend implementations must implement Adapt.adapt_storage(::NewBackend, x). Adapt.jl's fallback is the identity, so a backend that omits this method silently leaves data where it is. The recommended definition delegates to the backend's array type, so that adapt(backend, x) behaves exactly like adapt(BackendArray, x):
Adapt.adapt_storage(::CUDABackend, x) = adapt(CuArray, x)KernelInterface.allocate — Function
allocate(::Backend, Type, dims...; unified=false)::AbstractArrayAllocate an uninitialized array on the active device of the backend. unified=true allocates unified memory, accessible from the host and the device without explicit copies, if the backend supports it and throws otherwise. Use supports_unified to determine whether it is supported by a backend.
KernelInterface.zeros — Function
zeros(::Backend, Type, dims...; unified=false)::AbstractArrayAllocate an array with allocate and fill it with zeros.
This is generic: backends implement allocate (and fill! for their array type).
KernelInterface.ones — Function
ones(::Backend, Type, dims...; unified=false)::AbstractArrayAllocate an array with allocate and fill it with ones.
This is generic: backends implement allocate (and fill! for their array type).
KernelInterface.copyto! — Function
copyto!(::Backend, dest::AbstractArray, src::AbstractArray)::typeof(dest)Copy the elements of src to dest, ordered with respect to the other work on the calling task's queue: after work queued before the copy, and before work queued after it. Returns dest.
Either array can be a host array or an array of backend. dest and src must have the same length, otherwise an ArgumentError is thrown. Backends only have to support dense (contiguous) arrays with the same element type.
The copy may be asynchronous with respect to the host, but doesn't have to be: it can also block until it has completed. For a simple, synchronous copy, use Base.copyto!.
Because the copy may be asynchronous, the caller has to keep both arrays alive, and not access them from the host, until the copy has completed, e.g. by calling synchronize before using them. A GC.@preserve around copyto! only keeps them alive until the copy is queued:
arr = zeros(64)
GC.@preserve arr begin
copyto!(backend, arr, ...)
# other operations
synchronize(backend)
endOn some backends it may be necessary to first call pagelock! on host memory to enable fully asynchronous behavior w.r.t to the host.
KernelInterface.pagelock! — Function
pagelock!(::Backend, dest::AbstractArray)::Union{Nothing, Missing}Pagelock (pin) a host memory buffer for a backend device. This may be necessary for copyto! to perform asynchronously with respect to the host.
This function returns nothing, or missing if not implemented.
KernelInterface.unsafe_free! — Function
unsafe_free!(x::AbstractArray)Release the memory of an array for reuse by future allocations, reducing pressure on the allocator. The array may not be used afterwards.
This is a hint: releasing the memory is allowed to do nothing.
KernelInterface.functional — Function
functional(::Backend)::Union{Bool, Missing}Queries if the provided backend is functional. This may mean different things for different backends, but generally should mean that the necessary drivers and a compute device are available.
This function should return a Bool or missing if not implemented.
KernelInterface.versioninfo — Function
versioninfo(io::IO=stdout, backend::Backend)::NothingPrint information about backend to io. It is up to the backends to determine what is relevant.
KernelInterface.supports_unified — Function
supports_unified(::Backend)::BoolWhether allocate supports unified=true on the active device: memory that can be accessed from both the host and the device without explicit copies.
KernelInterface.supports_atomics — Function
supports_atomics(::Backend)::BoolWhether kernels on the active device support Atomix.jl's atomic operations: at least add and compare-and-swap on 32-bit integers and floats in global memory.
KernelInterface.supports_float64 — Function
supports_float64(::Backend)::BoolWhether kernels on the active device support Float64 values.
Devices and execution
KernelInterface.synchronize — Function
synchronize(::Backend)Block the calling task until all work it has queued on the active device of backend has completed.
Backend implementations must implement this function cooperatively, yielding to the Julia scheduler while waiting rather than blocking inside a driver call. It must not wait for work on other queues that the current queue is not ordered after, e.g., by wait_event. See the notes for backend implementations for why.
KernelAbstractions.@spawn — Macro
@spawn [threadpool] backend [device=id] exprRun expr on a new Julia task, like Threads.@spawn, and return the Task. Use it in place of Threads.@spawn to launch kernels from a task. It guarantees that
- the task runs on the device that was active in the spawning task, or on
devicewhen that argument is given; - the work the task queues on
backendruns after the work the spawning task had queued onbackendbefore calling@spawn; - once
wait(task)orfetch(task)returns successfully, all work the task queued onbackendhas completed, so its results may be used from any task.fetch(task)returns the value ofexpr. Ifexprthrows, the task is not synchronized: its queued work may still be running when the exception surfaces.
Everything else works as for Threads.@spawn: the optional threadpool argument (:default or :interactive) is forwarded, $x captures the value of x at spawn time, and an enclosing @sync waits for the task.
Example
A = KernelAbstractions.ones(backend, Float32, 1024)
mul2_kernel(backend, 64)(A, ndrange = length(A)) # queued by the current task
task = KernelAbstractions.@spawn backend begin
mul2_kernel(backend, 64)(A, ndrange = length(A)) # ordered after the launch above
sum(A)
end
fetch(task) == 4 * length(A)Choosing the device
Which device a task started with plain Threads.@spawn uses depends on the backend, and need not be the spawning task's. @spawn selects the spawning task's device, or the one given by device, an index into 1:ndevices(backend):
task = KernelAbstractions.@spawn backend device=2 begin
mul2_kernel(backend, 64)(B, ndrange = length(B))
endThe task's work on that device still runs after the spawning task's earlier work on its own device.
@spawn does not order the task's work against work the spawning task queues afterwards. Order conflicting uses of shared data by waiting for the task, or by spawning a new task after that work.
Queued work is ordered, but the spawning task's earlier work need not have completed when expr starts. Before passing that work's results to a consumer outside that ordering, e.g., an MPI call on a GPU buffer, call synchronize(backend) in expr, which also waits for that work. Some backends synchronize implicitly when the buffer's pointer is taken, but portable code should not rely on that. @spawn does not wait for asynchronous operations outside the backend, like MPI.Isend.
Prefer device= over calling device! in expr. To keep work ordered across a manual switch, call record_event before it and wait_event after it, as shown for wait_event.
Backend authors: see the notes for backend implementations for the protocol behind these guarantees.
KernelInterface.record_event — Function
record_event(backend::Backend)Capture the work the calling task has queued on backend's active device before this call, and return a handle for wait_event. Recording need not wait for that work to complete. The handle is only meant for wait_event.
The default implementation calls synchronize and returns nothing. Backends whose queue is task-local may override this to return an event recorded on the current task's queue instead, without blocking the host. Such a backend must then also implement wait_event for the returned type. See the notes for backend implementations.
KernelInterface.wait_event — Function
wait_event(backend::Backend, event)Order the work the calling task subsequently queues on backend's active device after the work captured by event, which was returned by record_event. Returning does not mean that the captured work has completed.
The wait applies to the queue of the device that is active when wait_event is called; switching devices adds no ordering. To order work across a device switch, select the device first and wait afterwards:
event = record_event(backend) # captures work on the current device
device!(backend, 2)
wait_event(backend, event) # orders this task's work on device 2 after itwait_event(::Backend, ::Nothing) is a no-op, matching the default record_event. A backend that returns another event type must implement wait_event for it, either by adding a dependency to the current task's queue or by waiting cooperatively as synchronize does. A backend with more than one device must accept an event recorded on another device, waiting cooperatively if the driver cannot add a cross-device dependency. See the notes for backend implementations for how this ordering should interact with implicit synchronization.
KernelInterface.device — Function
device(backend::Backend)::IntReturn the 1-based index of the currently active device for backend.
device(backend::Backend, A::AbstractArray)::IntReturn the 1-based index of the device that owns the memory of A, independently of the currently active device.
Backends supporting multiple devices must implement this for their array type. The fallback only works for a single device, and throws if ndevices reports more.
KernelInterface.ndevices — Function
ndevices(backend::Backend)::IntReturn the number of devices available to backend.
KernelInterface.device! — Function
device!(backend::Backend, id::Int)::NothingSelect the active device for backend. id is a 1-based device index; an id outside 1:ndevices(backend) throws an ArgumentError.
device! is not a synchronization point: work queued before the switch is not ordered with respect to work queued after it. To order across a switch, either synchronize beforehand, or bracket the switch with record_event and wait_event.
Example
device!(CUDABackend(), 2) # use the second CUDA deviceKernelInterface.priority! — Function
priority!(::Backend, prio::Symbol)::NothingSet the priority for the backend stream/queue. This is an optional feature that backends may or may not implement. If a backend shall support priorities it must accept :high, :normal, :low. Where :normal is the default.
Kernel handles
KernelAbstractions.Kernel — Type
Kernel{Backend, WorkgroupSize, NDRange, Func}Host-side handle for a kernel specialized on a backend, workgroup size, and ndrange.
Kernels are created by calling a @kernel function on a backend, for example my_kernel(CUDABackend(), 256). The returned object is callable:
kernel = my_kernel(backend, 64)
kernel(A, B, ndrange=length(A)) # launch asynchronously
synchronize(backend)Use workgroupsize, ndrange, and backend to inspect a kernel's static configuration.
Kernels are launched on any backend that implements KernelInterface; see the notes for backend implementations.
KernelAbstractions.workgroupsize — Function
workgroupsize(kernel::Kernel)Return the static workgroup size type parameter of kernel (StaticSize or DynamicSize).
KernelAbstractions.ndrange — Function
ndrange(ctx)Return the launch ndrange as a tuple.
KernelAbstractions.backend — Function
backend(kernel::Kernel)Return the Backend that kernel was constructed for.
Index loops
KernelAbstractions.foreach_index — Function
foreach_index(f, A::AbstractArray, Bs::AbstractArray...)
foreach_index(f, backend::Backend, indices)Call f(i) once for every index i of the arrays A, Bs..., or for every index i in indices, with one work item per index. Returns nothing; the iterations run asynchronously.
This is a for loop over indices without a kernel to write out: the body is an ordinary Julia function, which becomes the body of a kernel.
function scale!(y, x)
foreach_index(y, x) do i
@inbounds y[i] = 2 * x[i] + 1
end
return y
end
scale!(y, x)
synchronize(get_backend(y))The first form runs on the backend of the arrays, which must all have the same one, over eachindex(A, Bs...): f receives the index that a for i in eachindex(A, Bs...) loop would, a linear index if the arrays have IndexLinear style and a CartesianIndex otherwise. Pass every array that the body indexes with i, so that the index is valid for each of them. Ranges, CartesianIndices and LinearIndices, and views or reshapes of them, have no backend and run on that of the other arrays; if none of the arrays has a backend, use the second form.
The second form runs on backend, over the given indices: a range of integers such as 1:n or axes(A, 2), for which f receives an Int, or a CartesianIndices of such ranges, for which f receives a CartesianIndex. Indices that do not start at 1 are passed on as they are, so this iterates, e.g., the interior of a 2-D array A:
foreach_index(get_backend(A), CartesianIndices((2:size(A, 1)-1, 2:size(A, 2)-1))) do I
...
endThe iterations run concurrently and in no particular order, so they must not race on the same memory. The value f returns is ignored. Like any other kernel launch, foreach_index returns before the iterations have finished: call synchronize before reading their results on the host. Bounds checks are not elided, so write @inbounds in the body where it is warranted, as in a hand-written kernel.
The keyword argument workgroupsize sets the workgroup size of the launch; by default the backend chooses it.
Extended help
f is subject to the same restrictions as a kernel: every value it captures must be of a known type. Closing over a variable of the enclosing global scope leaves its type unknown and fails to compile (typically with unsupported dynamic function invocation), which is why the example above wraps the loop in a function.
For the same reason f cannot capture a variable that is assigned to after f is created, or in f, as Julia then boxes the variable. Bind the value to a new variable (e.g. with let) for f to capture, and accumulate results into an array, with an atomic update if the indices race.
On the CPU backend foreach_index also launches a kernel, compiled for every new f. For a loop that runs once, or over few indices, a threaded loop (Threads.@threads) is cheaper.
See also @kernel to write the kernel out, which is what to reach for when the body needs more of the kernel language than an index (workgroup-level indices, local memory, or synchronization).
Reflection
To look at the code a backend actually generates, wrap a kernel launch in one of the @device_code_* macros below. They work the same on the CPU backend and on GPU backends, and they are public, but not exported, so you must call them qualified:
KernelAbstractions.@device_code_llvm mul2(backend, 64)(A, ndrange=length(A))KernelAbstractions.@device_code_lowered — Macro
KernelAbstractions.@device_code_lowered [kwargs...] exEvaluate ex and, for every device kernel compiled along the way, show the lowered IR.
This is GPUCompiler.@device_code_lowered, re-exposed for convenience; see its documentation for the supported keyword arguments. It applies to any GPUCompiler-based backend, so wrapping a kernel launch works on the CPU backend and on GPU backends alike. Note that ex is really evaluated: the kernels it launches are compiled and run.
Examples
KernelAbstractions.@device_code_lowered my_kernel(backend, 64)(A, ndrange=length(A))KernelAbstractions.@device_code_typed — Macro
KernelAbstractions.@device_code_typed [kwargs...] exEvaluate ex and, for every device kernel compiled along the way, show the type-inferred IR.
This is GPUCompiler.@device_code_typed, re-exposed for convenience; see its documentation for the supported keyword arguments. It applies to any GPUCompiler-based backend, so wrapping a kernel launch works on the CPU backend and on GPU backends alike. Note that ex is really evaluated: the kernels it launches are compiled and run.
Examples
KernelAbstractions.@device_code_typed my_kernel(backend, 64)(A, ndrange=length(A))KernelAbstractions.@device_code_warntype — Macro
KernelAbstractions.@device_code_warntype [kwargs...] exEvaluate ex and, for every device kernel compiled along the way, show the type-inferred IR, highlighting type instabilities.
This is GPUCompiler.@device_code_warntype, re-exposed for convenience; see its documentation for the supported keyword arguments. It applies to any GPUCompiler-based backend, so wrapping a kernel launch works on the CPU backend and on GPU backends alike. Note that ex is really evaluated: the kernels it launches are compiled and run.
Examples
KernelAbstractions.@device_code_warntype my_kernel(backend, 64)(A, ndrange=length(A))KernelAbstractions.@device_code_llvm — Macro
KernelAbstractions.@device_code_llvm [kwargs...] exEvaluate ex and, for every device kernel compiled along the way, show the generated LLVM IR.
This is GPUCompiler.@device_code_llvm, re-exposed for convenience; see its documentation for the supported keyword arguments. It applies to any GPUCompiler-based backend, so wrapping a kernel launch works on the CPU backend and on GPU backends alike. Note that ex is really evaluated: the kernels it launches are compiled and run.
Examples
KernelAbstractions.@device_code_llvm my_kernel(backend, 64)(A, ndrange=length(A))KernelAbstractions.@device_code_native — Macro
KernelAbstractions.@device_code_native [kwargs...] exEvaluate ex and, for every device kernel compiled along the way, show the generated machine code.
This is GPUCompiler.@device_code_native, re-exposed for convenience; see its documentation for the supported keyword arguments. It applies to any GPUCompiler-based backend, so wrapping a kernel launch works on the CPU backend and on GPU backends alike. Note that ex is really evaluated: the kernels it launches are compiled and run.
Examples
KernelAbstractions.@device_code_native my_kernel(backend, 64)(A, ndrange=length(A))KernelAbstractions.@device_code — Macro
KernelAbstractions.@device_code [dir=...] [...] exEvaluate ex and dump all forms of code generated for the device kernels it compiles to the directory dir, or to a temporary directory if none is given.
This is GPUCompiler.@device_code, re-exposed for convenience; see its documentation for the supported keyword arguments. Like the other @device_code_* macros it applies to any GPUCompiler-based backend, and really evaluates ex.
Internal
The functionalities in this section are considered internal and not part of the public API contract. They are only documented here for developers and contributors of KernelAbstractions.jl, but should not be used by end users (and if they do, they should expect breakage without notice).
KernelAbstractions.partition — Function
partition(kernel, ndrange, workgroupsize)Partition the iteration space of kernel into workgroups.
Returns the blocked iteration space and whether dynamic bounds-checking is required for the last (possibly partial) workgroup. Primarily used by backend implementations and tests.
KernelAbstractions.@context — Macro
@context()Access the hidden context object used by KernelAbstractions.
function f(@context, a)
I = @index(Global, Linear)
a[I]
end
@kernel function my_kernel(a)
f(@context, a)
endKernelAbstractions.argconvert — Function
argconvert(kernel::Kernel, arg)Convert arg to the device-side representation expected by kernel's backend.
Backend implementations define methods for their array and scalar types. This is called automatically when a kernel is launched.
KernelAbstractions.NDIteration.DynamicSize — Type
DynamicSizeMarker type indicating that a kernel's workgroup size or ndrange is chosen at launch time.
KernelAbstractions.NDIteration.StaticSize — Type
StaticSize{S}Marker type encoding a compile-time workgroup size or ndrange as a tuple S. Each entry of S is an Int extent or, for an ndrange axis whose indices do not start at 1, a UnitRange{Int}.
KernelAbstractions.NDIteration.NDRange — Type
NDRangeEncodes a blocked iteration space. The mapping field relates blocked indices to ndrange indices: nothing for the identity, or a StaticOffset/DynamicOffset for an ndrange whose indices do not start at 1.
Example
ndrange = NDRange{2, DynamicSize, DynamicSize}(CartesianIndices((256, 256)), CartesianIndices((32, 32)))
for block in ndrange
for items in workitems(ndrange)
I = expand(ndrange, block, items)
checkbounds(Bool, A, I) || continue
@inbounds A[I] = 2*A[I]
end
endKernelAbstractions.NDIteration.StaticOffset — Type
StaticOffset{O}Compile-time offset O::NTuple{N, Int} added to the indices produced by an NDRange.
KernelAbstractions.NDIteration.DynamicOffset — Type
DynamicOffset{N}Runtime offset added to the indices produced by an NDRange.
KernelAbstractions.NDIteration.extents — Function
extents(ndrange)Number of indices along each axis of ndrange, given as a tuple of extents and/or ranges, a CartesianIndices, a single range, or an integer.
KernelAbstractions.NDIteration.offsets — Function
offsets(ndrange)Offset of the first index along each axis of ndrange relative to 1.
KernelAbstractions.NDIteration.linear_index — Function
linear_index(ndrange::CartesianIndices, I::CartesianIndex)Column-major position of I within ndrange, counted from 1.
linear_index(iterspace::NDRange, ndrange, groupidx::CartesianIndex, idx::CartesianIndex)Linear index of work item idx of workgroup groupidx, as returned by @index(Global, Linear). Defaults to the position of expand(iterspace, groupidx, idx) within ndrange. A custom mapping whose linear index is not a function of the expanded index alone, e.g. a list of indices whose linear index is the position in the list, can specialize this on its NDRange type.
KernelAbstractions.LinearLaunch — Type
LinearLaunch{T}()Launch configuration for a kernel launched on a 1-D grid: the x-components of the hardware group and local ids are the column-major positions in blocks(iterspace) and workitems(iterspace), like with the default launch (nothing).
@index computes in T (Int32 or Int), which has to hold the number of work-items in the padded iteration space. It still returns Ints.
KernelAbstractions.NDLaunch — Type
NDLaunch{T}()Launch configuration for a kernel launched on an N-d grid, where N = ndims(iterspace) is at most the number of grid dimensions of the backend (the length of KI.max_work_group_dims, i.e. 3): the grid consists of size(blocks(iterspace)) groups of size(workitems(iterspace)) work-items, so the x, y and z-components of the hardware group and local ids are the positions along the first N dimensions of the iteration space. This avoids decomposing linear ids into Cartesian positions, i.e., divisions.
The linear group and local indices are computed x-fastest, so they agree with the ordering of a LinearLaunch (and with how GPUs typically form sub-groups).
@index computes in T (Int32 or Int), which has to hold the number of work-items in the padded iteration space. It still returns Ints.
KernelAbstractions.select_launch — Function
select_launch(kernel::Kernel, workgroupsize, iterspace)::Union{LinearLaunch, NDLaunch}Choose how to launch kernel with the (possibly preliminary) iteration space iterspace and the given workgroupsize (nothing if it will be tuned), as returned by launch_config: an NDLaunch if the iteration space fits the grid of backend(kernel), i.e. its number of dimensions and per-dimension limits, and a LinearLaunch otherwise, both computing indices in Int32 if possible.
If the workgroup size will be tuned, the choice holds for any workgroup size the backend tunes afterwards, so the context and thus the compiled kernel are the same before and after tuning. That requires tuning with launch_workgroupsize and with at most KI.max_work_group_size(backend) work-items, and assumes that iterspace covers the padded iteration space with a single workgroup, as launch_config does.
Throws an ArgumentError if the iteration space has more than typemax(Int) work-items.
The result is not inferred concretely: launching with it takes a dynamic dispatch, unless the caller checks for the common case of an NDLaunch{Int32} first.
KernelAbstractions.launch_workgroupsize — Function
launch_workgroupsize(backend, launch, threads, ndrange)The workgroup size for launching threads work-items per workgroup over ndrange with launch (see select_launch). Backends that tune the workgroup size of a kernel launched with a LinearLaunch or NDLaunch have to use this.
KernelAbstractions.compiler_options — Function
compiler_options(kernel::Kernel)::NamedTupleBackend-specific compiler options for compiling kernel with KI.kernel_function, e.g. a hint derived from its static workgroup size (CUDA.jl passes maxthreads). Backends may implement this for their backend type; the default is no options.
KernelAbstractions.PrivateArray — Type
PrivateArray{T,S,N,L} <: StaticArray{S,T,N}Fixed-size array in per-work-item stack storage, as returned by @private. The storage is uninitialized, lives until the kernel returns, and is shared by all copies of the array object. Like a local array in C, every @private declaration has a single allocation: arrays created by the same declaration, e.g. in different iterations of a loop, share storage.
The type only implements indexing; everything else uses the generic AbstractArray implementations, or StaticArrays' if that package is loaded. See @private for what that means in a kernel. It cannot be constructed from values, so copy, zero and other methods that construct a new array of the same type are not supported.