Notes for backend implementations
The KernelInterface sibling package defines the core interface a backend must implement. A backend must implement a backend type that subtypes KernelInterface.Backend. This documentation contains the host and devices side functions that backends can define, as well as whether they are mandatory or not.
Semantics of KernelAbstractions.synchronize
KernelAbstractions.synchronize is required to be cooperative, with that we mean it can not block inside an external library, but instead must implement a cooperative wait that will yield the current task and return the scheduling slice to the Julia runtime.
This is of particular import to allow for overlapping of communication and computation with MPI, and for KernelAbstractions.@spawn, whose trailing synchronize would otherwise stall every task scheduled on the same thread instead of letting independent tasks run concurrently.
Task-local queues and KernelAbstractions.@spawn
Backends should give each Julia task its own queue, so that work from different tasks can execute concurrently. Separate queues do not by themselves order work.
KernelAbstractions.@spawn orders it with this protocol:
- The spawning task calls
record_event. - The new task selects its device with
device!, then callswait_eventbefore running the user's code. - If that code returns normally, the new task calls
synchronize, so that a successfulwait(task)implies its queued work has completed.
The default record_event synchronizes and returns nothing, for which wait_event does nothing. Backends can return an event instead, so that the spawning task doesn't wait; see the docstrings of both functions for what that requires. Backends with more than one device must implement device, ndevices and device!, and wait_event must accept an event recorded on another device, since @spawn backend device=id records on the spawning task's device.
Backends that track which queue last used an array, and wait for that queue before using the array on another one, should skip that wait when the current queue is already ordered after the array's last use: through an event recorded on the previous queue after that use and waited for by the current queue, or through a synchronize of the previous queue that completed that use. In particular, they should not wait, on the host or on the device, for work queued on the previous queue after that event or synchronization. Waits needed to make memory accessible or to keep it alive still apply, and uses through a pointer taken before that event or synchronization may be synchronized conservatively. This lets a spawned task use arrays it shares with its parent while the parent keeps queuing work, which is also why synchronize must not wait for work on other queues that the current queue is not ordered after.
Moving data with adapt
KernelAbstractions extends Adapt.jl so that adapt(backend, x) moves the arrays in x to backend, without the caller having to know the backend's array type. Every backend must support this by extending Adapt.adapt_storage for its backend type. The recommended definition delegates to the backend's array type, so that adapt(backend, x) and adapt(BackendArray, x) agree:
Adapt.adapt_storage(::CUDABackend, x) = adapt(CuArray, x)Launching @kernel kernels
KernelAbstractions launches @kernel kernels on any backend that implements KernelInterface: it partitions the ndrange, builds the kernel's hidden context (a KernelAbstractions.CompilerMetadata), compiles the kernel with KI.kernel_function, tunes the workgroup size, and launches it with KI.launch. A backend doesn't implement any of that itself, but it needs KernelInterface's typed index queries and an N-d KI.launch, and it must not override KernelAbstractions' index functions (see below). @private storage is implemented by an overlay in GPUCompiler.SHARED_METHOD_TABLE, so a GPUCompiler-based backend has to include that table: it should return its method tables from GPUCompiler.method_tables rather than override GPUCompiler.method_table_view. It may customize the launch through:
KI.launch_configuration: the workgroup size used when the kernel has no static or given one. It receives the number of work-items in thendrangeasnitems, e.g. to prefer more workgroups over larger ones.KernelAbstractions.compiler_options: compiler options for a kernel, e.g. a register hint derived from its static workgroup size.Adapt.adapt_storage(::KernelAbstractions.ConstAdaptor, x)for the backend's device arrays, which implements@Const. KernelAbstractions marks@Constarguments when launching a kernel, and they are constified by the backend's argument conversion, on the host. That conversion, byKI.argconvertas well as withinKI.launch, must therefore use Adapt.jl to recurse into the arguments.
@index computes its indices from how a kernel was launched: on a grid with the shape of the iteration space (NDLaunch, for up to as many dimensions as the backend's grid has), which doesn't need any divisions, or on a 1-D grid (LinearLaunch). Either way it computes in a narrow index type such as Int32 when the iteration space fits, which is why the typed KI.get_group_id and KI.get_local_id queries have to compute in that type too, as KernelInterface specifies. For the same reason, backends must not override __validindex or the __index_* functions.
A backend can still implement (obj::KernelAbstractions.Kernel{MyBackend})(args...; ndrange, workgroupsize) to launch kernels itself, e.g. while it is being ported to KernelInterface. That relies on KernelAbstractions internals: it has to choose the launch with select_launch, pass it to the kernel's context, and tune the workgroup size with launch_workgroupsize, as the generic launch in src/backend_launch.jl does.
Packages that customize the iteration space (with a custom partition and expand) don't need to do anything for these launches: the index functions only compute the global index directly for the iteration spaces KernelAbstractions creates itself, and call expand, in and linear_index otherwise.
Packages with an Adapt rule for CompilerMetadata must preserve its launch, e.g. by passing launch = KernelAbstractions.__launch(ctx) to the constructor. Otherwise the kernel computes its indices as if it had been launched on a 1-D grid, which gives wrong results for a kernel launched with an NDLaunch.