Scratch Memory
Many operations need scratch memory: merge and radix sorts swap between two buffers, reductions keep partial results per block, scans keep the totals of their tiles, and findall counts the selected elements of each block. By default each call allocates what it needs. To allocate it once and reuse it, make a workspace for the call and pass it through the operation's workspace keyword:
import AcceleratedKernels as AK
using CUDA
v = CuArray(rand(Float32, 1_000_000))
ws = AK.workspace(AK.sort!, v) # the same arguments as the call
for _ in 1:10
rand!(v)
AK.sort!(v; workspace=ws) # allocates no scratch memory
end
AK.workspace_size(AK.sum, v) # what the call needs, without allocating itA workspace belongs to one kind of call: the operation resolves its algorithm and computes its buffers as it would without one, and a workspace made for another backend or device, another algorithm (Auto() may choose differently for another length) or other buffer sizes is an ArgumentError. It holds no state between calls, so calls may reuse it in turn, but not at the same time on several tasks or streams.
AcceleratedKernels.workspace — Function
workspace(op, args...; kwargs...) -> WorkspaceAllocate the scratch memory of the call op(args...; kwargs...), to pass to that operation with the same arguments, or others that resolve to the same algorithm and need the same buffers, through its workspace keyword. The operation then allocates no scratch memory of its own; it still allocates its result where it returns a new array (sort, findall, mapreduce along dims, ...), and compiles kernels as usual.
import AcceleratedKernels as AK
using CUDA
v = CuArray(rand(Float32, 1_000_000))
ws = AK.workspace(AK.sort!, v)
for _ in 1:10
rand!(v)
AK.sort!(v; workspace=ws) # no scratch allocations
endThe workspace is checked on every call: a different backend or device, a different resolved algorithm (Auto() may choose another one for another length), also for the operations it calls, or different buffer sizes are an ArgumentError, and so is a workspace whose buffers alias the operation's arrays.
AcceleratedKernels.workspace_size — Function
workspace_size(op, args...; kwargs...) -> NamedTupleThe scratch buffers the call op(args...; kwargs...) needs, as a NamedTuple of (eltype, dims) pairs (nested for the operations it calls), without allocating anything. An operation with an algorithm whose call needs no scratch gives an empty NamedTuple; the launch wrappers (foreachindex, map!, reverse!, the searches) take no workspace.
julia> AK.workspace_size(AK.sort!, CuArray(rand(Float32, 10_000)); alg=AK.RadixSort())
(temp = (Float32, (10000,)), hist = (UInt32, (5120,)), scan = (prefixes = (UInt32, (3,)),),
key_range = (partials = (Tuple{UInt32, UInt32}, (40,)),))
julia> AK.workspace_size(AK.sum, CuArray(rand(Float32, 10_000)))
(partials = (Float32, (4,)),)AcceleratedKernels.Workspace — Type
WorkspaceScratch memory for one operation, made by workspace and passed to the operation with its workspace keyword. A workspace records the backend and device it was made for, the algorithm the operation resolved to, and its buffers; a call whose arguments need a different algorithm or other buffers throws an ArgumentError instead of using it.
A workspace holds no state between calls, so it can be reused by any number of calls, but not by calls that may run at the same time (on several tasks or streams). It covers the operations' device memory, with two exceptions: on the host, the threaded algorithms still allocate small per-task bookkeeping, and Base.sort! its own scratch for the per-task sorts of CPUThreads.SampleSort and for the slices of a sort along dims; and before Julia 1.12, a reduction of several arrays or of a Broadcasted object materializes it first.