Map
AcceleratedKernels.map! — Function
map!(
f, dst::AbstractArray, src::AbstractArray, srcs::AbstractArray...;
backend=nothing,
block_size::Int=256,
max_tasks::Int=Threads.nthreads(),
min_elems::Int=1,
) -> dstApply the function f to each element of src in parallel and store the result in dst. With more source arrays, f takes one element of each, as in Base.map!. dst and the sources must have the same number of elements, which are matched in column-major order, so their shapes may differ. dst may be one of the sources, but must not otherwise share memory with them (a shifted view of a source, say). backend is derived from dst and the sources; the other keywords are those of foreachindex.
On CPUs, multithreading only improves performance when complex computation hides the memory latency and the overhead of spawning tasks - that includes more complex functions and less cache-local array access patterns. For compute-bound tasks, it scales linearly with the number of threads.
Examples
using Metal
import AcceleratedKernels as AK
x = MtlArray(rand(Float32, 100_000))
y = similar(x)
AK.map!(y, x) do x_elem
T = typeof(x_elem)
T(2) * x_elem + T(1)
end
z = similar(x)
AK.map!(+, z, x, y)AcceleratedKernels.map — Function
map(f, src::AbstractArray, srcs::AbstractArray...; kwargs...)Apply the function f as map! does, storing the results in a new array with the shape and eltype of src (if f changes the eltype, allocate dst separately and call map!). The keywords are those of map!.