Map

AcceleratedKernels.map! — Function
map!(
    f, dst::AbstractArray, src::AbstractArray, srcs::AbstractArray...;
    backend=nothing,
    block_size::Int=256,
    max_tasks::Int=Threads.nthreads(),
    min_elems::Int=1,
) -> dst

Apply the function f to each element of src in parallel and store the result in dst. With more source arrays, f takes one element of each, as in Base.map!. dst and the sources must have the same number of elements, which are matched in column-major order, so their shapes may differ. dst may be one of the sources, but must not otherwise share memory with them (a shifted view of a source, say). backend is derived from dst and the sources; the other keywords are those of foreachindex.

On CPUs, multithreading only improves performance when complex computation hides the memory latency and the overhead of spawning tasks - that includes more complex functions and less cache-local array access patterns. For compute-bound tasks, it scales linearly with the number of threads.

Examples

using Metal
import AcceleratedKernels as AK

x = MtlArray(rand(Float32, 100_000))
y = similar(x)
AK.map!(y, x) do x_elem
    T = typeof(x_elem)
    T(2) * x_elem + T(1)
end

z = similar(x)
AK.map!(+, z, x, y)
source
AcceleratedKernels.map — Function
map(f, src::AbstractArray, srcs::AbstractArray...; kwargs...)

Apply the function f as map! does, storing the results in a new array with the shape and eltype of src (if f changes the eltype, allocate dst separately and call map!). The keywords are those of map!.

source