Reference

Contents

Index

OpenCL.llvm_to_spirv_backend — Constant
OpenCL.llvm_to_spirv_backend :: Symbol

The backend used to compile LLVM IR to SPIRV (:llvm (default), or :khronos). This is a load-time preference, as compiler configurations are cached across compilations:

using OpenCL, Preferences
set_preferences!(OpenCL, "llvm_to_spirv_backend" => "khronos")

Changing it requires restarting Julia. The :khronos backend additionally requires SPIRV_LLVM_Translator_jll (or SPIRV_LLVM_Translator_unified_jll) to be loaded.

source
OpenCL.Const — Type
Const(A::CLDeviceArray)

Mark a CLDeviceArray as constant/read-only. The invariant guaranteed is that you will not modify an CLDeviceArray for the duration of the current kernel. Loads from it are emitted as !invariant.load, so violating that invariant is undefined behavior.

Warning

Experimental API. Subject to change without deprecation.

source
OpenCL.KernelException — Type
KernelException

An exception thrown during kernel execution on dev, detected when synchronizing (cl.finish, a blocking copy, Array(x), ...).

How much is known about the exception depends on the debug level the kernel was compiled with (the session's -g level, or @opencl debug_level=):

  • 0: only that an exception was thrown;
  • 1: additionally its type name and reason, for exceptions thrown by Julia's runtime (bounds errors, domain errors, ...);
  • 2: additionally the position of the faulting work-item (work_item is its local id, work_group the id of its work-group), the name of any other exception, and a device-side backtrace as (function, file, line) tuples.

dev identifies the device whose mailbox reported the exception. Fields that were not recorded are empty strings, empty vectors, or all-zero tuples.

source
OpenCL.OpenCLResults — Type
OpenCLResults

Cached compilation results for an OpenCL kernel job, managed by GPUCompiler.cached_results. Fields are populated through the compile pipeline: obj (SPIR-V bytes) + entry + device_rng after codegen, and kernels after the session-local link onto an OpenCL context. The first three are session-portable (cached through precompilation, except when GPUCompiler marks the job session-dependent and wipes its entries before image serialization); kernels is session-local and never populated during precompilation. obj === nothing identifies a job that has not been compiled yet.

kernels is a small linear cache of (cl.Context, cl.Kernel) pairs. The cache partition already covers everything that affects codegen via GPUCompiler.cache_owner, so the only runtime-visible dimension left is the OpenCL context that owns the linked cl.Kernel. A linear scan with === is fastest in the common case (n=1) and stays cheap for the rare workload that bounces between a handful of contexts on the same device.

source
Base.rand — Method
Random.rand(rng::Philox2x32, UInt32)

Generate a byte of random data using the on-device Tausworthe generator.

source
Base.resize! — Method

resize!(a::CLVector, n::Integer)

Resize a to contain n elements. If n is smaller than the current collection length, the first n elements will be retained. If n is larger, the new elements are not guaranteed to be initialized.

source
Base.unsafe_wrap — Method
unsafe_wrap(Array, arr::CLArray)

Wrap a Julia Array around the buffer that backs a CLArray. This is only possible if the GPU array is backed by host memory, such as unified (host or shared) memory, or shared virtual memory.

source
OpenCL.check_exceptions — Method
check_exceptions(queue::cl.CmdQueue)

Synchronize queue and every other queue recorded against its device's mailbox, then check whether a kernel threw an exception and, if so, rethrow it host-side as a KernelException.

The exception mailbox is shared by queues targeting the same device in a context, so this may wait for and surface an exception from another queue on that device.

source
OpenCL.format — Method

Format string using dict-like variables, replacing all accurancies of %(key) with value.

Example: s = "Hello, %(name)" format(s, name="Tom") ==> "Hello, Tom"

source
OpenCL.has_feature — Method
has_feature(name::Symbol) -> Bool

Compile-time query (device side): does the kernel's target device support optional feature name (see FEATURES)? Folds to a constant, so if has_feature(:subgroups) … else … end keeps only the live branch. Host-side, use feature_supported(dev, name).

source
OpenCL.kernel_convert — Function
kernel_convert(x)

This function is called for every argument to be passed to a kernel, allowing it to be converted to a GPU-friendly format. By default, the function does nothing and returns the input object x as-is.

Do not add methods to this function, but instead extend the underlying Adapt.jl package and register methods for the the OpenCL.KernelAdaptor type.

source
OpenCL.program_backend! — Method
program_backend!(mode::Symbol)
program_backend!(f::Function, mode::Symbol)

Select how kernels are fed to the driver for the current task: :auto (default), :spirv (SPIR-V IL), or :opencl (OpenCL C source). The second form applies mode only for the duration of f.

source
OpenCL.program_backend — Method
program_backend() -> Symbol

The requested program backend for the current task (:auto, :spirv, or :opencl).

source
OpenCL.return_type — Method
OpenCL.return_type(f, tt) -> r::Type

Return a type r such that f(args...)::r where args::tt.

source
Random.seed! — Function
Random.seed!(rng::Philox2x32, seed::Integer, [counter::Integer=0])

Seed the on-device Philox2x32 generator with an UInt32 number. Should be called by at least one thread per warp.

source
OpenCL.@opencl — Macro
@opencl [kwargs...] func(args...)

High-level interface for executing code on an OpenCL device.

The @opencl macro should prefix a call, with func a callable function or object that should return nothing. It will be compiled to an OpenCL kernel upon first use, and to a certain extent arguments will be converted and managed automatically using kernel_convert. Finally, the kernel is launched on the current queue.

There are a few keyword arguments that influence the behavior of @opencl:

  • launch: whether to launch this kernel, defaults to true. If false, the returned kernel object should be launched by calling it and passing arguments again.
  • name: the name of the kernel in the generated code. Defaults to an automatically- generated name.
  • debug_level: how much a device-side exception reports, from 0 (only that one was thrown) to 2 (its type, reason, position and backtrace). Defaults to the session's -g level; see KernelException.
  • global_size, local_size: the launch configuration, as in OpenCL.
source
OpenCL.cl.CLPtr — Type
CLPtr{T}

A memory address that refers to data of type T that is accessible from q device. A CLPtr is ABI compatible with regular Ptr objects, e.g. it can be used to ccall a function that expects a Ptr to device memory, but it prevents erroneous conversions between the two.

source
OpenCL.cl.PtrOrCLPtr — Type
PtrOrCLPtr{T}

A special pointer type, ABI-compatible with both Ptr and CLPtr, for use in ccall expressions to convert values to either a device or a host type (in that order). This is required for APIs which accept pointers that either point to host or device memory.

source
OpenCL.cl.UnifiedDeviceMemory — Type
UnifiedDeviceMemory

A buffer of device memory, owned by a specific device. Generally, may only be accessed by the device that owns it.

source
OpenCL.cl.UnifiedHostMemory — Type
UnifiedHostMemory

A buffer of memory on the host. May be accessed by the host, and all devices within the host driver. Frequently used as staging areas to transfer data to or from devices.

source
OpenCL.cl.clcall — Method
clcall(kernel, types, args...; kwargs...)

Set the arguments of kernel and enqueue it as one atomic operation with respect to other uses of the same kernel object. Manual use of cl.set_arg!, cl.set_args!, or cl.enqueue_kernel is not internally synchronized; hold lock(kernel) across the complete argument-setup and enqueue sequence.

source
OpenCL.cl.driver_uuid — Method
driver_uuid(d::Device)

Return the universally unique identifier of the driver backing the device as a Base.UUID, or missing if the device does not support the cl_khr_device_uuid extension. Devices using the same driver report the same driver UUID.

source
OpenCL.cl.pci_bus_info — Method
pci_bus_info(d::Device)

Return the PCI address of the device as a named tuple (; domain, bus, device, func), or missing if the device does not support the cl_khr_pci_bus_info extension.

source
OpenCL.cl.uuid — Method
uuid(d::Device)

Return the universally unique identifier of the device as a Base.UUID, or missing if the device does not support the cl_khr_device_uuid extension.

source
OpenCL.cl.@checked — Macro
@checked function foo(...)
    rv = ...
    return rv
end

Macro for wrapping a function definition returning a status code. Two versions of the function will be generated: foo, with the function body wrapped by an invocation of the check function (to be implemented by the caller of this macro), and unchecked_foo where no such invocation is present and the status code is returned to the caller.

source