代码之家  ›  专栏  ›  技术社区  ›  splicer

如何在OpenCL中使用本地内存?

  •  39
  • splicer  · 技术社区  · 16 年前

    我最近一直在玩OpenCL,我能够编写只使用全局内存的简单内核。现在我想开始使用本地内存,但我似乎不知道如何使用 get_local_size() 和 get_local_id()

    例如,假设我想将Apple的OpenCL Hello World示例内核转换为使用本地内存的内核。你会怎么做?以下是原始内核源代码:

    __kernel square(
        __global float *input,
        __global float *output,
        const unsigned int count)
    {
        int i = get_global_id(0);
        if (i < count)
            output[i] = input[i] * input[i];
    }
    

    3 回复  |  直到 10 年前
        1
  •  32
  •   Tom    11 年前

    看看英伟达或AMD SDK的样本,他们应该指出你在正确的方向。例如,矩阵转置将使用本地内存。

    使用平方内核,可以将数据暂存在中间缓冲区中。记住传入附加参数。

    __kernel square(
        __global float *input,
        __global float *output,
        __local float *temp,
        const unsigned int count)
    {
        int gtid = get_global_id(0);
        int ltid = get_local_id(0);
        if (gtid < count)
        {
            temp[ltid] = input[gtid];
            // if the threads were reading data from other threads, then we would
            // want a barrier here to ensure the write completes before the read
            output[gtid] =  temp[ltid] * temp[ltid];
        }
    }
    
        2
  •  29
  •   Rick-Rainer Ludwig    15 年前

    __local float localBuffer[1024];
    

    由于clSetKernelArg调用较少,这将删除代码。

        3
  •  5
  •   Hunter Wang    13 年前

    在OpenCL中,本地内存意味着在一个工作组中的所有工作项之间共享数据。在使用本地内存数据之前,通常需要执行一个barrier调用(例如,一个工作项想要读取由其他工作项写入的本地内存数据)。屏障的硬件成本很高。请记住,本地内存应该用于重复的数据读/写。应尽可能避免银行冲突。

    推荐文章