代码之家  ›  专栏  ›  技术社区  ›  Nikita Sivukhin

CPU缓存理解

  •  4
  • Nikita Sivukhin  · 技术社区  · 7 年前

    int64_t MemoryAccessAllElements(const int64_t *data, size_t length) {
        for (size_t id = 0; id < length; id++) {
            volatile int64_t ignored = data[id];
        }
        return 0;
    }
    int64_t MemoryAccessEvery4th(const int64_t *data, size_t length) {
        for (size_t id = 0; id < length; id += 4) {
            volatile int64_t ignored = data[id];
        }
        return 0;
    }
    

    然后我得到下一个结果(结果由google基准测试平均,对于大型数组,大约有10次迭代,对于较小的数组,需要执行更多的工作):

    Benchmark results

    这张照片上发生了很多不同的事情,不幸的是,我无法解释图表中的所有变化。

    CPU Caches:                                                                                              
      L1 Data 32K (x1), 8 way associative
      L1 Instruction 32K (x1), 8 way associative
      L2 Unified 256K (x1), 8 way associative
      L3 Unified 30720K (x1), 20 way associative
    

    在这些图片中,我们可以看到图形行为的许多变化:

    1. 64字节数组大小之后会出现一个峰值,这可以通过以下事实来解释:缓存线大小为64字节长,并且当数组大小超过64字节时,我们又会遇到一次一级缓存未命中(这可以归类为强制缓存未命中)

    但是有很多关于结果的问题我无法解释:

    1. 为什么要延迟 MemoryAccessEvery4th
    2. 为什么我们可以看到另一个高峰 MemoryAccessAllElements 大约512字节?这是一个有趣的点,因为此时我们开始访问多个缓存线集(一个缓存线集中有8*64字节)。但它真的是由这一事件引起的吗?如果是,如何解释?
    3. 为什么我们可以看到在基准测试时传递二级缓存大小后延迟增加 内存访问次数4 记忆附件 ?

    gallery of processor cache effects 和 what every programmer should know about memory

    有人能帮我理解CPU缓存的内部过程吗?

    升级版本: 我使用以下代码来衡量内存访问的性能:

    #include <benchmark/benchmark.h>
    using namespace benchmark;
    void InitializeWithRandomNumbers(long long *array, size_t length) {
        auto random = Random(0);
        for (size_t id = 0; id < length; id++) {
            array[id] = static_cast<long long>(random.NextLong(0, 1LL << 60));
        }
    }
    static void MemoryAccessAllElements_Benchmark(State &state) {
        size_t size = static_cast<size_t>(state.range(0));
        auto array = new long long[size];
        InitializeWithRandomNumbers(array, size);
        for (auto _ : state) {
            DoNotOptimize(MemoryAccessAllElements(array, size));
        }
        delete[] array;
    }
    static void CustomizeBenchmark(benchmark::internal::Benchmark *benchmark) {
        for (int size = 2; size <= (1 << 24); size *= 2) {
            benchmark->Arg(size);
        }
    }
    BENCHMARK(MemoryAccessAllElements_Benchmark)->Apply(CustomizeBenchmark);
    BENCHMARK_MAIN();
    

    您可以在中找到稍微不同的示例 the repository

    0 回复  |  直到 5 年前