代码之家  ›  专栏  ›  技术社区  ›  markzzz

如何从具有向量数组位置的数组(内存中)中进行最佳读取?

  •  0
  • markzzz  · 技术社区  · 5 年前

    我有这样一个代码:

    const rack::simd::float_4 pos = phase * waveTable.mLength;
    const rack::simd::int32_4 pos0 = pos;
    const rack::simd::float_4 frac = pos - (rack::simd::float_4)pos0;
    rack::simd::float_4 v0;
    rack::simd::float_4 v1;
    for (int v = 0; v < 4; v++) {
        v0[v] = waveTable.mSamples[pos0[v]];
        v1[v] = waveTable.mSamples[pos0[v] + 1]; // mSamples size is waveTable.mLength + 1 for interpolation wraparound
    }
    
    oversampleBuffer[i] = v0 + (v1 - v0) * frac;
    

    这需要 phase (归一化)和插值(线性插值)两个样本,存储在 waveTable.mSamples (每个都是单个浮点数)。

    位置在内部 rack::simd::float_4 ,它们基本上是4个对齐的浮子,定义为 __m128 . 经过一些基准测试后,这部分代码需要一些时间(我想是因为缺少了很多缓存)。

    是用 -march=nocona ,所以我可以使用MMX、SSE、SSE2和SSE3。

    您将如何优化此代码?谢谢

    0 回复  |  直到 5 年前
        1
  •  3
  •   Peter Cordes    5 年前

    由于多种原因,你的代码效率不高。

    1. 您正在使用标量代码设置SIMD向量的各个通道。处理器不能做到这一点,但编译器假装可以。不幸的是,这些编译器实现的解决方法很慢,通常是通过向内存发送一个rountrip并返回来实现的。

    2. 一般来说,您应该避免编写长度为2或4的非常小的循环。有时编译器展开,你就没事了,但其他时候它们没有,CPU预测错误的分支太多了。

    3. 最后,处理器可以用一条指令加载64位值。您正在从表中加载连续的对,可以使用64位加载,而不是两个32位加载。

    这是一个未经测试的固定版本。这假设您正在为PC构建,即使用SSE SIMD。

    // Load a vector with rsi[ i0 ], rsi[ i0 + 1 ], rsi[ i1 ], rsi[ i1 + 1 ]
    inline __m128 loadFloats( const float* rsi, int i0, int i1 )
    {
        // Casting load indices to unsigned, otherwise compiler will emit sign extension instructions
        __m128d res = _mm_load_sd( (const double*)( rsi + (uint32_t)i0 ) );
        res = _mm_loadh_pd( res, (const double*)( rsi + (uint32_t)i1 ) );
        return _mm_castpd_ps( res );
    }
    
    __m128 interpolate4( const float* waveTableData, uint32_t waveTableLength, __m128 phase )
    {
        // Convert wave table length into floats.
        // Consider doing that outside of the inner loop, and passing the __m128.
        const __m128 length = _mm_set1_ps( (float)waveTableLength );
    
        // Compute integer indices, and the fraction
        const __m128 pos = _mm_mul_ps( phase, length );
        const __m128 posFloor = _mm_floor_ps( pos );    // BTW this one needs SSE 4.1, workarounds are expensive
        const __m128 frac = _mm_sub_ps( pos, posFloor );
        const __m128i posInt = _mm_cvtps_epi32( posFloor );
    
        // Abuse 64-bit load instructions to load pairs of values from the table.
        // If you have AVX2, can use _mm256_i32gather_pd instead, will load all 8 floats with 1 (slow) instruction.
        const __m128 s01 = loadFloats( waveTableData, _mm_cvtsi128_si32( posInt ), _mm_extract_epi32( posInt, 1 ) );
        const __m128 s23 = loadFloats( waveTableData, _mm_extract_epi32( posInt, 2 ), _mm_extract_epi32( posInt, 3 ) );
    
        // Shuffle into the correct order, i.e. gather even/odd elements from the vectors
        const __m128 v0 = _mm_shuffle_ps( s01, s23, _MM_SHUFFLE( 2, 0, 2, 0 ) );
        const __m128 v1 = _mm_shuffle_ps( s01, s23, _MM_SHUFFLE( 3, 1, 3, 1 ) );
    
        // Finally, linear interpolation between these vectors.
        const __m128 diff = _mm_sub_ps( v1, v0 );
        return _mm_add_ps( v0, _mm_mul_ps( frac, diff ) );
    }
    

    组装 looks good 现代编译器甚至自动使用FMA, when available (默认情况下为GCC,clang-with -ffp-contract=fast 跨C语句进行收缩,而不仅仅是在一个表达式中。)


    刚刚看到更新。考虑将目标切换到SSE 4.1。Steam硬件调查显示 market penetration is 98.76% 。如果您仍然支持Pentium 4等史前CPU _mm_floor_ps is in DirectXMath ,而不是 _mm_extract_epi32 你可以使用 _mm_srli_si128 + _mm_cvtsi128_si32 .

    即使你需要支持像SSE3这样的旧基线, -mtune=generic 甚至 -mtune=haswell 这可能是个好主意 -march=nocona ,仍然可以做出对一系列CPU(而不仅仅是Pentium 4)有利的内联和其他代码生成选择。

    推荐文章