代码之家  ›  专栏  ›  技术社区  ›  Marc

Bucketing算法

  •  3
  • Marc  · 技术社区  · 17 年前

    我有一些代码可以工作,但有点瓶颈,我一直在努力找出如何加速它。它在一个循环中,我不知道如何矢量化它。

    我有一个2D数组VAL,它表示timeseries数据。行是日期,列是不同的序列。我试图按月对数据进行存储,以便对其执行各种操作(求和、平均值等)。这是我目前的代码:

    allDts; %Dates/times for vals.  Size is [size(vals, 1), 1]
    vals;
    [Y M] = datevec(allDts);
    fomDates = unique(datenum(Y, M, 1)); %first of the month dates
    
    [Y M] = datevec(fomDates);
    nextFomDates = datenum(Y, M, DateUtil.monthLength(Y, M)+1);
    
    newVals = nan(length(fomDates), size(vals, 2)); %preallocate for speed
    
    for k = 1:length(fomDates);
    

        idx = (allDts >= fomDates(k)) & (allDts < nextFomDates(k));
        bucketed = vals(idx, :);
        newVals(k, :) = nansum(bucketed);
    end %for
    

    有什么想法吗?提前谢谢。

    2 回复  |  直到 14 年前
        1
  •  2
  •   Community Mohan Dere    9 年前

    这是一个很难矢量化的问题。我可以建议一种使用 CELLFUN this other SO question 总是 工作速度比循环快。这可能是非常具体的问题,这是最好的选择。有了这个免责声明,我将建议您尝试两种解决方案:CELLFUN版本和对for-loop版本的修改,这两种版本可能运行得更快。

    [Y,M] = datevec(allDts);
    monthStart = datenum(Y,M,1);  % Start date of each month
    [monthStart,sortIndex] = sort(monthStart);  % Sort the start dates
    [uniqueStarts,uniqueIndex] = unique(monthStart);  % Get unique start dates
    
    valCell = mat2cell(vals(sortIndex,:),diff([0 uniqueIndex]));
    newVals = cellfun(@nansum,valCell,'UniformOutput',false);
    

    呼吁 MAT2CELL 瓦尔斯 瓦尔塞尔 . 变量 将是长度为的单元格数组 努梅尔(唯一开始) 南森 关于矩阵的对应单元 瓦尔塞尔 .

    FOR-LOOP解决方案:

    [Y,M] = datevec(allDts);
    monthStart = datenum(Y,M,1);  % Start date of each month
    [monthStart,sortIndex] = sort(monthStart);  % Sort the start dates
    [uniqueStarts,uniqueIndex] = unique(monthStart);  % Get unique start dates
    
    vals = vals(sortIndex,:);  % Sort the values according to start date
    nMonths = numel(uniqueStarts);
    uniqueIndex = [0 uniqueIndex];
    newVals = nan(nMonths,size(vals,2));  % Preallocate
    for iMonth = 1:nMonths,
      index = (uniqueIndex(iMonth)+1):uniqueIndex(iMonth+1);
      newVals(iMonth,:) = nansum(vals(index,:));
    end