代码之家  ›  专栏  ›  技术社区  ›  Alejandro Alcalde

根据Scala-flink中的另一个数据集过滤数据集

  •  1
  • Alejandro Alcalde  · 技术社区  · 8 年前

    我正在尝试复制以下python代码:

    cond_entropy_x = np.array([entropy(x[y == v]) for v in uy])
    

    x 和 y 是向量,和 uy 是的 ,例如 0,1 .

    在弗林克,我有:

    val uy = y.distinct.collect
    val condHx = for (i ← uy)
        yield entropy(x.filterWithBcVariable(y)((_, yy) ⇒ yy == i))
    

    然而,似乎 filterWithBcVariable 是的 ,只需要第一个。

    for (i ← values) yield y.join(x).where(a ⇒ a).equalTo(_ ⇒ i)
    

    但我的记性没了。

    我怎么过滤 十 是的 ?

    像这样的 x.zip(y) 会这样做,但不支持。

    有什么想法吗?

    1 回复  |  直到 8 年前
        1
  •  0
  •   Alejandro Alcalde    8 年前

    我想出了一个解决办法,也许不是最好的,但至少它起作用了。

    现在,不是过去 x y 分开的 DataSets ,我路过一个 DataSet[LabeledVector] 只有一列:

    val xy = input.map(lv ⇒ LabeledVector(lv.label, DenseVector(lv.vector(0))))
    

    xy 我的职责是:

    def conditionalEntropy(xy: DataSet[LabeledVector]): Double = {
        // Get the label
        val y = xy map (_.label)
        // Get probs for the label
        val p = probs(y).toArray.asBreeze
        // Get unique values in label
        val values = y.distinct.collect
        // Compute Conditional Entropy
        val condH = for (i ← values)
          yield entropy(xy.filter(_.label == i))
        p.dot(seq2Breeze(condH))
      }