如何在sparkml(二进制)分类中指定“正类”?(或者可能是:一个
MulticlassClassificationEvaluator
假设我们正在训练一个模型,以达到二元分类问题中的目标精度,如。。。
label_idxer = StringIndexer(inputCol="response",
outputCol="label").fit(df_spark)
# we fit so we can get the "labels" attribute to inform reconversion stage
feature_idxer = StringIndexer(inputCols=cat_features,
outputCols=[f"{f}_IDX" for f in cat_features],
handleInvalid="keep")
onehotencoder = OneHotEncoder(inputCols=feature_idxer.getOutputCols(),
outputCols=[f"{f}_OHE" for f in feature_idxer.getOutputCols()])
assembler = VectorAssembler(inputCols=(num_features + onehotencoder.getOutputCols()),
outputCol="features")
rf = RandomForestClassifier(labelCol=label_idxer.getOutputCol(),
featuresCol=assembler.getOutputCol(),
seed=123456789)
label_converter = IndexToString(inputCol=rf.getPredictionCol(),
outputCol="prediction_label",
labels=label_idxer.labels)
pipeline = Pipeline(stages=[label_idxer, feature_idxer, onehotencoder,
assembler,
rf,
label_converter]) # type: pyspark.ml.pipeline.PipelineModel
crossval = CrossValidator(estimator=pipeline,
evaluator=MulticlassClassificationEvaluator(
labelCol=rf.getLabelCol(),
predictionCol=rf.getPredictionCol(),
metricName="weightedPrecision"),
numFolds=3)
(train_u, test_u) = dff.randomSplit([0.8, 0.2])
model = crossval.fit(train_u)
我知道。。。
Precision = TP / (TP + FP)
…但是如何将特定的类标签指定为精度目标的“正类”?(目前,IDK在培训中实际使用了哪个响应值,以及如何判断)。