Feature Transformation -- Imputer (Estimator)

Imputation estimator for completing missing values, either using the mean or the median of the columns in which the missing values are located. The input columns should be of numeric type. This function requires Spark 2.2.0+.

ft_imputer(x, input_cols = NULL, output_cols = NULL,
  missing_value = NULL, strategy = "mean",
  uid = random_string("imputer_"), ...)

Arguments

x

A spark_connection, ml_pipeline, or a tbl_spark.

input_cols

The names of the input columns

output_cols

The names of the output columns.

missing_value

The placeholder for the missing values. All occurrences of missing_value will be imputed. Note that null values are always treated as missing.

strategy

The imputation strategy. Currently only "mean" and "median" are supported. If "mean", then replace missing values using the mean value of the feature. If "median", then replace missing values using the approximate median value of the feature. Default: mean

uid

A character string used to uniquely identify the feature transformer.

...

Optional arguments; currently unused.

Value

The object returned depends on the class of x.

  • spark_connection: When x is a spark_connection, the function returns a ml_transformer, a ml_estimator, or one of their subclasses. The object contains a pointer to a Spark Transformer or Estimator object and can be used to compose Pipeline objects.

  • ml_pipeline: When x is a ml_pipeline, the function returns a ml_pipeline with the transformer or estimator appended to the pipeline.

  • tbl_spark: When x is a tbl_spark, a transformer is constructed then immediately applied to the input tbl_spark, returning a tbl_spark

Details

In the case where x is a tbl_spark, the estimator fits against x to obtain a transformer, which is then immediately used to transform x, returning a tbl_spark.

See also