Splits incoming data into batches and processes them in parallel as an asynchronous map/reduce job. Enqueues a JsExecMapReduceProcessable task via DbAsyncProcessor, which runs mapFn to split the data into batches and reduceFn to combine the results back together. Input can be a file hash string, typically produced by a PersistAsTableStep, or an InputStream of CSV or Excel data (for Excel, only the first worksheet is used), which is persisted as an uploaded table before the job is enqueued. The incoming arguments are always passed through unchanged to the next step once the job has been enqueued.
Implements: PipelineStep
Properties
| Property | Returns | Description |
|---|---|---|
| description | String | Human-readable summary of this step, shown to administrators managing the pipeline: that it processes data in parallel as an asynchronous map/reduce job. |
| formPath | String | Path to the Velocity template used by the pipeline management UI to edit this step's configuration. |
| mapFn | String | Name of the script function used to split the input into batches for parallel processing. |
| nextStep | PipelineStep | The step that the incoming arguments are passed to once the map/reduce job has been enqueued. |
| reduceFn | String | Name of the script function used to combine the results of the processed batches back together. |
Methods
getFormPath()
Returns: String
Path to the Velocity template used by the pipeline management UI to edit this step's configuration.
getDescription()
Returns: String
Human-readable summary of this step, shown to administrators managing the pipeline: that it processes data in parallel as an asynchronous map/reduce job.
getMapFn()
Returns: String
Name of the script function used to split the input into batches for parallel processing.
getReduceFn()
Returns: String
Name of the script function used to combine the results of the processed batches back together.
getNextStep()
Returns: PipelineStep
The step that the incoming arguments are passed to once the map/reduce job has been enqueued.