> For the complete documentation index, see [llms.txt](https://docs.pentaho.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.pentaho.com/install/9.3-install/pentaho-configuration/tasks-to-be-performed-by-an-it-administrator/set-up-the-adaptive-execution-layer-ael/advanced-topics/spark-tuning-landing-page-cp/configuring-application-tuning-parameters-for-spark/optimizing-spark-tuning/step-1-set-the-spark-parameters-on-the-cluster.md).

# Step 1: Set the Spark parameters on the cluster

If possible, ensure that no other jobs are running on the cluster.

Use the following steps to optimize Spark tuning globally.

1. Set the Spark parameters as described in [Set the Spark parameters globally](/install/9.3-install/pentaho-configuration/tasks-to-be-performed-by-an-it-administrator/set-up-the-adaptive-execution-layer-ael/advanced-topics/spark-tuning-landing-page-cp/configuring-application-tuning-parameters-for-spark/spark-tuning-process/application-tuning-parameters-for-transformations/set-the-spark-parameters-globally.md).
2. Run a single step PDI transformation on the cluster using a small number of executors per node and record the number of minutes it takes for the run to complete.
3. Increment the number of executors per node by 1, and then rerun the PDI transformation and record the time it takes to complete.
4. Repeat step 3 for as many executors per node that you want to verify.

   The following table shows an example of the results using the **Sort** PDI transformation step.

   | PDI step | Run number | Executors per node | Job duration |
   | -------- | ---------- | ------------------ | ------------ |
   | **Sort** | 1          | `2`                | 37 minutes   |
   | **Sort** | 2          | `3`                | 42 minutes   |
   | **Sort** | 3          | `4`                | 38 minutes   |
   | **Sort** | 4          | `5`                | 40 minutes   |
5. Evaluate the results of the runs, then choose the fastest, most repeatable run.

   The performance results of your executed transformations are available on the YARN ResourceManager and Spark History Server.
6. Set the values for the Spark parameters in the `application.properties` file according to the findings in step 5.
7. Rerun the transformation with the selected value several times to verify the results.

   The global tuning for the Spark application is recorded in the **Logging** tab in the Execution Results panel of PDI.

   ![PDI logging of Spark application parameters](/files/5wpSsbjFwSKRFDYYXHkd)

Spark parameters are set on the cluster or Pentaho Server as a baseline and apply to all users and all transformations. If needed, proceed to [Step 2: Adjust the Spark parameters in the transformation](/install/9.3-install/pentaho-configuration/tasks-to-be-performed-by-an-it-administrator/set-up-the-adaptive-execution-layer-ael/advanced-topics/spark-tuning-landing-page-cp/configuring-application-tuning-parameters-for-spark/optimizing-spark-tuning/step-2-adjust-the-spark-parameters-in-the-transformation.md) to tune Spark for your transformation. For example, additional tuning may be required to run the KTR if it runs slowly or consumes excessive resources.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.pentaho.com/install/9.3-install/pentaho-configuration/tasks-to-be-performed-by-an-it-administrator/set-up-the-adaptive-execution-layer-ael/advanced-topics/spark-tuning-landing-page-cp/configuring-application-tuning-parameters-for-spark/optimizing-spark-tuning/step-1-set-the-spark-parameters-on-the-cluster.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
