-
Notifications
You must be signed in to change notification settings - Fork 496
Add polaris spark client webpage under unreleased #1503
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from 5 commits
888ba37
ae805ef
bf27cae
f448b6d
0b4ee98
842c132
2d19c57
492459d
f898e93
6d09e16
8c16437
d1d6e8a
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -30,6 +30,12 @@ and depends on iceberg-spark-runtime 1.8.1. | |
|
|
||
| # Build Plugin Jar | ||
| A task createPolarisSparkJar is added to build a jar for the Polaris Spark plugin, the jar is named as: | ||
| `polaris-iceberg-<icebergVersion>-spark-runtime-<sparkVersion>_<scalaVersion>-<polarisVersion>.jar`. For example: | ||
| `polaris-iceberg-1.8.1-spark-runtime-3.5_2.12-0.10.0-beta-incubating-SNAPSHOT.jar`. | ||
|
|
||
| - `./gradlew :polaris-spark-3.5_2.12:createPolarisSparkJar` -- build jar for Spark 3.5 with Scala version 2.12. | ||
| - `./gradlew :polaris-spark-3.5_2.13:createPolarisSparkJar` -- build jar for Spark 3.5 with Scala version 2.13. | ||
|
|
||
| The result jar is located at plugins/spark/v3.5/build/<scala_version>/libs after the build. | ||
|
|
||
| # Start Spark with Local Polaris Service using built Jar | ||
|
|
@@ -51,13 +57,12 @@ bin/spark-shell \ | |
| --conf spark.sql.extensions=org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions,io.delta.sql.DeltaSparkSessionExtension \ | ||
| --conf spark.sql.catalog.spark_catalog=org.apache.spark.sql.delta.catalog.DeltaCatalog \ | ||
| --conf spark.sql.catalog.<catalog-name>.warehouse=<catalog-name> \ | ||
| --conf spark.sql.catalog.<catalog-name>.header.X-Iceberg-Access-Delegation=true \ | ||
| --conf spark.sql.catalog.<catalog-name>.header.X-Iceberg-Access-Delegation=vended-credentials \ | ||
| --conf spark.sql.catalog.<catalog-name>=org.apache.polaris.spark.SparkCatalog \ | ||
| --conf spark.sql.catalog.<catalog-name>.uri=http://localhost:8181/api/catalog \ | ||
| --conf spark.sql.catalog.<catalog-name>.credential="root:secret" \ | ||
| --conf spark.sql.catalog.<catalog-name>.scope='PRINCIPAL_ROLE:ALL' \ | ||
| --conf spark.sql.catalog.<catalog-name>.token-refresh-enabled=true \ | ||
| --conf spark.sql.catalog.<catalog-name>.type=rest \ | ||
| --conf spark.sql.sources.useV1SourceList='' | ||
| ``` | ||
|
|
||
|
|
@@ -72,24 +77,26 @@ bin/spark-shell \ | |
| --conf spark.sql.extensions=org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions,io.delta.sql.DeltaSparkSessionExtension \ | ||
| --conf spark.sql.catalog.spark_catalog=org.apache.spark.sql.delta.catalog.DeltaCatalog \ | ||
| --conf spark.sql.catalog.polaris.warehouse=<catalog-name> \ | ||
| --conf spark.sql.catalog.polaris.header.X-Iceberg-Access-Delegation=true \ | ||
| --conf spark.sql.catalog.polaris.header.X-Iceberg-Access-Delegation=vended-credentials \ | ||
| --conf spark.sql.catalog.polaris=org.apache.polaris.spark.SparkCatalog \ | ||
| --conf spark.sql.catalog.polaris.uri=http://localhost:8181/api/catalog \ | ||
| --conf spark.sql.catalog.polaris.credential="root:secret" \ | ||
| --conf spark.sql.catalog.polaris.scope='PRINCIPAL_ROLE:ALL' \ | ||
| --conf spark.sql.catalog.polaris.token-refresh-enabled=true \ | ||
| --conf spark.sql.catalog.polaris.type=rest \ | ||
| --conf spark.sql.sources.useV1SourceList='' | ||
| ``` | ||
|
|
||
| # Limitations | ||
| The Polaris Spark client supports catalog management for both Iceberg and Delta tables, it routes all Iceberg table | ||
| requests to the Iceberg REST endpoints, and routes all Delta table requests to the Generic Table REST endpoints. | ||
|
|
||
| Following describes the current limitations of the Polaris Spark client: | ||
| The Spark Client requires at least delta 3.2.1 to work with Delta tables, which requires at least Apache Spark 3.5.3. | ||
| Following describes the current functionality limitations of the Polaris Spark client: | ||
| 1) Create table as select (CTAS) is not supported for Delta tables. As a result, the `saveAsTable` method of `Dataframe` | ||
| is also not supported, since it relies on the CTAS support. | ||
| 2) Create a Delta table without explicit location is not supported. | ||
| 3) Rename a Delta table is not supported. | ||
| 4) ALTER TABLE ... SET LOCATION/SET FILEFORMAT/ADD PARTITION is not supported for DELTA table. | ||
| 5) For other non-iceberg tables like csv, there is no specific guarantee provided today. | ||
| 5) For other non-Iceberg tables like csv, there is no specific guarantee provided today. | ||
| 6) TABLE_WRITE_DATA privilege is not supported for Delta Table. | ||
| 7) Credential Vending is not supported for Delta Table. | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Does this imply that writes to Delta Tables are not supported?
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. It's more complicated than that; it implies exactly what it says. With Delta, the catalog doesn't take over the responsibility of the writes. But that doesn't mean that the client can't write to the Delta table. However, it can't use vended credentials to do so.
Contributor
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Actually, those are our Polaris service limitations, not really client limitations, it might be better to introduce another generic table support page and put that limitation there. I removed it from the polaris spark client for now |
||
| Original file line number | Diff line number | Diff line change | ||||
|---|---|---|---|---|---|---|
| @@ -0,0 +1,156 @@ | ||||||
| --- | ||||||
| # | ||||||
| # Licensed to the Apache Software Foundation (ASF) under one | ||||||
| # or more contributor license agreements. See the NOTICE file | ||||||
| # distributed with this work for additional information | ||||||
| # regarding copyright ownership. The ASF licenses this file | ||||||
| # to you under the Apache License, Version 2.0 (the | ||||||
| # "License"); you may not use this file except in compliance | ||||||
| # with the License. You may obtain a copy of the License at | ||||||
| # | ||||||
| # http://www.apache.org/licenses/LICENSE-2.0 | ||||||
| # | ||||||
| # Unless required by applicable law or agreed to in writing, | ||||||
| # software distributed under the License is distributed on an | ||||||
| # "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY | ||||||
| # KIND, either express or implied. See the License for the | ||||||
| # specific language governing permissions and limitations | ||||||
| # under the License. | ||||||
| # | ||||||
| Title: Polaris Spark Client | ||||||
| type: docs | ||||||
| weight: 400 | ||||||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. nit: "Entities" page is also at weight 400. We should, ideally, not have multiple pages at the same weight so that it's not confusing. But in other thoughts: Do we think this belongs between "Entities" and "Telemetry" in the drop down? Personally, I think putting it between "Configuring Polaris" and "Deploying in Production" makes more sense. Or even creating a new folder for how to use new features.
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. +1 to a new folder
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. We could create a new folder once we got more than one client, e.g. spark and Trino clients.
Contributor
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. I think a new folder seems to much for now, i moved after Deploying in Production, since the spark client can only be used after the polaris is deployed |
||||||
| --- | ||||||
|
|
||||||
| Apache Polaris now provides Catalog support for Generic Tables (non-iceberg tables), please check out | ||||||
| the [Catalog API Spec]({{% ref "polaris-catalog-service" %}}) for Generic Table API specs. | ||||||
|
|
||||||
| Along with the Generic Table Catalog support, Polaris is also releasing a Spark Client, which helps to | ||||||
| provide an end-to-end solution for Apache Spark to manage Delta tables using Polaris. | ||||||
|
|
||||||
| Note the Polaris Spark Client is able to handle both Iceberg and Delta tables, not just Delta. | ||||||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Q: do parquet/csv tables work?
Contributor
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. I briefly tested scv manually before, basic operations like create, insert, drop works, alter doesn't work. I haven't tested parquet yet.
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. We should probably file an issue for this |
||||||
|
|
||||||
| This page documents how to build and use the Polaris Spark Client directly with the source repo. | ||||||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Can we talk about how to use the jar(line 89 to line 145) first, then goes to build details? I believe most of users mainly care about the usage once we release the client jar.
Contributor
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. I adjusted the order, and have section Start Spark against a deployed Polaris serviceand Connecting with Spark using local Polaris Spark client |
||||||
|
|
||||||
| ## Prerequisite | ||||||
| 1. Check out the polaris repo | ||||||
| ```shell | ||||||
| cd ~ | ||||||
| git clone https://github.com/apache/polaris.git | ||||||
| ``` | ||||||
| 2. Spark with version >= 3.5.3 and <= 3.5.5, recommended with 3.5.5. | ||||||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more.
Suggested change
Contributor
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. This is not needed for Quick Start, so i moved it under 'Start Spark against a deployed Polaris service' now |
||||||
| ```shell | ||||||
| cd ~ | ||||||
| wget https://archive.apache.org/dist/spark/spark-3.5.5/spark-3.5.5-bin-hadoop3.tgz | ||||||
| mkdir spark-3.5 | ||||||
| tar xzvf spark-3.5.5-bin-hadoop3.tgz -C spark-3.5 --strip-components=1 | ||||||
| cd spark-3.5 | ||||||
| ``` | ||||||
|
|
||||||
| All Spark Client code is available under `plugins/spark` of the polaris repo. | ||||||
|
|
||||||
| ## Quick Start with Local Polaris Service | ||||||
| If you want to quickly try out the functionality with a local Polaris service, you can follow the instructions | ||||||
| in `plugins/spark/v3.5/getting-started/README.md`. | ||||||
|
|
||||||
| The getting-started will start two containers: | ||||||
| 1) The `polaris` service for running Apache Polaris using an in-memory metastore | ||||||
| 2) The `jupyter` service for running Jupyter notebook with PySpark (Spark 3.5.5 is used) | ||||||
|
|
||||||
| The notebook `SparkPolaris.ipynb` provided under `plugins/spark/v3.5/getting-started/notebooks` provides examples | ||||||
| with basic commands, includes: | ||||||
| 1) Connect to Polaris using Python client to create a Catalog and Roles | ||||||
| 2) Start Spark session using the Polaris Spark Client | ||||||
| 3) Using Spark to perform table operations for both Delta and Iceberg | ||||||
|
|
||||||
| ## Start Spark against a deployed Polaris Service | ||||||
| If you want to start Spark with a deployed Polaris service, you can follow the instructions below. | ||||||
|
|
||||||
| Before starting, make sure the service deployed is up-to-date, and that Spark 3.5 with at least version 3.5.3 is installed. | ||||||
|
|
||||||
| ### Build Spark Client Jars | ||||||
| The polaris-spark project provides a task createPolarisSparkJar to help building jars for the Polaris Spark client, | ||||||
| The built jar is named as: | ||||||
| `polaris-iceberg-<icebergVersion>-spark-runtime-<sparkVersion>_<scalaVersion>-<polarisVersion>.jar`. | ||||||
|
|
||||||
| For example: `polaris-iceberg-1.8.1-spark-runtime-3.5_2.12-0.10.0-beta-incubating-SNAPSHOT.jar`. | ||||||
|
|
||||||
| Run the following commands to build a Spark Client jar that is compatible with Spark 3.5 and Scala 2.12. | ||||||
| ```shell | ||||||
| cd ~/polaris | ||||||
| ./gradlew :polaris-spark-3.5_2.12:createPolarisSparkJar | ||||||
| ``` | ||||||
| If you want to build a Scala 2.13 compatible jar, you can use the following command: | ||||||
| - `./gradlew :polaris-spark-3.5_2.13:createPolarisSparkJar` | ||||||
|
|
||||||
| The result jar is located at `plugins/spark/v3.5/build/<scala_version>/libs` after the build. You can also copy the | ||||||
| corresponding jar to any location your Spark will have access. | ||||||
|
|
||||||
| ### Connecting with Spark Using the built jar | ||||||
| The following CLI command can be used to start the spark with connection to the deployed Polaris service using | ||||||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. "start Apache Spark with a connection"
Contributor
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Capitalized the Spark, since I have Apache Spark at the very beginning, i don't think we need to repeat Apache Spark everywhere |
||||||
| the Polaris Spark client jar. | ||||||
|
|
||||||
| ```shell | ||||||
| bin/spark-shell \ | ||||||
| --jars <path-to-spark-client-jar> \ | ||||||
| --packages org.apache.hadoop:hadoop-aws:3.4.0,io.delta:delta-spark_2.12:3.3.1 \ | ||||||
| --conf spark.sql.extensions=org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions,io.delta.sql.DeltaSparkSessionExtension \ | ||||||
| --conf spark.sql.catalog.spark_catalog=org.apache.spark.sql.delta.catalog.DeltaCatalog \ | ||||||
| --conf spark.sql.catalog.<spark-catalog-name>.warehouse=<polaris-catalog-name> \ | ||||||
| --conf spark.sql.catalog.<spark-catalog-name>.header.X-Iceberg-Access-Delegation=vended-credentials \ | ||||||
| --conf spark.sql.catalog.<spark-catalog-name>=org.apache.polaris.spark.SparkCatalog \ | ||||||
| --conf spark.sql.catalog.<spark-catalog-name>.uri=<polaris-service-uri> \ | ||||||
| --conf spark.sql.catalog.<spark-catalog-name>.credential='<client-id>:<client-secret>' \ | ||||||
| --conf spark.sql.catalog.<spark-catalog-name>.scope='PRINCIPAL_ROLE:ALL' \ | ||||||
| --conf spark.sql.catalog.<spark-catalog-name>.token-refresh-enabled=true | ||||||
| ``` | ||||||
|
|
||||||
| Replace `path-to-spark-client-jar` to where the built jar is located. The `spark-catalog-name` is the catalog name you | ||||||
| will use with spark, and `polaris-catalog-name` is the catalog name used by Polaris service, for simplicity, you can use | ||||||
| the same name. Replace the `polaris-service-uri`, `client-id` and `client-secret` accordingly, you can refer to | ||||||
| [Using Polaris]({{% ref "getting-started/using-polaris" %}}) for more details about those fields. | ||||||
|
|
||||||
| Or you can create a spark session start the connection, following is an example with pyspark | ||||||
| ```python | ||||||
| from pyspark.sql import SparkSession | ||||||
|
|
||||||
| spark = SparkSession.builder | ||||||
| .config("spark.jars", <path-to-spark-client-jar>) | ||||||
| .config("spark.jars.packages", "org.apache.hadoop:hadoop-aws:3.3.4,io.delta:delta-spark_2.12:3.3.1") | ||||||
| .config("spark.sql.catalog.spark_catalog", "org.apache.spark.sql.delta.catalog.DeltaCatalog") | ||||||
| .config("spark.sql.extensions", "org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions,io.delta.sql.DeltaSparkSessionExtension") | ||||||
| .config("spark.sql.catalog.<spark-catalog-name>", "org.apache.polaris.spark.SparkCatalog") | ||||||
| .config("spark.sql.catalog.<spark-catalog-name>.uri", <polaris-service-uri>) | ||||||
| .config("spark.sql.catalog.<spark-catalog-name>.token-refresh-enabled", "true") | ||||||
| .config("spark.sql.catalog.<spark-catalog-name>.credential", "<client-id>:<client_secret>") | ||||||
| .config("spark.sql.catalog.<spark-catalog-name>.warehouse", <polaris_catalog_name>) | ||||||
| .config("spark.sql.catalog.polaris.scope", 'PRINCIPAL_ROLE:ALL') | ||||||
| .config("spark.sql.catalog.polaris.header.X-Iceberg-Access-Delegation", 'vended-credentials') | ||||||
| .getOrCreate() | ||||||
| ``` | ||||||
| Similar as the CLI command, make sure the corresponding fields are replaced correctly. | ||||||
|
|
||||||
| ### Create tables with Spark | ||||||
| After the spark is started, you can use it to create and access Iceberg and Delta table like what you are doing before, | ||||||
| for example: | ||||||
| ```python | ||||||
| spark.sql("USE polaris") | ||||||
| spark.sql("CREATE NAMESPACE IF NOT EXISTS DELTA_NS") | ||||||
| spark.sql("CREATE NAMESPACE IF NOT EXISTS DELTA_NS.PUBLIC") | ||||||
| spark.sql("USE NAMESPACE DELTA_NS.PUBLIC") | ||||||
| spark.sql("""CREATE TABLE IF NOT EXISTS PEOPLE ( | ||||||
| id int, name string) | ||||||
| USING delta LOCATION 'file:///tmp/delta_tables/people'; | ||||||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. In general, we've switched to using "/var/tmp/" instead of "/tmp/" in Getting Started due to tmp dir's GC sometimes being quick to dump |
||||||
| """) | ||||||
| ``` | ||||||
|
|
||||||
| ## Limitations | ||||||
| The Polaris Spark client has the following functionality limitations: | ||||||
| 1) Create table as select (CTAS) is not supported for Delta tables. As a result, the `saveAsTable` method of `Dataframe` | ||||||
| is also not supported, since it relies on the CTAS support. | ||||||
| 2) Create a Delta table without explicit location is not supported. | ||||||
| 3) Rename a Delta table is not supported. | ||||||
| 4) ALTER TABLE ... SET LOCATION/SET FILEFORMAT/ADD PARTITION is not supported for DELTA table. | ||||||
| 5) For other non-Iceberg tables like csv, there is no specific guarantee provided today. | ||||||
| 6) TABLE_WRITE_DATA privileges is not supported for Delta Table. | ||||||
| 7) Credential Vending is not supported for Delta Table. | ||||||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
What does this mean? If we mean that it may work but we don't officially support it, let's word it accordingly.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
What does "officially support" mean?
It may, or may not work... we don't provide any guarantee about whether or not it works. I think we might just want to not even mention it.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
I think let me make it very explicit that it is not supported for now.