[SPARK-25132][SQL] Case-insensitive field resolution when reading from Parquet #22148

seancxmao · 2018-08-20T03:24:49Z

What changes were proposed in this pull request?

Spark SQL returns NULL for a column whose Hive metastore schema and Parquet schema are in different letter cases, regardless of spark.sql.caseSensitive set to true or false. This PR aims to add case-insensitive field resolution for ParquetFileFormat.

Do case-insensitive resolution only if Spark is in case-insensitive mode.
Field resolution should fail if there is ambiguity, i.e. more than one field is matched.

How was this patch tested?

Unit tests added.

yucai · 2018-08-20T03:34:21Z

LGTM.
@cloud-fan @gatorsmile Could you kindly help trigger Jenkins and review?

…m Parquet

HyukjinKwon · 2018-08-20T03:59:50Z

...e/src/main/scala/org/apache/spark/sql/execution/datasources/parquet/ParquetReadSupport.scala

+          .get(f.name)
+          .map(clipParquetType(_, f.dataType, caseSensitive))
+          .getOrElse(toParquet.convertField(f))
+      }


nit: I would remove this brace per https://github.com/databricks/scala-style-guide#anonymous-methods

cloud-fan · 2018-08-20T06:05:38Z

ok to test

cloud-fan · 2018-08-20T06:09:03Z

...e/src/main/scala/org/apache/spark/sql/execution/datasources/parquet/ParquetReadSupport.scala

+    } else {
+      // Do case-insensitive resolution only if in case-insensitive mode
+      val caseInsensitiveParquetFieldMap =
+        parquetRecord.getFields.asScala.groupBy(_.getName.toLowerCase)


nit: toLowerCase(Locale.ROOT)

cloud-fan · 2018-08-20T06:12:51Z

...e/src/main/scala/org/apache/spark/sql/execution/datasources/parquet/ParquetReadSupport.scala

+            if (parquetTypes.size > 1) {
+              // Need to fail if there is ambiguity, i.e. more than one field is matched
+              val parquetTypesString = parquetTypes.map(_.getName).mkString("[", ", ", "]")
+              throw new AnalysisException(s"""Found duplicate field(s) "${f.name}": """ +


This is triggered at runtime at executor side, we should probably use RuntimeException here.

cloud-fan · 2018-08-20T06:15:49Z

sql/core/src/test/scala/org/apache/spark/sql/FileBasedDataSourceSuite.scala

+        withSQLConf(SQLConf.CASE_SENSITIVE.key -> "true") {
+          data.write.format(format).mode("overwrite").save(tableDir)
+        }
+        sql(s"CREATE TABLE $tableName (a LONG, b LONG) USING $format LOCATION '$tableDir'")


not related to this PR, but it makes me think that case-sensitivity should be a global or at least table level config, otherwise the behavior is a little confusing. cc @gatorsmile

table-level conf is reasonable. Let us do it in 3.0?

cloud-fan · 2018-08-20T06:16:31Z

LGTM except a few minor comments

SparkQA · 2018-08-20T07:05:02Z

Test build #94943 has finished for PR 22148 at commit 9261beb.

This patch fails due to an unknown error code, -9.
This patch merges cleanly.
This patch adds no public classes.

…to toLowerCase

wangyum · 2018-08-20T07:12:20Z

retest this please

HyukjinKwon

LGTM too

yucai · 2018-08-20T09:22:03Z

retest this please

SparkQA · 2018-08-20T10:07:32Z

Test build #94945 has finished for PR 22148 at commit c8279d2.

This patch fails Spark unit tests.
This patch merges cleanly.
This patch adds no public classes.

HyukjinKwon · 2018-08-20T12:15:37Z

retest this please

SparkQA · 2018-08-20T16:16:42Z

Test build #94957 has finished for PR 22148 at commit 0176d29.

This patch fails PySpark unit tests.
This patch merges cleanly.
This patch adds no public classes.

SparkQA · 2018-08-20T16:42:41Z

Test build #94954 has finished for PR 22148 at commit 0176d29.

This patch passes all tests.
This patch merges cleanly.
This patch adds no public classes.

HyukjinKwon · 2018-08-21T02:34:02Z

Merged to master.

seancxmao · 2018-08-21T05:52:42Z

Thanks!

gatorsmile · 2018-08-21T22:49:34Z

@seancxmao Please submit a follow-up PR to document the behavior changes in the migration guide of Spark SQL?

… for ORC native data source table persisted in metastore ## What changes were proposed in this pull request? Apache Spark doesn't create Hive table with duplicated fields in both case-sensitive and case-insensitive mode. However, if Spark creates ORC files in case-sensitive mode first and create Hive table on that location, where it's created. In this situation, field resolution should fail in case-insensitive mode. Otherwise, we don't know which columns will be returned or filtered. Previously, SPARK-25132 fixed the same issue in Parquet. Here is a simple example: ``` val data = spark.range(5).selectExpr("id as a", "id * 2 as A") spark.conf.set("spark.sql.caseSensitive", true) data.write.format("orc").mode("overwrite").save("/user/hive/warehouse/orc_data") sql("CREATE TABLE orc_data_source (A LONG) USING orc LOCATION '/user/hive/warehouse/orc_data'") spark.conf.set("spark.sql.caseSensitive", false) sql("select A from orc_data_source").show +---+ | A| +---+ | 3| | 2| | 4| | 1| | 0| +---+ ``` See #22148 for more details about parquet data source reader. ## How was this patch tested? Unit tests added. Closes #22262 from seancxmao/SPARK-25175. Authored-by: seancxmao <[email protected]> Signed-off-by: Dongjoon Hyun <[email protected]> (cherry picked from commit a0aed47) Signed-off-by: Dongjoon Hyun <[email protected]>

… for ORC native data source table persisted in metastore ## What changes were proposed in this pull request? Apache Spark doesn't create Hive table with duplicated fields in both case-sensitive and case-insensitive mode. However, if Spark creates ORC files in case-sensitive mode first and create Hive table on that location, where it's created. In this situation, field resolution should fail in case-insensitive mode. Otherwise, we don't know which columns will be returned or filtered. Previously, SPARK-25132 fixed the same issue in Parquet. Here is a simple example: ``` val data = spark.range(5).selectExpr("id as a", "id * 2 as A") spark.conf.set("spark.sql.caseSensitive", true) data.write.format("orc").mode("overwrite").save("/user/hive/warehouse/orc_data") sql("CREATE TABLE orc_data_source (A LONG) USING orc LOCATION '/user/hive/warehouse/orc_data'") spark.conf.set("spark.sql.caseSensitive", false) sql("select A from orc_data_source").show +---+ | A| +---+ | 3| | 2| | 4| | 1| | 0| +---+ ``` See #22148 for more details about parquet data source reader. ## How was this patch tested? Unit tests added. Closes #22262 from seancxmao/SPARK-25175. Authored-by: seancxmao <[email protected]> Signed-off-by: Dongjoon Hyun <[email protected]>

…ive field resolution when reading from Parquet ## What changes were proposed in this pull request? #22148 introduces a behavior change. According to discussion at #22184, this PR updates migration guide when upgrade from Spark 2.3 to 2.4. ## How was this patch tested? N/A Closes #23238 from seancxmao/SPARK-25132-doc-2.4. Authored-by: seancxmao <[email protected]> Signed-off-by: Dongjoon Hyun <[email protected]> (cherry picked from commit 55276d3) Signed-off-by: Dongjoon Hyun <[email protected]>

…ive field resolution when reading from Parquet ## What changes were proposed in this pull request? #22148 introduces a behavior change. According to discussion at #22184, this PR updates migration guide when upgrade from Spark 2.3 to 2.4. ## How was this patch tested? N/A Closes #23238 from seancxmao/SPARK-25132-doc-2.4. Authored-by: seancxmao <[email protected]> Signed-off-by: Dongjoon Hyun <[email protected]>

…ive field resolution when reading from Parquet ## What changes were proposed in this pull request? apache#22148 introduces a behavior change. According to discussion at apache#22184, this PR updates migration guide when upgrade from Spark 2.3 to 2.4. ## How was this patch tested? N/A Closes apache#23238 from seancxmao/SPARK-25132-doc-2.4. Authored-by: seancxmao <[email protected]> Signed-off-by: Dongjoon Hyun <[email protected]>

…ive field resolution when reading from Parquet ## What changes were proposed in this pull request? apache#22148 introduces a behavior change. According to discussion at apache#22184, this PR updates migration guide when upgrade from Spark 2.3 to 2.4. ## How was this patch tested? N/A Closes apache#23238 from seancxmao/SPARK-25132-doc-2.4. Authored-by: seancxmao <[email protected]> Signed-off-by: Dongjoon Hyun <[email protected]> (cherry picked from commit 55276d3) Signed-off-by: Dongjoon Hyun <[email protected]>

seancxmao force-pushed the SPARK-25132-Parquet branch from e0a5553 to 600c3ad Compare August 20, 2018 03:36

[SPARK-25132][SQL] Case-insensitive field resolution when reading fro…

1600190

…m Parquet

seancxmao force-pushed the SPARK-25132-Parquet branch from 600c3ad to 1600190 Compare August 20, 2018 03:41

HyukjinKwon reviewed Aug 20, 2018

View reviewed changes

seancxmao added 2 commits August 20, 2018 13:09

remove unnecessary braces of anonymous methods

ce4c935

move input parameters to the same line as map function

9261beb

cloud-fan reviewed Aug 20, 2018

View reviewed changes

use RuntimeException instead of AnalysisException;\n add Locale.ROOT …

c8279d2

…to toLowerCase

HyukjinKwon approved these changes Aug 20, 2018

View reviewed changes

ParquetSchemaTest should assertThrows[RuntimeException] also

0176d29

asfgit closed this in f984ec7 Aug 21, 2018

This was referenced Aug 22, 2018

[SPARK-25132][SQL][BACKPORT-2.3] Case-insensitive field resolution when reading from Parquet #22183

Closed

[SPARK-25132][SQL][DOC] Add migration doc for case-insensitive field resolution when reading from Parquet #22184

Closed

seancxmao mentioned this pull request Aug 29, 2018

[SPARK-25175][SQL] Field resolution should fail if there is ambiguity for ORC native data source table persisted in metastore #22262

Closed

seancxmao mentioned this pull request Sep 5, 2018

[SPARK-25391][SQL] Make behaviors consistent when converting parquet hive table to parquet data source #22343

Closed

seancxmao mentioned this pull request Dec 5, 2018

[SPARK-25132][SQL][FOLLOWUP][DOC] Add migration doc for case-insensitive field resolution when reading from Parquet #23238

Closed

cloud-fan mentioned this pull request Mar 9, 2019

[SPARK-27119][SQL] Do not infer schema when reading Hive serde table with native data source #24041

Closed

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[SPARK-25132][SQL] Case-insensitive field resolution when reading from Parquet #22148

[SPARK-25132][SQL] Case-insensitive field resolution when reading from Parquet #22148

seancxmao commented Aug 20, 2018

yucai commented Aug 20, 2018

HyukjinKwon Aug 20, 2018

cloud-fan commented Aug 20, 2018

cloud-fan Aug 20, 2018

cloud-fan Aug 20, 2018 •

edited

Loading

cloud-fan Aug 20, 2018

gatorsmile Aug 21, 2018

cloud-fan commented Aug 20, 2018

SparkQA commented Aug 20, 2018

wangyum commented Aug 20, 2018

HyukjinKwon left a comment

yucai commented Aug 20, 2018

SparkQA commented Aug 20, 2018

HyukjinKwon commented Aug 20, 2018

SparkQA commented Aug 20, 2018

SparkQA commented Aug 20, 2018

HyukjinKwon commented Aug 21, 2018

seancxmao commented Aug 21, 2018

gatorsmile commented Aug 21, 2018

[SPARK-25132][SQL] Case-insensitive field resolution when reading from Parquet #22148

[SPARK-25132][SQL] Case-insensitive field resolution when reading from Parquet #22148

Conversation

seancxmao commented Aug 20, 2018

What changes were proposed in this pull request?

How was this patch tested?

yucai commented Aug 20, 2018

HyukjinKwon Aug 20, 2018

Choose a reason for hiding this comment

cloud-fan commented Aug 20, 2018

cloud-fan Aug 20, 2018

Choose a reason for hiding this comment

cloud-fan Aug 20, 2018 • edited Loading

Choose a reason for hiding this comment

cloud-fan Aug 20, 2018

Choose a reason for hiding this comment

gatorsmile Aug 21, 2018

Choose a reason for hiding this comment

cloud-fan commented Aug 20, 2018

SparkQA commented Aug 20, 2018

wangyum commented Aug 20, 2018

HyukjinKwon left a comment

Choose a reason for hiding this comment

yucai commented Aug 20, 2018

SparkQA commented Aug 20, 2018

HyukjinKwon commented Aug 20, 2018

SparkQA commented Aug 20, 2018

SparkQA commented Aug 20, 2018

HyukjinKwon commented Aug 21, 2018

seancxmao commented Aug 21, 2018

gatorsmile commented Aug 21, 2018

cloud-fan Aug 20, 2018 •

edited

Loading