Skip to content
Closed
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -117,7 +117,16 @@ class LegacySimpleDateFormatter(pattern: String, locale: Locale) extends LegacyD
object DateFormatter {
import LegacyDateFormats._

val defaultLocale: Locale = Locale.US
/**
* This is change from Locale.US to GB, because:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's make the doc shorter

Before Spark 3.0, the first day-of-week is always Monday. Since Spark 3.0, it depends on the locale.
We pick GB as the default locale instead of US, to be compatible with Spark 2.x, as US locale uses
Sunday as the first day-of-week. See SPARK-31879.

* The first day-of-week varies by culture.
* For example, the US uses Sunday, while the United Kingdom and the ISO-8601 standard use Monday.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this the only difference between US and en-GB?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems that we don't need care about the other differences, e.g. currency symbol
FYI,
http://www.localeplanet.com/java/en-GB/index.html and http://www.localeplanet.com/java/en-US/index.html

the timeZone is not the same, but it's a separate field we don't get it from Locale right?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is there any localized timezone related pattern letter?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have O, OOOO and ZZZZ that are localized, but there are decided by the zoneID(spark.sql.session.timeZone), not related to Locate here.

*
* Using `US` makes functions which rely on the Locale to express the first day of week
* inconsistent with Spark 2.4
* see https://issues.apache.org/jira/browse/SPARK-31879
*/
val defaultLocale: Locale = new Locale("en", "GB")

val defaultPattern: String = "yyyy-MM-dd"

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -278,7 +278,16 @@ object LegacyDateFormats extends Enumeration {
object TimestampFormatter {
import LegacyDateFormats._

val defaultLocale: Locale = Locale.US
/**
* This is change from Locale.US to GB, because:
* The first day-of-week varies by culture.
* For example, the US uses Sunday, while the United Kingdom and the ISO-8601 standard use Monday.
*
* Using `US` makes functions which rely on the Locale to express the first day of week
* inconsistent with Spark 2.4
* see https://issues.apache.org/jira/browse/SPARK-31879
*/
val defaultLocale: Locale = new Locale("en", "GB")

def defaultPattern(): String = s"${DateFormatter.defaultPattern} HH:mm:ss"

Expand Down
2 changes: 2 additions & 0 deletions sql/core/src/test/resources/sql-tests/inputs/datetime.sql
Original file line number Diff line number Diff line change
Expand Up @@ -164,3 +164,5 @@ select from_csv('26/October/2015', 'date Date', map('dateFormat', 'dd/MMMMM/yyyy
select from_unixtime(1, 'yyyyyyyyyyy-MM-dd');
select date_format(timestamp '2018-11-17 13:33:33', 'yyyyyyyyyy-MM-dd HH:mm:ss');
select date_format(date '2018-11-17', 'yyyyyyyyyyy-MM-dd');

select to_timestamp('2020-01-01', 'YYYY-ww-uu');

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we use formatting in the test? so this doesn't conflict with #28706

Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
-- Automatically generated by SQLQueryTestSuite
-- Number of queries: 119
-- Number of queries: 120


-- !query
Expand Down Expand Up @@ -1025,3 +1025,11 @@ struct<>
-- !query output
org.apache.spark.SparkUpgradeException
You may get a different result due to the upgrading of Spark 3.0: Fail to recognize 'yyyyyyyyyyy-MM-dd' pattern in the DateTimeFormatter. 1) You can set spark.sql.legacy.timeParserPolicy to LEGACY to restore the behavior before Spark 3.0. 2) You can form a valid datetime pattern with the guide from https://spark.apache.org/docs/latest/sql-ref-datetime-pattern.html


-- !query
select to_timestamp('2020-01-01', 'YYYY-ww-uu')
-- !query schema
struct<to_timestamp(2020-01-01, YYYY-ww-uu):timestamp>
-- !query output
2019-12-30 00:00:00
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
-- Automatically generated by SQLQueryTestSuite
-- Number of queries: 119
-- Number of queries: 120


-- !query
Expand Down Expand Up @@ -980,3 +980,11 @@ select date_format(date '2018-11-17', 'yyyyyyyyyyy-MM-dd')
struct<date_format(CAST(DATE '2018-11-17' AS TIMESTAMP), yyyyyyyyyyy-MM-dd):string>
-- !query output
00000002018-11-17


-- !query
select to_timestamp('2020-01-01', 'YYYY-ww-uu')
-- !query schema
struct<to_timestamp(2020-01-01, YYYY-ww-uu):timestamp>
-- !query output
2019-12-30 00:00:00
10 changes: 9 additions & 1 deletion sql/core/src/test/resources/sql-tests/results/datetime.sql.out
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
-- Automatically generated by SQLQueryTestSuite
-- Number of queries: 119
-- Number of queries: 120


-- !query
Expand Down Expand Up @@ -997,3 +997,11 @@ struct<>
-- !query output
org.apache.spark.SparkUpgradeException
You may get a different result due to the upgrading of Spark 3.0: Fail to recognize 'yyyyyyyyyyy-MM-dd' pattern in the DateTimeFormatter. 1) You can set spark.sql.legacy.timeParserPolicy to LEGACY to restore the behavior before Spark 3.0. 2) You can form a valid datetime pattern with the guide from https://spark.apache.org/docs/latest/sql-ref-datetime-pattern.html


-- !query
select to_timestamp('2020-01-01', 'YYYY-ww-uu')
-- !query schema
struct<to_timestamp(2020-01-01, YYYY-ww-uu):timestamp>
-- !query output
2019-12-30 00:00:00