[SPARK-8357] Fix unsafe memory leak on empty inputs in GeneratedAggregate #7560

JoshRosen · 2015-07-21T06:37:21Z

This patch fixes a managed memory leak in GeneratedAggregate. The leak occurs when the unsafe aggregation path is used to perform grouped aggregation on an empty input; in this case, GeneratedAggregate allocates an UnsafeFixedWidthAggregationMap that is never cleaned up because next() is never called on the aggregate result iterator.

This patch fixes this by short-circuiting on empty inputs.

This patch is an updated version of #6810.

Closes #6810.

…ty input

JoshRosen · 2015-07-21T06:39:35Z

sql/core/src/main/scala/org/apache/spark/sql/execution/GeneratedAggregate.scala

+      if (!iter.hasNext) {
+        // This is an empty input, so return early so that we do not allocate data structures
+        // that won't be cleaned up (see SPARK-8357).
+        if (groupingExpressions.isEmpty) {


Here, I made a slight simplification compared to @navis's original patch: if groupingExpressions is empty and the input is empty, then always return an empty aggregation buffer. @navis's patch contained an additional branch here which would skip this output if partial = true, but I think that is an unnecessary performance optimization given that the non-generated-Aggregate operator still outputs an empty row even on empty inputs. Removing this branch means fewer cases to have to test.

JoshRosen · 2015-07-21T07:16:28Z

/cc @rxin, I'm planning to merge this pending tests to help unblock other work on enabling Unsafe by default.

SparkQA · 2015-07-21T08:03:03Z

Test build #37921 has finished for PR 7560 at commit 3486ce4.

This patch fails Spark unit tests.
This patch merges cleanly.
This patch adds no public classes.

JoshRosen · 2015-07-21T08:05:53Z

Huh, looks like this failed a HiveCompatibilitySuite test for a query that doesn't even use the GeneratedAggregate operator: https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder/37921/testReport/org.apache.spark.sql.hive.execution/HiveCompatibilitySuite/partcols1/

JoshRosen · 2015-07-21T08:05:56Z

Jenkins, retest this please.

SparkQA · 2015-07-21T09:27:35Z

Test build #37933 has finished for PR 7560 at commit 3486ce4.

This patch fails Spark unit tests.
This patch merges cleanly.
This patch adds the following public classes (experimental):
- trait ExpectsInputTypes extends Expression
- trait ImplicitCastInputTypes extends ExpectsInputTypes
- trait Unevaluable extends Expression
- trait Nondeterministic extends Expression
- trait CodegenFallback extends Expression

JoshRosen · 2015-07-21T15:43:36Z

Jenkins, retest this please.

JoshRosen · 2015-07-21T15:43:44Z

(Looks like master might be broken?)

SparkQA · 2015-07-21T17:24:53Z

Test build #37957 has finished for PR 7560 at commit 3486ce4.

This patch fails PySpark unit tests.
This patch merges cleanly.
This patch adds no public classes.

JoshRosen · 2015-07-21T18:42:29Z

I think that the test failure here is unrelated, so I'm going to give this one final look and then will merge it into master.

This pull request enables Unsafe mode by default in Spark SQL. In order to do this, we had to fix a number of small issues: **List of fixed blockers**: - [x] Make some default buffer sizes configurable so that HiveCompatibilitySuite can run properly (#7741). - [x] Memory leak on grouped aggregation of empty input (fixed by #7560 to fix this) - [x] Update planner to also check whether codegen is enabled before planning unsafe operators. - [x] Investigate failing HiveThriftBinaryServerSuite test. This turns out to be caused by a ClassCastException that occurs when Exchange tries to apply an interpreted RowOrdering to an UnsafeRow when range partitioning an RDD. This could be fixed by #7408, but a shorter-term fix is to just skip the Unsafe exchange path when RangePartitioner is used. - [x] Memory leak exceptions masking exceptions that actually caused tasks to fail (will be fixed by #7603). - [x] ~~https://issues.apache.org/jira/browse/SPARK-9162, to implement code generation for ScalaUDF. This is necessary for `UDFSuite` to pass. For now, I've just ignored this test in order to try to find other problems while we wait for a fix.~~ This is no longer necessary as of #7682. - [x] Memory leaks from Limit after UnsafeExternalSort cause the memory leak detector to fail tests. This is a huge problem in the HiveCompatibilitySuite (fixed by f4ac642a4e5b2a7931c5e04e086bb10e263b1db6). - [x] Tests in `AggregationQuerySuite` are failing due to NaN-handling issues in UnsafeRow, which were fixed in #7736. - [x] `org.apache.spark.sql.ColumnExpressionSuite.rand` needs to be updated so that the planner check also matches `TungstenProject`. - [x] After having lowered the buffer sizes to 4MB so that most of HiveCompatibilitySuite runs: - [x] Wrong answer in `join_1to1` (fixed by #7680) - [x] Wrong answer in `join_nulls` (fixed by #7680) - [x] Managed memory OOM / leak in `lateral_view` - [x] Seems to hang indefinitely in `partcols1`. This might be a deadlock in script transformation or a bug in error-handling code? The hang was fixed by #7710. - [x] Error while freeing memory in `partcols1`: will be fixed by #7734. - [x] After fixing the `partcols1` hang, it appears that a number of later tests have issues as well. - [x] Fix thread-safety bug in codegen fallback expression evaluation (#7759). Author: Josh Rosen <[email protected]> Closes #7564 from JoshRosen/unsafe-by-default and squashes the following commits: 83c0c56 [Josh Rosen] Merge remote-tracking branch 'origin/master' into unsafe-by-default f4cc859 [Josh Rosen] Merge remote-tracking branch 'origin/master' into unsafe-by-default 963f567 [Josh Rosen] Reduce buffer size for R tests d6986de [Josh Rosen] Lower page size in PySpark tests 013b9da [Josh Rosen] Also match TungstenProject in checkNumProjects 5d0b2d3 [Josh Rosen] Add task completion callback to avoid leak in limit after sort ea250da [Josh Rosen] Disable unsafe Exchange path when RangePartitioning is used 715517b [Josh Rosen] Enable Unsafe by default

navis and others added 13 commits June 29, 2015 09:44

[SPARK-8357] [SQL] Memory leakage on unsafe aggregation path with emp…

1b07556

…ty input

added comments

d396589

added a test as suggested by JoshRosen

15c5afc

fixed test fails

4d326b9

addressed comments

51178e8

Rolled-back test-conf cleanup & fixed possible CCE & added more tests

1a02a55

used new conf apis

735972f

fixed format & added test for CCE case

143e1ef

addressed comments

c5419b3

Back out Projection changes.

adc8239

Merge remote-tracking branch 'origin/master' into SPARK-8357

3c7db0f

Revert SparkPlan change:

c649310

Some minor cleanup

3486ce4

JoshRosen reviewed Jul 21, 2015
View reviewed changes

This was referenced Jul 21, 2015

[SPARK-8357] [SQL] Memory leakage on unsafe aggregation path with empty input #6810

Closed

[SPARK-8850] [SQL] Enable Unsafe mode by default #7564

Closed

asfgit closed this in 9ba7c64 Jul 21, 2015

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[SPARK-8357] Fix unsafe memory leak on empty inputs in GeneratedAggregate #7560

[SPARK-8357] Fix unsafe memory leak on empty inputs in GeneratedAggregate #7560

JoshRosen commented Jul 21, 2015

JoshRosen Jul 21, 2015

JoshRosen commented Jul 21, 2015

SparkQA commented Jul 21, 2015

JoshRosen commented Jul 21, 2015

JoshRosen commented Jul 21, 2015

SparkQA commented Jul 21, 2015

JoshRosen commented Jul 21, 2015

JoshRosen commented Jul 21, 2015

SparkQA commented Jul 21, 2015

JoshRosen commented Jul 21, 2015

[SPARK-8357] Fix unsafe memory leak on empty inputs in GeneratedAggregate #7560

[SPARK-8357] Fix unsafe memory leak on empty inputs in GeneratedAggregate #7560

Conversation

JoshRosen commented Jul 21, 2015

JoshRosen Jul 21, 2015

Choose a reason for hiding this comment

JoshRosen commented Jul 21, 2015

SparkQA commented Jul 21, 2015

JoshRosen commented Jul 21, 2015

JoshRosen commented Jul 21, 2015

SparkQA commented Jul 21, 2015

JoshRosen commented Jul 21, 2015

JoshRosen commented Jul 21, 2015

SparkQA commented Jul 21, 2015

JoshRosen commented Jul 21, 2015