feat: support array_insert by SemyonSinchenko · Pull Request #1073 · apache/datafusion-comet

SemyonSinchenko · 2024-11-09T16:13:44Z

Which issue does this PR close?

Related to #1042

array_insert: SELECT array_insert(array(1, 2, 3, 4), 5, 5)

Rationale for this change

As described in #1042

What changes are included in this PR?

QueryPlanSerde.scala: I added an additional case for the array insert;
expr.proto: I added a new message for the ArrayInsert;
planner.rs: I added a case for the array_insert;
list.rs:
- I added a new ArrayInsert struct;
- I implemented PhysicalExpr, Display and PartialExpr for it;
- The main logic of insertion is in fn array_insert

How are these changes tested?

At the moment I added a simple tests for fn array_insert and a test for QueryPlanSerde.

SemyonSinchenko · 2024-11-09T17:47:44Z

As I was able to realize, array_insert does not supported in datafusion. Is the list.rs a good place to have an implementation of ArrayInsert and PhysicalExpr for it?

SemyonSinchenko · 2024-11-11T14:34:56Z

@andygrove Sorry for tagging but I have questions about the ticket (array_insert).

[RESOLVED] array_insert was added in spark 3.4, so all the 3.3.x tests are obviously failed. I checked and it looks like the EoL for 3.3 is about the end of 2024. Technically I think I can try to workaround tests in 3.3.x by reflection API, my question is mostly should I do it due to soon EoL of the 3.3.x?
array_insert is not supported in DataFusion. I made an implementation (and it looks like it works, except negative indices and corner cases). Is the list.rs a good place for it? Or should I move my code somewhere else?
Spark does not support anything except Int32 for position argument, is it OK if I will support only int32 too? In theory, other types can be supported too, but I'm still trying to realize how to achieve it and it may become complex...

Thanks in advance! That is my first serious attempt to contribute to the project, so sorry If I'm annoying.

+ fix tests for spark < 3.4

- added test for the negative index - added test for the legacy spark mode

andygrove · 2024-11-13T22:00:06Z

Thanks, @SemyonSinchenko. I think it's fine to skip the test for Spark 3.3. I plan on reviewing this PR in more detail tomorrow, but it looks good from an initial read.

codecov-commenter · 2024-11-14T06:58:20Z

Codecov Report

Attention: Patch coverage is 59.09091% with 9 lines in your changes missing coverage. Please review.

Project coverage is 34.27%. Comparing base (845b654) to head (8b58d8d).
Report is 18 commits behind head on main.

Files with missing lines	Patch %	Lines
.../scala/org/apache/comet/serde/QueryPlanSerde.scala	59.09%	7 Missing and 2 partials ⚠️

Additional details and impacted files

@@             Coverage Diff              @@
##               main    #1073      +/-   ##
============================================
- Coverage     34.46%   34.27%   -0.20%     
  Complexity      888      888              
============================================
  Files           113      113              
  Lines         43580    43355     -225     
  Branches       9658     9488     -170     
============================================
- Hits          15021    14860     -161     
- Misses        25507    25596      +89     
+ Partials       3052     2899     -153

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

SemyonSinchenko · 2024-11-16T13:42:15Z

This PR is ready for the review. The only failed check failed due to the internal GHA error:

GitHub Actions has encountered an internal error when running your job.

Co-authored-by: Andy Grove <agrove@apache.org>

NoeB · 2024-11-19T07:34:09Z

I am not sure if this should be done together with this PR but it would some add "free" tests. Spark introduced with 3.5 array_prepend which it implements with array_insert and starting 4.0 it also implements array_append with array_insert. If you want you can copy the array_append tests and replace array_append with array_prepend for Spark 3.5+ and enable the array_append tests for spark 4.0. You can also ignore this comment if you do not agree or if it leads to unrelated errors.

- fixes; - tests; - comments in the code;

In one case there is a zero in index and test fails due to spark error

SemyonSinchenko · 2024-11-20T18:01:58Z

Thanks for comments and suggestions!

This PR is ready for the review again.

What were changed from the last round of the review:

I added tests for the array_prepend (spark 3.5+) and enabled the test for array_append (spark 4.0+);
I added tests for the corner cases:
- Index is negative;
- Index is bigger than the array length;
- Index is negative and it's abs is greater than the array length;
- Test of the fallback to spark (udf as child);
- Value to insert is actually null;
I fixed the native part behavior for some of the corner cases;

These tests pointed my attention to the uncovered parts of the code that I fixed and I also added couple of additional tests to the native part. I revisited how spark do array_insert and it is more tricky than I realized at the first glance. I added few additional comments to the native part of the implementation and fixed the behavior.

At the moment this PR is tested in multiple ways:

Basic tests in native part that are useful because easy to debug and very fast to run;
Basic tests in native part for the so called "legacy mode" in spark;
Basic tests on the scala side;
Logical tests on small data for corner cases (negative, positive, long, short, null, etc.) on the scala side;
Tests in array_prepend and array_append that are calling array_insert under the hood (NULLS, different data types, etc.);
Test of the fallback to the Spark in case when one of children is not supported by Comet;

So, it looks to me, that all the possible cases are covered and the behavior is the same like in spark.

SemyonSinchenko · 2024-11-20T18:02:14Z

I am not sure if this should be done together with this PR but it would some add "free" tests. Spark introduced with 3.5 array_prepend which it implements with array_insert and starting 4.0 it also implements array_append with array_insert. If you want you can copy the array_append tests and replace array_append with array_prepend for Spark 3.5+ and enable the array_append tests for spark 4.0. You can also ignore this comment if you do not agree or if it leads to unrelated errors.

Done!

andygrove · 2024-11-20T22:45:58Z

+        let src_element_type = match src_value.data_type() {
+            DataType::List(field) => field.data_type(),
+            DataType::LargeList(field) => field.data_type(),
+            data_type => {
+                return Err(DataFusionError::Internal(format!(
+                    "Unexpected src array type in ArrayInsert: {:?}",
+                    data_type
+                )))
+            }


minor nit: this logic for extracting a list type is repeated a few times and could be factored out into a function

@andygrove Thanks for the suggestion!
I moved a checking of the array type (and the exception logic) to the method:

pub fn array_type(&self, data_type: &DataType) -> DataFusionResult<DataType> { match data_type { DataType::List(field) => Ok(DataType::List(Arc::clone(field))), DataType::LargeList(field) => Ok(DataType::LargeList(Arc::clone(field))), data_type => { return Err(DataFusionError::Internal(format!( "Unexpected src array type in ArrayInsert: {:?}", data_type ))) } } }

It allows at least to avoid returning the same error multiple time. Is it what you suggested? Or should I move this method to a helper function and refactor also GerArrayStructField to use such a function?

P.S. Sorry for the stupid question... But can you please explain to me why we always check both List and LargeList, while Apache Spark only supports i32 indexes for arrays (max length is Integer.MAX_VALUE - 15), which is the case of List to my understanding? All the code in the list.rs might become a bit simpler if we make it non-generic (it also makes implementation of other missing methods like array_zip simpler).

That's a good question. I wonder if the existing code for LargeList is actually being tested. It would be interesting to try removing it and see if there are any regressions. It makes sense to only handle List if Spark only supports i32 indexes.

The difference I think is that a LargeList can store more than Integer.MAX_VALUE entries in all rows in a single batch, so if you have multiple Spark rows all with the max num of rows supported, it wouldn't fit into an Arrow List array. That would probably need to be supported elsewhere, but it may be worth keeping the LargeList handling around in case that scenario is supported? And other DataFusion expressions might return a LargeList even if it doesn't come directly from Spark? Does the native Parquet reader ever use a LargeList?

Thanks for the explanation!
You are right, I will close #1118 then

andygrove

LGTM. Thanks @SemyonSinchenko!

Part of the implementation of array_insert

f583e5c

SemyonSinchenko added 5 commits November 11, 2024 09:11

Missing methods

e870c21

Working version

ac7a2b3

Reformat code

9d9518e

Fix code-style

6e0d5f4

Add comments about spark's implementation.

e4b5e4c

SemyonSinchenko added 2 commits November 13, 2024 12:34

Implement negative indices

19230bf

+ fix tests for spark < 3.4

Fix code-style

58ecb82

SemyonSinchenko changed the title ~~[WIP][DO-NOT-MERGE] feat: support array_insert~~ [WIP] feat: support array_insert Nov 13, 2024

SemyonSinchenko changed the title ~~[WIP] feat: support array_insert~~ feat: support array_insert Nov 13, 2024

Fix scalastyle

a248567

SemyonSinchenko changed the title ~~feat: support array_insert~~ [WIP] feat: support array_insert Nov 13, 2024

SemyonSinchenko added 2 commits November 13, 2024 13:13

Fix tests for spark < 3.4

e4349f5

Fixes & tests

0d38ef0

- added test for the negative index - added test for the legacy spark mode

SemyonSinchenko changed the title ~~[WIP] feat: support array_insert~~ feat: support array_insert Nov 13, 2024

SemyonSinchenko marked this pull request as ready for review November 13, 2024 18:04

andygrove reviewed Nov 13, 2024

View reviewed changes

Comment thread spark/src/test/scala/org/apache/comet/CometExpressionSuite.scala

SemyonSinchenko added 2 commits November 14, 2024 06:55

Use assume(isSpark34Plus) in tests

c7f26f9

Merge remote-tracking branch 'refs/remotes/origin/main'

8b58d8d

Test else-branch & improve coverage

f832cf0

andygrove reviewed Nov 18, 2024

View reviewed changes

Comment thread native/spark-expr/src/list.rs Outdated

andygrove reviewed Nov 18, 2024

View reviewed changes

Comment thread spark/src/test/scala/org/apache/comet/CometExpressionSuite.scala Outdated

Update native/spark-expr/src/list.rs

6e41858

Co-authored-by: Andy Grove <agrove@apache.org>

Merge main + add tests

659ab7a

- fixes; - tests; - comments in the code;

SemyonSinchenko added 2 commits November 19, 2024 19:33

Fix fallback test

4770fce

In one case there is a zero in index and test fails due to spark error

Adjust the behaviour for the NULL case to Spark

e9ef941

SemyonSinchenko closed this Nov 20, 2024

SemyonSinchenko reopened this Nov 20, 2024

SemyonSinchenko requested a review from andygrove November 20, 2024 18:02

andygrove reviewed Nov 20, 2024

View reviewed changes

andygrove approved these changes Nov 20, 2024

View reviewed changes

SemyonSinchenko added 2 commits November 21, 2024 10:25

Move the logic of type checking to the method

6431ad9

Fix code-style

e02d20f

andygrove merged commit 9990b34 into apache:main Nov 22, 2024

SemyonSinchenko deleted the array-insert branch November 23, 2024 09:02

SemyonSinchenko mentioned this pull request Nov 24, 2024

chore: Make list.rs non generic & simplify the code #1118

Closed

andygrove mentioned this pull request Jan 7, 2025

[EPIC] Add support for all array expressions #1042

Closed

21 tasks

Conversation

SemyonSinchenko commented Nov 9, 2024 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Which issue does this PR close?

Rationale for this change

What changes are included in this PR?

How are these changes tested?

Uh oh!

SemyonSinchenko commented Nov 9, 2024

Uh oh!

SemyonSinchenko commented Nov 11, 2024 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

Uh oh!

andygrove commented Nov 13, 2024

Uh oh!

codecov-commenter commented Nov 14, 2024 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Codecov Report

Uh oh!

SemyonSinchenko commented Nov 16, 2024

Uh oh!

Uh oh!

Uh oh!

NoeB commented Nov 19, 2024

Uh oh!

SemyonSinchenko commented Nov 20, 2024

Uh oh!

SemyonSinchenko commented Nov 20, 2024

Uh oh!

andygrove Nov 20, 2024

Choose a reason for hiding this comment

Uh oh!

SemyonSinchenko Nov 21, 2024 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Choose a reason for hiding this comment

Uh oh!

andygrove Nov 22, 2024

Choose a reason for hiding this comment

Uh oh!

Kimahriman Nov 28, 2024 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Choose a reason for hiding this comment

Uh oh!

SemyonSinchenko Nov 28, 2024

Choose a reason for hiding this comment

Uh oh!

andygrove left a comment

Choose a reason for hiding this comment

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

5 participants

SemyonSinchenko commented Nov 9, 2024 •

edited

Loading

SemyonSinchenko commented Nov 11, 2024 •

edited

Loading

codecov-commenter commented Nov 14, 2024 •

edited

Loading

SemyonSinchenko Nov 21, 2024 •

edited

Loading

Kimahriman Nov 28, 2024 •

edited

Loading