Usage of "private" Spark APIs
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 20/100
- Issue type
- Bug
- Clarity
- Needs clarification
- Activity status
- Stale
- Tech stack
- scala, spark
- Domain
- data-engineering, distributed-systems
Research direction
Start by reviewing Frameless's use of Spark's org.apache.spark.sql.catalyst.expressions.objects.StaticInvoke and the Spark 2.2 compatibility assumptions described here. Compare the reported Spark 2.2, Spark 2.3, and Databricks runtime behavior; the issue is resolved only when the project has a decided way to handle or document these private API incompatibilities.
Written by the indexing model from the issue text.
Description
Frameless is depending on portions of Spark for which there is no binary compatibility commitment. For example, Frameless uses StaticInvoke, which is part of the org.apache.spark.sql.catalyst.expressions.objects package. If you look at the (bountiful) mima exclusions in Spark, the entire org.apache.spark.sql.catalyst package is not checked for binary compatibility.
I don't really consider this a bug with frameless, but I wanted to at least raise it as a concern as it recently bit us at work.
backstory for those who care
At work we use Databricks runtime 3.5. Databricks claims that this runtime uses Spark 2.2. However, we ran into a bewildering issue with a binary incompatibility between Frameless and the runtime Spark version (related to org.apache.spark.sql.catalyst.expressions.objects.StaticInvoke). After quite a bit of investigation, we realized that the Databricks runtime doesn't actually include Spark 2.2 proper, but a private fork of it that has some incompatible changes. It has a backported change from Spark 2.3 that is incompatible with Spark 2.2 (and the version of Frameless that is built against Spark 2.2). We can work around this particular issue by moving to Spark 2.3 and the Databricks 4.0 runtime, but it's tough to know what other incompatibilities could be lurking in the private forks, and I could envision other people running into similar issues (especially if they can't move to Spark 2.3).
- Dominant language
- Scala
- Stars
- 895
- Forks
- 135
- PR merge metrics
- No merged PRs in 30d
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from typelevel/frameless
-
Scala 3 buildsOpenenhancement feature help wanted
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
Difficulty 3/5 1-2 days Newbie friendliness 35/100
-
enhancement feature
Difficulty 5/5 Over a week Newbie friendliness 25/100
-
Difficulty 5/5 Over a week Newbie friendliness 30/100
All issues in typelevel/frameless
Similar issues
-
area:expressions area:ffi bug good first issue priority:medium
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
apache/datafusion-comet#6592 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
chipsalliance/chisel#5504 ·
Maintainers usually reply within 1 day
-
documentation good first issue
Difficulty 2/5 Half a day Newbie friendliness 75/100
Maintainers usually reply within 1 day
-
Homebrew formula 2.1.26: 'cs completions bash' fails (exit 127) because bin/cs is not executableOpen
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
smart-data-lake/smart-data-lake#1166 ·
Maintainers usually reply within 1 day