spark
https://github.com/apache/spark
Scala
Apache Spark - A unified analytics engine for large-scale data processing
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Scala not yet supported76 Subscribers
View all SubscribersAdd a CodeTriage badge to spark
Help out
- Issues
- [WIP][POC][PYTHON] Optimize PySpark broadcast serialization with Arrow
- [SQL][CONNECT] SparkConnectClient may drop metadata (auth headers) on retried/reattached RPCs if metadata is a non-list iterable
- [SPARK-58518][SQL] Do not duplicate input paths when globbing is disabled
- [SPARK-57785][SQL][CONNECT] harden: spark connect client's reattachment mechanism a... in...
- [WIP][ML] Avoid duplicate StringIndexer skip lookups
- [SPARK-58522][PYTHON] Accept a single tuple of columns in Window and TableArg partitionBy/orderBy
- [SPARK-58536] Fix stale docstring for getDefaultFinalStatus in cluster mode
- [SPARK-36284][CORE][SHUFFLE] Add shuffle checksum support for push-based shuffle
- [SPARK-58520][SQL] Document dynamic table options in SQL reference
- [SPARK-58519][SQL] Document UPDATE and DELETE FROM statements
- Docs
- Scala not yet supported