spark
https://github.com/apache/spark
Scala
Apache Spark - A unified analytics engine for large-scale data processing
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Scala not yet supported76 Subscribers
View all SubscribersAdd a CodeTriage badge to spark
Help out
- Issues
- [WIP][ML] Avoid duplicate StringIndexer skip lookups
- [SPARK-58522][PYTHON] Accept a single tuple of columns in Window and TableArg partitionBy/orderBy
- [SPARK-58536] Fix stale docstring for getDefaultFinalStatus in cluster mode
- [SPARK-58544][SQL] Fix vector distance and norm functions returning wrong results from intermediate float overflow
- [SPARK-36284][CORE][SHUFFLE] Add shuffle checksum support for push-based shuffle
- [SPARK-58551][PYTHON] Python Data Sources Limit Pushdown API
- [SPARK-58556][CONNECT][UI] Show ML cache status in Spark Connect UI
- [SPARK-58559][PYTHON] Package all of sbin in PySpark classic distribution
- [SPARK-58520][SQL] Document dynamic table options in SQL reference
- [SPARK-58523][SQL] Add Catalyst runtime filtering interface for DSv2 scans
- Docs
- Scala not yet supported