spark
https://github.com/apache/spark
Scala
Apache Spark - A unified analytics engine for large-scale data processing
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Scala not yet supported76 Subscribers
View all SubscribersAdd a CodeTriage badge to spark
Help out
- Issues
- [SPARK-57275][CONNECT] Validate row count after consuming all arrow batches
- [SPARK-43847][PYTHON] Throw structured error when reading Protobuf descriptor file fails
- [SPARK-40437][SS][PYTHON] Support string representation of durationMs in GroupState.setTimeoutDuration
- [SPARK-51579][SQL] Avoid EOFException in CSV Parsing by Appending Line Terminator
- [SPARK-57091][SQL] Add BroadcastNearestByJoinExec to avoid cross-product materialization
- [SPARK-57055][SQL][DOCS] Document non-binary collation gap in DataFrameStatFunctions.bloomFilter
- [SPARK-57052][SS] Add state row format validation to multiGet in RocksDBStateStoreProvider
- [SPARK-57415][SQL] Parquet vectorized reader performance improvements (umbrella)
- [SPARK-55791][PYTHON] Fix pandas-on-Spark equality comparisons under ANSI mode
- [SPARK-56897][SQL] Reduce per-value allocations in DELTA_BYTE_ARRAY Parquet decoder
- Docs
- Scala not yet supported