spark
https://github.com/apache/spark
Scala
Apache Spark - A unified analytics engine for large-scale data processing
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Scala not yet supported76 Subscribers
View all SubscribersAdd a CodeTriage badge to spark
Help out
- Issues
- [SPARK-46164][PYTHON] Add include/exclude parameters to DataFrame.describe in pandas API on Spark
- Fix UNBOUND_SQL_PARAMETER regression in PySpark 4.1.1
- [SPARK-56757][PYTHON] Refactor scalar iterator Pandas UDF worker path
- [SPARK-56795][SQL] Reject column types unsupported by the data source on CREATE TABLE / ALTER TABLE
- Is Spark limited to split the Parquet read granularity by Row Group level only?
- [SPARK-56726][CONNECT] Add Dataset.getNumPartitions to Spark Connect client
- [SPARK-56734][CORE] Optimize RocksDBPersistenceEngine with Column Families and zero-allocation prefix matching
- Predicate Pushdown in Spark Structured Streaming (DataSource V2).
- Support Filter pushdown in Spark Structured Streaming
- Error: Illegal Parquet type: FIXED_LEN_BYTE_ARRAY (UUID)
- Docs
- Scala not yet supported