# Apache Spark

> Unified engine for large-scale data

Distributed processing for SQL, streaming, machine learning and graphs.

- Source: https://github.com/apache/spark
- Homepage: https://spark.apache.org/
- License: Apache-2.0
- Language: Scala
- Stars: 44111
- Forks: 29400
- Contributors: 2402
- Last commit: 2026-10-02
- Purpose: Data & analytics
- Runs on: Library / SDK
- For: Enterprise

## Worth score: 74/100

How much DigGitHub recommends it, from activity, adoption, docs, license and security signals.

- Popularity: 23/25
- Momentum: 10/20
- Maintenance: 15/25
- Community: 15/15
- Readiness: 11/15

## Alternatives

- [scikit-learn](https://diggithub.com/scikit-learn/scikit-learn.md): Machine learning in Python
- [Scrapy](https://diggithub.com/scrapy/scrapy.md): Fast, high-level web crawling framework for Python
- [CCXT](https://diggithub.com/ccxt/ccxt.md): Unified API for 100+ crypto exchanges
- [Prefect](https://diggithub.com/PrefectHQ/prefect.md): Workflow orchestration for Python
- [TradingAgents](https://diggithub.com/TauricResearch/TradingAgents.md): Multi-agent LLM framework for financial trading research
- [DuckDB](https://diggithub.com/duckdb/duckdb.md): In-process analytical SQL database
- [Polars](https://diggithub.com/pola-rs/polars.md): Extremely fast DataFrame engine
- [Streamlit](https://diggithub.com/streamlit/streamlit.md): Build data apps in pure Python

---

Source page: https://diggithub.com/apache/spark
Updated: 2026-10-03
