Quickstart¶
Get up and running with Spark Connect in minutes. This guide walks you through your first Rust query.
Prerequisites¶
Before you start, ensure you have a running Spark Connect server. See Configuration and Connection for how to start one locally:
The server listens on sc://localhost:15002 by default.
Your First Query¶
Connect to the server and run a simple query:
use spark_connect::SparkSession;
fn main() -> Result<(), Box<dyn std::error::Error>> {
let spark = SparkSession::builder()
.remote("sc://localhost:15002")
.get_or_create()?;
let df = spark.range(10)?;
df.show(20)?;
println!("Count: {}", df.count()?);
Ok(())
}
Filtering and Selection¶
Add a filter and select specific columns:
use spark_connect::functions as f;
let df = spark.range(100)?;
df
.filter(f::col("id").gt(lit(50)))
.select(vec![f::col("id")])
.show(20)?;
Aggregation and Grouping¶
Compute aggregations over grouped data:
use spark_connect::{functions as f, lit};
let df = spark.range(20)?;
df
.with_column("category", f::col("id") % lit(3))
.group_by(vec![f::col("category")])
.agg(vec![
f::sum(f::col("id")).alias("total").expression().clone(),
f::avg(f::col("id")).alias("average").expression().clone(),
])
.show(20)?;
Next Steps¶
- DataFrames - column operations, joins, window functions
- SQL - run SQL queries directly
- Reading and Writing - work with CSV, Parquet, Delta, and other formats
- Configuration and Connection - connect to remote servers, set session config