Lakestream Overview
Fluss provides the streaming table layer and lakehouse integration for a Lakestream foundation. A Flink tiering job maintains the lake representation, and supported readers use the recorded tiering progress to access data across both layers. For the definitions and diagrams, see Streamhouse and Lakestream.
Metadata, Tiering, and Reads
Three mechanisms keep the layers coordinated:
- Shared metadata identifies the table, defines its schema, and maps its physical representations and progress. Streaming buckets and lake partitions need a defined relationship; their layouts need not be identical.
- Managed tiering maintains the lake representation from streaming data. In Fluss, the Tiering Service runs as an Apache Flink job. It writes and commits lake data, then records the associated streaming progress. This is storage maintenance; application transformations run separately in compute engines.
- Supported union reads use a committed lake representation and its associated streaming progress to read the additional data needed from Fluss. Append-only reads must avoid gaps and duplicate records caused by overlap. Mutable-table reads must also respect versions, updates, and deletes.
A lake-only query observes committed lake data. A Union Read accesses the relevant data across both layers through an integration that understands their coordination. A lake-compatible engine alone does not automatically provide union reads. Read behavior depends on the table type, lake format, engine, and execution mode; the architecture does not imply a simultaneous snapshot across all tables and engines.
Supported Integrations
Application writes enter through supported Fluss table interfaces, while tiering maintains the lake representation. Schema changes must follow the supported Fluss workflow so both representations remain coordinated.
Fluss integrates with Paimon, Iceberg, Hudi, and Lance. Check the documentation for your deployed release and the selected format and engine for supported table types, reads, writes, and schema evolution.
Getting Started
- Try Lakestream: Follow the Lakestream Quickstart using Docker.
- Understand Maintenance: Learn how the Tiering Service maintains the lake representation.
- Choose an Access Path: Explore Union Read and lake-only reads.
- Deploy: Follow Deploying Lakestream.