Skip to main content
Open Source · Apache 2.0

Streaming Storage for Real-Time Analytics & AI

Apache Fluss is an open-source, lakehouse-native streaming storage system. It enables Lakestream: a shared table foundation coordinating fresh streaming data and historical lakehouse data for the Streamhouse architecture.

-- Register Apache Fluss as a Flink catalog
CREATE CATALOG fluss_catalog WITH (
'type' = 'fluss',
'bootstrap.servers' = 'coordinator-server:9123'
);
USE CATALOG fluss_catalog;
-- Create a primary-key table
CREATE TABLE pk_table (
shop_id BIGINT,
user_id BIGINT,
num_orders INT,
PRIMARY KEY (shop_id, user_id) NOT ENFORCED
) WITH ('bucket.num' = '4');
INSERT INTO pk_table VALUES (1234, 1234, 1);
SELECT * FROM pk_table WHERE shop_id = 1234;
Architecture

Unlocking the Streamhouse Architecture

Streamhouse brings streaming, serving, and analytics together through shared tables and independent compute engines. Fluss enables its Lakestream storage foundation by coordinating fresh streaming data with historical lakehouse data.

Explore Streamhouse and Lakestream
Apache Fluss architectureCDC streams, event streams, and AI workloads feed the Fluss hot tier through ingestion. The Fluss Gateway provides HTTP/REST and AI agent access. A Coordinator Server manages metadata, placement, and failover for Tablet Servers with Log Tables and PK Tables, exposing sub-second freshness, a columnar log, changelog streams, and latest state. Lakestream coordinates Fluss and lakehouse storage through shared metadata and a Flink tiering service. Lake formats include Apache Paimon, Apache Iceberg, Apache Hudi, and Lance. Read patterns include incremental streaming reads, batch snapshot and full scans, key-value and prefix lookups, and union reads across hot and cold data in a single query. Query engines can query Fluss directly through supported integrations. Engines shown are Apache Flink, Apache Spark, StarRocks (planned), Apache DataFusion (work in progress), Apache Doris (work in progress), DuckDB (experimental), and Trino (work in progress). Data access patterns include column pruning, partition pruning, and predicate pushdowns.01 · SOURCES02 · APACHE FLUSS · HOT TIER03 · READ PATTERNSCDC StreamsPostgres · MySQLOracle · MongoDBEvent StreamsDevices · WebMobileAI WorkloadsFeatures · EmbeddingsMultimodal · AgentsFluss GatewayHTTP / RESTAI Agent AccessIngestionApache Flink / SparkApache Fluss ClientsFLUSS CLUSTERSub-second freshness · Columnar logChangelog stream · Latest stateCoordinator ServerMetadata · Placement · FailoverTablet ServerNode 01Log TablePK TableTablet ServerNode 02Log TablePK TableTablet ServerNode 03Log TablePK TableTablet ServerNode NLog TablePK TableLAKESTREAMOne logical table · Two freshness layersTiering ServiceFlink job · Compaction & commitsShared metadata · Committed progress04 · LAKEHOUSE NATIVELAKEHOUSE · COLD TIEROpen formats · Long retention · Query-engine nativeApache PaimonApache IcebergApache HudiLanceConsume / QueryStreaming ReadsChangelog stream · IncrementalBatch ReadsSnapshot scan · Full scanLookupsKV lookups · Prefix lookupsUnion ReadHot & Cold Data · Single queryQuery Flussdirectly05 · QUERY ENGINESApache FlinkApache SparkApache DataFusionWIPApache DorisWIPStarRocksPlannedDuckDBExperimentalTrinoWIP06 · DATA ACCESS PATTERNSColumn PruningPartition PruningPredicate Pushdowns
The multiple-systems tax

Shared tables reduce repeated data maintenance.

A common stream can feed several systems that each ingest, reconstruct, and synchronize equivalent data. Streamhouse lets compatible workloads reuse maintained tables on a Lakestream foundation. Compute engines still run transformations; specialized stores remain useful when a workload needs them.

Separate systems · repeated maintenance
Message broker
Kafka, for event transport.
Stream processor
Flink or Spark, for derived features and aggregations.
Online store
Redis or DynamoDB, for sub-millisecond lookup.
Offline store
Iceberg or Parquet on S3, for training and history.
Sync layer
Bespoke pipelines and freshness monitors that drift silently.

Repeated ingestion · Equivalent state · Synchronization work

Streamhouse · shared table foundation
Apache Fluss
Streaming tables and lakehouse integration enabling Lakestream beneath independent engines.
  • Streaming Log · Durable, replayable, offset-ordered streams
  • PK Lookup · Sub-millisecond key/value serving
  • Lakestream · Coordinated streaming and lakehouse layers of one logical table
  • Shared Tables · Reusable results maintained by independent compute engines
  • Multi-Modal · Lance integration for vectors and ML context
  • Audit Trail · Change data feed, replayable by design

Shared tables · Managed tiering · Compatible access paths

Six capability pillars

The benefits, grounded in the architecture.

Apache Fluss enables shared table maintenance and access across streaming and lakehouse storage, with independent engines providing computation, queries, and serving.

Streamhouse Architecture

Reusable tables for streaming, serving, and analytics.

Independent engines and applications share maintained datasets through supported interfaces, reducing repeated ingestion and reconstruction of equivalent data.

Architectural basisShared table schemas, row semantics, and supported access paths.

Lakestream Foundation

Fresh and historical data as one logical table.

Streaming and lakehouse representations remain coordinated through shared metadata, managed tiering, and supported reads across their different freshness layers.

Architectural basisTable metadata, committed tiering progress, and Union Read integrations.

Compute / Storage Separation

Independent engines operating on shared tables.

Engines run transformations and publish reusable result tables. Private execution state, including windows and timers, remains the responsibility of each computation.

Architectural basisSeparate compute and storage services with supported table interfaces.

Columnar Streaming Analytics

Pruning that compounds.

Server-side projection, predicate pushdown, and partition pruning on Arrow-format streams compound into order-of-magnitude I/O and network savings.

Architectural basisARROW log format and the compound pruning stack on the TabletServer.

Feature & Context Stores

Maintained data for ML serving and AI context.

Applications retrieve shared features and context through supported interfaces while owning their retrieval policies, decision logs, and any specialized indexes.

Architectural basisPrimary-key lookups, streaming reads, and lake format integrations.

Ecosystem Openness

Documented formats and supported APIs.

Compatible engines read streams, committed lake data, or both through a supported union integration. Capabilities depend on the engine and lake format.

Architectural basisStreaming APIs, open lake formats, and catalog integrations.

Apache Fluss vs Apache Kafka

Where Streams Meet The Lakehouse

Kafka is the streaming transport. Fluss is the streaming storage. If your need is large-scale stream processing with Flink, real-time analytics, AI/ML, or a sub-second lakehouse, Fluss is the shared streaming storage substrate behind all of them. Read the full breakdown to see which fits your stack.

Community

Built in the open, governed by the ASF.

Apache Fluss is developed openly by a global community of contributors. Join the discussion, file an issue, or send a patch.

Apache 2.0
Open-source license
ASF
Apache Software Foundation governance