Skip to main content

15 posts tagged with "Engineering"

Architecture, system internals, and technical deep dives.

View All Tags

The Streamhouse And Lakestream With Apache Fluss

Giannis Polyzos
PMC member of Apache Fluss
Anton Borisov
PMC member of Apache Fluss

Banner

Most streaming architectures are designed around movement: get changes out of a source, process them, and deliver them somewhere useful. That works well until the same data has to support several different workloads. At that point, the architecture is no longer only about moving events, but also about deciding where a reusable state should live and who should maintain it.

This is the problem behind Streamhouse. It organizes workloads around shared logical tables rather than around a collection of independently maintained destinations.

Lakestream provides the stream–lake storage foundation underneath that model, coordinating fresh streaming data and historical lakehouse data as parts of one logical table.

Tiering Service Deep Dive Part 3: In Production

Giannis Polyzos
PMC member of Apache Fluss

Banner

Part 1 and Part 2 built up everything you need to know about how tiering behaves: the mental model, the dials, the queue dynamics, the scale-out story. This part is about what to do with all of that. What breaks at runtime, and which of those failures self-heal versus need operator action. The design mistakes that look fine on day one but come back to bite you on day two. And the operator's daily view: which five numbers tell you whether tiering is healthy on a Tuesday afternoon, where each one comes from, and why two of them can only come from your Flink-side dashboards.

Tiering Service Deep Dive, 3-parts:

  • Part 1 - The Mental Model: how one tiering round actually works, from timer fire to lake commit.
  • Part 2 - Tuning: per-table dials, multi-table dynamics, and scaling out.
  • Part 3 - In Production: failure modes, design pitfalls, and the dashboard that tells you everything is fine.

Tiering Service Deep Dive Part 2: Tuning

Giannis Polyzos
PMC member of Apache Fluss

Banner

Part 1 built the mental model. What tiering is, who does what, how the round runs end-to-end.

This part adds the dials. Buckets and splits determine how a round parallelizes. Log and PK tables behave so differently on round one that the difference deserves its own treatment. The freshness setting, the one knob most users actually touch, does two different jobs that share the same value. Once a single job is handling many tables, queue position starts to dominate effective freshness more than any per-table setting. And once that happens, you have a deployment-shape decision: stay with one job, or scale out. By the end, you'll know which levers matter most and how to use them.

Tiering Service Deep Dive, 3-parts:

  • Part 1 - The Mental Model: how one tiering round actually works, from timer fire to lake commit.
  • Part 2 - Tuning: per-table dials, multi-table dynamics, and scaling out.
  • Part 3 - In Production: failure modes, design pitfalls, and monitoring.

Tiering Service Deep Dive Part 1: The Mental Model

Giannis Polyzos
PMC member of Apache Fluss

Banner

If you're new to Fluss, the lake-tiering story is one of those topics where every explanation seems to assume you already know how it works. This three-part walkthrough aims to bring some clarity to the confusing parts of the system, and to help you understand how it works in practice.

Part 1 builds the mental model from scratch and by the end of it you'll be able to describe, step by step, what happens between the moment a tiering timer fires and the moment a lake snapshot is committed.

Part 2 and Part 3 take that mental model and add the dials (parallelism, table kinds, freshness, multi-table behavior, scale-out) and then put it into a real production deployment (failures, pitfalls, monitoring).

Tiering Service Deep Dive, 3-parts:

  • Part 1 - The Mental Model: how one tiering round actually works, from timer fire to lake commit.
  • Part 2 - Tuning: per-table dials, multi-table dynamics, and scaling out.
  • Part 3 - In Production: failure modes, design pitfalls, and monitoring.

The Storage Hierarchy: Hot, Remote, and Lake

Giannis Polyzos
PMC member of Apache Fluss

Banner

Apache Fluss stores data in three places: local disk on the tablet server, remote object storage like S3, and the lakehouse. Which place holds which data at any given moment, and what is responsible for moving it between them, is the foundation everything else rests on. Your capacity plan depends on it. Your latency targets depend on it. Your disaster-recovery story depends on it. So does your ability to predict, in advance, that a particular configuration change is going to fill up local disk a week later.

How Apache Fluss Achieves True Pruning in Streaming Storage

Yunhong Zheng
PMC member of Apache Fluss

Banner

TL;DR:

Apache Kafka's "column pruning" is actually pseudo-pruning. All fields still cross the network, and clients discard unwanted ones after the fact. Apache Fluss redesigns the storage format, server-side read path, and write-side batching strategy from the ground up with Arrow IPC columnar storage, zero-copy server-side pruning, and client-side pre-shuffle batching. The result: pruning 90% of columns yields a 10x read throughput improvement, with performance scaling linearly with the pruning ratio.

Why Apache Fluss Chose Rust for Its Multi-Language SDK

Luo Yuxia
PMC member of Apache Fluss
Keith Lee
PMC member of Apache Fluss
Anton Borisov
PMC member of Apache Fluss

Banner

If you maintain a data system that only speaks Java, you will eventually hear from someone who doesn't. A Python team building a feature store. A C++ service that needs sub-millisecond writes. An AI agent that wants to call your system through a tool binding. They all need the same capabilities (writes, reads, lookups) and none of them want to spin up a JVM to get them.

Apache Fluss, streaming storage for real-time analytics and AI, hit this exact inflection point. The Java client works well for Flink-based compute, where the JVM is already the world you live in. But outside that world, asking consumers to run a JVM sidecar just to write a record or look up a key creates friction that compounds across every service, every pipeline, every agent in the stack.

We could have written a separate client for each language. Maintain five copies of the wire protocol, five implementations of the batching logic, five sets of retry semantics and idempotence tracking. That path scales linearly with languages and ends predictably: the Java client gets features first, the Python client gets them six months later with slightly different edge-case behavior, and the C++ client is perpetually "almost done."

We took a different path and tried to leverage the lessons of the great.

What does Apache Fluss mean in the context of AI?

Giannis Polyzos
PMC member of Apache Fluss

The Data Foundation for Real-Time Intelligent Systems​

Apache Fluss (Incubating) started as streaming storage for real-time analytics, built to work closely with stream processors like Apache Flink. Its focus has always been on freshness, efficient analytical access, and continuous data, making fast-changing streams directly usable without forcing them through batch-oriented systems or log-only pipelines.

Over the last year, Fluss has expanded beyond this original framing. You’ll now see it described as streaming storage for real-time analytics and AI. This change reflects how data systems are being used today: more workloads depend on continuously updated data, low-latency access to evolving state, and the ability to reason over context as it changes.

In this context, “AI” does not mean training or serving models inside Fluss. It refers to the class of intelligent systems that rely on fresh features, evolving context, and real-time state to make decisions continuously. Whether those systems use traditional machine learning models, newer AI techniques, or a combination of both, they all depend on the same data foundations.

This shift explains the recent evolution of Apache Fluss. Investments in stateless compute, richer data types with zero-copy schema evolution, and vector support through Lance were driven by a single question:

What does a data foundation need to look like to support real-time intelligent systems reliably at scale?

The rest of this post answers that question. We’ll explain what AI means when viewed through the lens of Apache Fluss, and why a streaming-first foundation for features, context, and state is central to building the next generation of intelligent systems.

Fluss × Iceberg (Part 1): Why Your Lakehouse Isn’t a Streamhouse Yet

Mehul Batra
PMC member of Apache Fluss
Luo Yuxia
PMC member of Apache Fluss

As software and data engineers, we've witnessed Apache Iceberg revolutionize analytical data lakes with ACID transactions, time travel, and schema evolution. Yet when we try to push Iceberg into real-time workloads such as sub-second streaming queries, high-frequency CDC updates, and primary key semantics, we hit fundamental architectural walls. This blog explores how Fluss × Iceberg integration works and delivers a true real-time lakehouse.

Apache Fluss represents a new architectural approach: the Streamhouse for real-time lakehouses. Instead of stitching together separate streaming and batch systems, the Streamhouse unifies them under a single architecture. In this model, Apache Iceberg continues to serve exactly the role it was designed for: a highly efficient, scalable cold storage layer for analytics, while Fluss fills the missing piece: a hot streaming storage layer with sub-second latency, columnar storage, and built-in primary-key semantics.

After working on Fluss–Iceberg lakehouse integration and deploying this architecture at a massive scale, including Alibaba's 3 PB production deployment processing 40 GB/s, we're ready to share the architectural lessons learned. Specifically, why existing systems fall short, how Fluss and Iceberg naturally complement each other, and what this means for finally building true real-time lakehouses.

Banner

Primary Key Tables: Unifying Log and Cache for 🚀 Streaming

Giannis Polyzos
PMC member of Apache Fluss

Modern data platforms have traditionally relied on two foundational components: a log for durable, ordered event storage and a cache for low-latency access. Common architectures include combinations such as Kafka with Redis, or Debezium feeding changes into a key-value store. While these patterns underpin a significant portion of production infrastructure, they also introduce complexity, fragility, and operational overhead.

Apache Fluss (Incubating) addresses this challenge with an elegant solution: Primary Key Tables (PK Tables). These persistent state tables provide the same semantics as running both a log and a cache, without needing two separate systems. Every write produces a durable log entry and an immediately consistent key-value update. Snapshots and log replay guarantee deterministic recovery, while clients benefit from the simplicity of interacting with one system for reads, writes, and queries.

In this post, we will explore how Fluss PK Tables work, why unifying log and cache into a persistent design is a critical advancement, and how this model resolves long-standing challenges of maintaining consistency across multiple systems.

Tiering Service Deep Dive

Yang Guo
Contributor of Apache Fluss

Background​

At the core of Fluss’s Lakehouse architecture sits the Tiering Service: a smart, policy-driven data pipeline that seamlessly bridges your real-time Fluss cluster and your cost-efficient lakehouse storage. It continuously ingests fresh events from the fluss cluster, automatically migrating older or less-frequently accessed data into colder storage tiers without interrupting ongoing queries. By balancing hot, warm, and cold storage according to configurable rules, the Tiering Service ensures that recent data remains instantly queryable while historical records are archived economically.

In this blog post we will take a deep dive and explore how Fluss’s Tiering Service orchestrates data movement, preserves consistency, and empowers scalable, high-performance analytics at optimized costs.

Understanding Partial Updates

Giannis Polyzos
PMC member of Apache Fluss

Banner

Traditional streaming data pipelines often need to join many tables or streams on a primary key to create a wide view. For example, imagine you’re building a real-time recommendation engine for an e-commerce platform. To serve highly personalized recommendations, your system needs a complete 360° view of each user, including: user preferences, past purchases, clickstream behavior, cart activity, product reviews, support tickets, ad impressions, and loyalty status.

That’s at least 8 different data sources, each producing updates independently.

Towards A Unified Streaming & Lakehouse Architecture

Luo Yuxia
PMC member of Apache Fluss

The unification of Lakehouse and streaming storage represents a major trend in the future development of modern data lakes and streaming storage systems. Designed specifically for real-time analytics, Fluss has embraced a unified Streaming and Lakehouse architecture from its inception, enabling seamless integration into existing Lakehouse architectures.

Fluss is designed to address the demands of real-time analytics with the following key capabilities:

  • Real-Time Stream Reading and Writing: Supports millisecond-level end-to-end latency.
  • Columnar Stream: Optimizes storage and query efficiency.
  • Streaming Updates: Enables low-latency updates to data streams.
  • Changelog Generation: Supports changelog generation and consumption.
  • Real-Time Lookup Queries: Facilitates instant lookup queries on primary keys.
  • Streaming & Lakehouse Unification: Seamlessly integrates streaming and lakehouse storage for unified data processing.

Introducing Fluss: Streaming Storage for Real-Time Analytics

Jark Wu
PMC Chair of Apache Fluss

We have discussed the challenges of using Kafka for real-time analytics in our previous blog post. Today, we are excited to introduce Fluss, a cutting-edge streaming storage system designed to power real-time analytics. We are going to explore Fluss's architecture, design principles, key features, and how it addresses the challenges of using Kafka for real-time analytics.

Why Fluss? Top 4 Challenges of Using Kafka for Real-Time Analytics

Jark Wu
PMC Chair of Apache Fluss

The industry is undergoing a clear and significant shift as big data computing transitions from offline to real-time processing. This transition is revolutionizing various sectors, including the E-commerce, automotive networking, finance, and beyond, where real-time data applications are becoming integral to operations. This evolution enables organizations to unlock greater value by leveraging real-time insights to drive business impact and enhance decision-making.